HiPLEX: Hierarchical Policy Factorization for Full Duplex Speech Language Models
HiPLEX factorizes full-duplex speech policies so reinforcement learning can improve timing and content separately.
HiPLEX is a reinforcement-learning framework that factorizes a pretrained full-duplex text policy into a control policy choosing pad, epad, or con, and a conditional content policy that emits a token only when con is selected. Timing advantages use event-causal masks from generated speech, while an LLM-judge semantic advantage trains the content factor through the existing text head. Across three Moshi seeds on Full-Duplex-Bench v1, it reduces takeovers during pauses and backchannels and shortens post-interruption latency versus GRPO with comparable judged quality. On Moshi and PersonaPlex it matches pooled human turn-timing and backchannel-rate marginals better than GRPO.