HiPLEX: Hierarchical Policy Factorization for Full Duplex Speech Language Models
HiPLEX factorizes full-duplex speech policies so reinforcement learning can improve timing and content separately.
HiPLEX is a reinforcement-learning framework that factorizes a pretrained full-duplex text policy into a control policy choosing pad, epad, or con, and a conditional content policy that emits a token only when con is selected. Timing advantages use event-causal masks from generated speech, while an LLM-judge semantic advantage trains the content factor through the existing text head. Across three Moshi seeds on Full-Duplex-Bench v1, it reduces takeovers during pauses and backchannels and shortens post-interruption latency versus GRPO with comparable judged quality. On Moshi and PersonaPlex it matches pooled human turn-timing and backchannel-rate marginals better than GRPO.
- Splits a full-duplex policy into when-to-speak and what-to-say factors
- Routes timing and semantic advantages to separate policy factors
- Lowers pause and backchannel takeovers versus GRPO on Moshi
- Better matches human turn-timing and backchannel rates than GRPO
Full article217 words · extracted from huggingface.co · click to collapse
As human--AI interactions become more conversational, full-duplex speech language models capable of natural real-time dialogue are growing in importance. Beyond generating appropriate responses, these models must coordinate turn-taking, backchanneling, and floor management in real time. Reinforcement learning (RL) provides a way to refine these behaviors through direct feedback on interaction outcomes. However, existing RL methods either apply timing feedback to a token policy or optimize semantic content, leaving the joint improvement of timing and content unresolved. We introduce HiPLEX, an RL framework that factorizes a pretrained full-duplex text policy into a control policy that decides when to emit content and a conditional content policy that decides what to emit. The first factor selects among 'pad', 'epad', and 'con'. The second selects a token only when 'con' is chosen. This hierarchy describes conditional actions within each frame and uses the model's existing text head. We route timing advantages to the token-group factor through event-causal masks derived from generated speech episodes, and route an LLM-judge semantic advantage to the conditional content factor. Across three Moshi seeds on Full-Duplex-Bench v1, HiPLEX reduces takeover rates during natural user pauses and backchannel opportunities, and shortens post-interruption response latency relative to GRPO, while maintaining comparable judged interruption-response quality. On Moshi and PersonaPlex, HiPLEX better matches pooled human turn-timing and backchannel-rate marginals than GRPO.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2610.07727