Conditional Trajectory Peaks: Single-Pass Multimodal Policies over Action Chunks
Conditional Trajectory Peaks models multimodal action chunks in one pass, matching pi0.5 at lower latency.
Conditional Trajectory Peaks is a single-pass imitation policy that jointly predicts action-chunk candidates, probability masses, and trajectory scales, with DAPS for peak specialization and ETBT for cross-chunk consistency. It reports 91.40% coverage on Push-T, 100.0%, 79.72%, and 84.44% success on D3IL Avoiding, Aligning, and Sorting-2, and 97.25% average success on LIBERO. In real dual-arm tests it succeeded in all 50 two-plate trials and matched pi0.5 success on bottle uprighting and pen placement while cutting inference latency from 218.24 ms to 75.80 ms.
- CTP predicts action-chunk candidates, probability masses, and scales in one pass.
- LIBERO average success was 97.25%; Push-T coverage was 91.40%.
- A real dual-arm two-plate task succeeded in all 50 trials.
- Latency fell from 218.24 ms to 75.80 ms versus pi0.5.
Full article173 words · extracted from huggingface.co · click to collapse
Multimodal imitation learning requires diverse executable futures under the same observation and consistent behavior across replanning cycles. We present Conditional Trajectory Peaks (CTP), a single-pass policy framework that jointly predicts complete action-chunk candidates, probability masses, and trajectory scales. Distribution-Aware Peak Specialization (DAPS) specializes trajectory peaks using trajectory-level posterior responsibilities and mass- and scale-modulated overlap constraints. Evidence-Gated Trajectory Belief Transport (ETBT) maintains cross-chunk consistency through geometric correspondence between exchangeable candidate sets, while allowing current policy evidence to override historical constraints. CTP achieves a coverage score of 91.40% on Push-T; success rates of 100.0%, 79.72%, and 84.44% on D3IL Avoiding, Aligning, and Sorting-2, respectively. On LIBERO, CTP achieves an average success rate of 97.25%. In real-world dual-arm experiments, CTP preserves both placement modes in a two-plate task, succeeding in all 50 trials. On bottle uprighting and pen placement into a holder, it maintains success rates comparable to π_{0.5} while reducing policy inference latency from 218.24 ms to 75.80 ms. These results demonstrate that single-pass trajectory modeling can combine multimodal behavior, closed-loop consistency, and efficient inference.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2610.06104