Underwater C3-JEPA: An Object-Centric Cross-View World Model for ROV Salvage
Researchers introduce Underwater C3-JEPA, a control-conditioned world model for near-field ROV salvage.
The paper introduces Underwater C3-JEPA, an object-centric multi-view world model for near-field heavy-load underwater ROV salvage. From synchronized RGB views and vehicle controls, it predicts latent task-object state through contact and hydrodynamic lag without contact sensors. Cross-camera attention, weak binding, and SIGReg shape a representation the authors say transfers more task-relevant information than a reconstruction-free latent baseline. Validation on real underwater video recovered a withheld camera's object state and stayed ahead of persistence, supporting MPC evaluation and imagined-rollout training.
- C3-JEPA predicts task-object state from multi-view RGB and vehicle controls.
- It uses cross-camera attention, weak binding, and SIGReg without contact sensors.
- Features transfer more task information than a reconstruction-free latent baseline.
- The interface supports MPC evaluation and imagined-rollout agent training.
- Real underwater video recovered a withheld camera's object state ahead of persistence.
Full article155 words · extracted from arxiv.org · click to collapse
We present Underwater C$^{3}$-JEPA (cross-view, control-conditioned, context-extended), an object-centric multi-view predictive world model for near-field heavy-load underwater ROV salvage. Without contact sensors, it predicts in latent space how the task-object state evolves through contact interaction and under the hydrodynamic lag of the vehicle, from synchronized multi-view RGB observations and vehicle control signals. C$^{3}$-JEPA encodes multi-camera observations into task-object and context tokens, fuses cross-camera evidence through held-out-view attention, and directly predicts future states conditioned on control. Weak binding anchors the target and gripper at low annotation cost, while SIGReg sharpens the geometric representation. Experiments show that the learned representation transfers substantially more task-relevant information to downstream probes than a reconstruction-free latent baseline, while keeping the predictor lightweight. The resulting predictive interface supports model-predictive-control (MPC) candidate evaluation and imagined-rollout behavior-agent training. Validation on real underwater video shows the same architecture recovering a withheld camera's object state and staying ahead of persistence, so the recipe transfers beyond simulation.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.30214