4D-HOF: Hand-Object Flow Matching for Feed-Forward 4D Interaction Reconstruction
4D-HOF reconstructs hand-object interactions feed-forward by flow-matching foundation-model estimates with physical test-time guidance.
4D-HOF is a feed-forward framework for 4D hand-object interaction reconstruction that starts from coarse states produced by vision foundation models. A conditional flow matching model transports those states toward an interaction manifold, correcting translation, rotation, and alignment without per-sequence optimization. Physical constraints and observed 2D evidence steer the generative trajectory at test time instead of a separate post-hoc optimizer. Trained on diverse data, the method reports state-of-the-art accuracy and stability on out-of-domain benchmarks.
- Feed-forward reconstruction avoids costly per-sequence optimization.
- Conditional flow matching corrects translation, rotation, and alignment errors.
- Test-time guidance uses physics and 2D evidence inside generation.
- Reports state-of-the-art results on out-of-domain benchmarks.
Full article169 words · extracted from arxiv.org · click to collapse
Existing methods for 4D hand-object reconstruction often rely on costly per-sequence optimization, while generative approaches typically synthesize interactions from random noise, which can lead to unstable interaction prediction. We introduce 4D-HOF, a feed-forward framework that reconstructs 4D hand-object interactions from coarse but informative estimates produced by vision foundation models. Concretely, we learn a conditional flow matching model that transports foundation-model-derived hand-object states toward an interaction manifold, allowing the model to correct errors in translation, rotation, and alignment in a feed-forward manner. A key advantage of our generative formulation is that it naturally enables test-time guidance within the transport process. Rather than applying a separate post-hoc optimization after reconstruction, we directly steer the evolving generative states using physical interaction constraints and observed 2D evidence, allowing the reconstruction to be refined as part of the generative process itself. By training the generative model on diverse datasets, 4D-HOF generalizes robustly to challenging in-the-wild scenarios. Experiments on out-of-domain benchmarks show that 4D-HOF achieves state-of-the-art performance, producing more stable and accurate 4D hand-object reconstructions.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2610.08782