Geometric and Semantic Coupling for Interaction Understanding in 3D Scenes
Segment-Snap combines geometric motion priors and semantic part-handle predictors for 3D interaction understanding.
Segment-Snap couples geometric and semantic cues to describe movable parts, handles, and their motion in 3D scenes. A geometric decoder uses planar and upright priors and predicted handle locations to choose hinge lines, while a joint predictor adds handle candidates and part-based class corrections. On Articulate3D validation, handle guidance raises motion-gated average precision from 13.74% to 40.98% at fixed masks and axes, and full context raises handle AP to 30.99%.
- Segment-Snap links movable parts, handles, and motion without a learned motion regressor.
- Handle guidance raises motion-gated AP from 13.74% to 40.98% on Articulate3D.
- Extra handle candidates lift handle AP from 24.63% to 29.65%; full context reaches 30.99%.
- Each geometric and semantic transfer is applied once, with no iterative feedback.
Full article156 words · extracted from huggingface.co · click to collapse
Interaction understanding in 3D scenes requires a joint description of movable parts, their motion, and the regions through which they can be operated. We present Segment-Snap, which connects these outputs through the physical relationship between parts and handles. Learned predictors identify broad part surfaces and small handles. A geometric decoder uses planar and upright priors to constrain motion, then selects hinge lines using predicted handle locations, without training a motion regressor. Conversely, a joint part-and-handle predictor supplies additional handle candidates, whose motion classes are refined using containing parts. Each information transfer is applied once, without iterative feedback. On Articulate3D validation, handle guidance raises motion-gated AP from 13.74% to 40.98% at fixed masks and axes. Additional handle candidates raise handle AP from 24.63% to 29.65%; part-based class correction adds 0.98 points, and full context reaches 30.99%. Repeated training, learned-decoder controls and paired visualizations establish the benefits and limitations of combining geometric and semantic evidence for interaction understanding.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2609.25247