Where Should I Join? Robot Group Joining via Language-Guided Goal Prediction
A language-guided method predicts where a robot should join conversations, queues, and audiences.
The paper formulates language-grounded robot group joining: given an observation and a natural-language description, the robot identifies relevant members and predicts socially compliant joining poses. Candidate subsets are generated by recursive spectral partitioning and ranked with a language-conditioned image-geometry model, after which a goal predictor uses human-formation priors to score feasible poses. Experiments on conversations, queues, and audiences report competitive grounding accuracy, sub-second inference, and better joining-pose prediction than baselines, including real-robot tests in static and changing interactions.
- Robot infers which group to join from vision and a language description.
- Recursive spectral partitioning proposes candidate member subsets for ranking.
- Goal predictor outputs a multimodal energy-orientation map of joining poses.
- Real-robot trials covered static and dynamically changing group interactions.
Full article169 words · extracted from arxiv.org · click to collapse
Social navigation typically assumes a specified goal and focuses on reaching it while respecting social conventions, whereas robot group joining requires predicting where to join based on the group's real-time activity and formation. This is a highly semantic task, yet an important capability for applications such as robotic guide dogs and autonomous mobility scooters. We formulate language-grounded robot group joining: given an observation and a natural-language description of a target group, the robot identifies the relevant group members and predicts socially compliant joining poses. For grounding, we generate structured candidate subsets through recursive spectral partitioning and rank them with a language-conditioned image--geometry model. Given the grounded group, a goal predictor leverages human-formation priors to produce a multimodal energy--orientation map over feasible robot poses. Experiments on conversations, queues, and audiences across varying group sizes, crowd densities, and visual ambiguities show that our method achieves competitive grounding accuracy with sub-second inference and outperforms all baselines in joining-pose prediction. Real-robot experiments further demonstrate group joining in both static and dynamically changing interactions.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.28467