EgoTools: Towards Tool-Centric Reasoning in Real-World Egocentric Videos
EgoTools adds 100 hours of egocentric tool-use video and a 1,000-question benchmark where current VLMs still struggle.
EgoTools is a suite for egocentric tool-use understanding. EgoTools-Data provides 100 hours of tool-centric recordings with synchronized audio, dense captions, reasoning-heavy narrations, and supplementary 3D information, while EgoTools-Bench has 1,000 QA pairs across perception, geometry, procedure, and causal reasoning. Gemini-3.1-Pro reaches 66.9% overall accuracy but only 51.7% on Perception and Grounding. Full supervised fine-tuning on EgoTools-Data raises Qwen3-VL-8B-Instruct from 50.0% to 60.9% on the full benchmark under strict source-video separation.
- 100 hours of tool-centric egocentric video with captions, audio, and 3D data
- 1,000 QA pairs covering perception, geometry, procedure, and causation
- Gemini-3.1-Pro scores 66.9% overall but 51.7% on grounding
- Fine-tuning lifts Qwen3-VL-8B-Instruct from 50.0% to 60.9%
Full article243 words · extracted from huggingface.co · click to collapse
Real-world embodied tasks, from everyday activities to professional procedures, require agents to act under physical constraints while tracking evolving object and task states. Tool use sits at the heart of such tasks, as many everyday and professional activities are tool-mediated. Understanding them requires reasoning about affordances, hand-tool-object geometry, procedural progress, and causal effects on target objects. Yet despite strong performance on perception-oriented video tasks such as captioning and general video QA, current multimodal video models remain limited in this form of tool-centric embodied reasoning. Progress in this direction has been limited by the lack of real-world egocentric data and diagnostic benchmarks. To address this gap, we introduce EgoTools, the first comprehensive suite for egocentric tool-use understanding. It consists of two complementary components: EgoTools-Data, a large-scale corpus of 100 hours of tool-centric egocentric recordings with synchronized audio, dense captions, reasoning-heavy narrations, and supplementary 3D information; and EgoTools-Bench, a diagnostic benchmark of 1,000 QA pairs across four tracks that cover tool-use understanding from perception and geometry to procedure and causal reasoning. Experimental results show that current models still struggle to ground tool use in visual evidence: Gemini-3.1-Pro achieves 66.9% overall accuracy but only 51.7% on Perception & Grounding. Beyond evaluation, we validate EgoTools-Data as a training resource. On the full 1,000-question benchmark, full supervised fine-tuning improves Qwen3-VL-8B-Instruct from 50.0% to 60.9%, under strict source-video separation. Together, these results establish EgoTools as a unified resource for both training and diagnostic evaluation of real-world egocentric tool-use understanding.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2609.39378