Generative Tutorial: Towards Live Contextualized Visual Instructions for Physical Tasks
Researchers prototype an AR system that generates live, workspace-specific visual instructions for physical tasks.
The paper introduces Generative Tutorial, a framework for live visual instruction that depicts intended actions inside the user's own workspace. A formative review of image and video generation across 15 physical tasks informed an augmented-reality prototype that uses observed context and predicted outcomes. In a 24-participant lab study, the system improved task quality, perceived workspace correspondence, and step-confirmation time versus pre-authored guidance.
- Framework generates goal images and demos inside the user's own workspace
- Formative tests covered image and video generation across 15 physical tasks
- A 24-person study beat pre-authored guidance on quality and confirmation time
- Generation errors and visual resemblance shaped trust and interpretation
Full article149 words · extracted from arxiv.org · click to collapse
Visual instructions for physical tasks are typically authored in one context and followed in another, requiring users to translate demonstrated tools, materials, and spatial relationships into their own environment. We introduce Generative Tutorial, a conceptual framework for live visual instruction that depicts intended outcomes and actions within the user's environment and task flow. A formative evaluation of state-of-the-art image and video generation identifies failures and potential benefits across 15 physical tasks. Drawing on these findings, we build an augmented-reality prototype system that proactively generates goal images and demonstration videos using observed workspace context and predicted visual outcomes of preceding actions. A 24-participant lab study found higher task performance quality, greater perceived workspace correspondence, and shorter step-confirmation intervals with the system than with pre-authored guidance. Qualitative findings highlighted how contextual resemblance shapes trust, how generation errors affect interpretation, and how guidance delivery should adapt to users' needs, informing future designs.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.24955