Through Human Eyes and Machine Eyes: Understanding View Mismatch in Video See-Through Extended Reality
Researchers show Meta Quest 3 screenshots diverge from human vision, creating prompt-injection and privacy risks for XR AI.
Researchers formalize the gap between rectangular screenshots captured by video see-through XR headsets and the smaller, non-rectangular region a wearer can actually see. A pilot measurement on Meta Quest 3 found a clear boundary mismatch. Four case studies show that screenshot-based sensing can feed vision-language models hidden or missing context, creating prompt-injection, privacy-leakage, bias, and reliability failures in AI-integrated XR.
- Meta Quest 3 screenshots are rectangular; the human-visible field is smaller and non-rectangular.
- Authors define co-visible, system-only, and human-only regions between the two views.
- Mismatch can cause prompt injection, privacy leakage, and biased or missing model context.
- Findings come from a pilot boundary measurement and four representative case studies.
Full article179 words · extracted from arxiv.org · click to collapse
Video see-through extended reality (VST XR) systems commonly use headset screenshots or captured frames as proxies for the user's first-person visual context. However, the system-captured view and the user's effective visible field do not necessarily coincide: a screenshot records a rectangular machine-readable frame, whereas the user's effective visible region can be more constrained and non-rectangular. This paper studies this human-system view mismatch in VST XR. We formalize the relationship between the system-captured region and the human-visible region by defining their co-visible, system-only, and human-only regions. \rev{We then conduct a pilot-level boundary measurement on Meta Quest 3, revealing a clear mismatch between the rectangular screenshot frame and the approximate human-visible boundary. Building on this model, we analyze how view mismatch can affect screenshot-based XR sensing and downstream vision-language model tasks. Through four representative case studies, we illustrate potential risks and failure modes including prompt injection, privacy leakage, human-invisible information bias, and missing human-visible information. Our results show that view mismatch is not only a geometric artifact, but can also introduce security, privacy, and reliability concerns for AI-integrated VST XR systems.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.29173