GAZEleak: Passcode Inference Against Eye-tracking XR Devices Through External Observation
GAZEleak recovers XR gaze passcodes from room video of head motion without touching the device.
GAZEleak is a physical side-channel that infers gaze-entered six-digit passcodes on mixed-reality headsets such as Apple Vision Pro from external video of head motion alone. On a preliminary front-view set from three aware author-subjects, the true code was in the top ten guesses for 10 of 18 codes under a cross-person protocol. The least exposed subject never reached the top ten, though the median guessed rank was 12,786 versus 500,000 for a random ordering.
- Adversary only films the wearer; no software is installed on the headset.
- Gaze shifts produce sub-degree head motions recoverable from commodity video.
- True six-digit code ranked in the top ten for 10 of 18 tests.
- Success is strongly subject-dependent, and the evaluation is preliminary.
Full article250 words · extracted from arxiv.org · click to collapse
Mixed-reality headsets such as Apple Vision Pro replace the touch screen with gaze-as-pointer interaction: the wearer looks at a target and confirms with an air pinch. Because the display is inside the headset and the eye tracker is walled off from third-party software, such input is widely assumed to be unobservable to bystanders---a built-in defense against the shoulder-surfing that plagues phones and laptops. We present GAZEleak, a side-channel attack that recovers gaze-driven input from external video of head motion alone, under a strictly local, physical-observer threat model: the adversary only films the wearer from across the room and installs no software on the device. The attack exploits the centrally coupled eye-head motor program: gaze shifts recruit small, target-dependent head reorientations that project into sub-degree pose changes recoverable from commodity video. GAZEleak implements a measurement-based, sparse-optical-flow inference pipeline for users' 6-digit device passcodes. On a preliminary front-view dataset from three author-subjects who were aware of the attack hypothesis, GAZEleak places the true code within the top ten guesses for 56% (10 of 18) of test codes under a cross-person protocol with no labeled victim data, and for every code of the most exposed subject. The performance is subject-dependent: no passcode from the least exposed subject reaches the top ten, although its median guessed passcode rank is 12,786 rather than 500,000 expected from an uninformative ordering. These results provide preliminary evidence that gaze-coupled head motion can expose passcode information under controlled conditions, while motivating broader evaluation across users, behaviors, and capture settings.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.35040