When Temporal Perturbations Act Like Sensor Biases: Label-Free Auditing of Wearable Activity Recognizers
SpectrumAudit shows many wearable activity models lose accuracy mainly from sensor DC offsets, not temporal signal changes.
Researchers introduce SpectrumAudit, a label-free audit of whether wearable activity-recognition models rely on persistent sensor offsets rather than temporal patterns. Across 27 victims from three datasets and three backbones, selected waveforms caused robust-accuracy losses of 2.87 to 40.83 points. The DC component was more damaging than the zero-mean residual on 24 of 27 models and recovered at least 90 percent of the drop on 22 of 27, with all five failures on WISDM. On held-out UTD-MHAD, accuracy fell 13.49 points and macro-F1 11.68 points, versus a 0.66-point gain for matched random changes.
- SpectrumAudit uses phase-randomized, label-sealed stimuli on held-out calibration windows.
- DC offsets outweighed zero-mean residuals on 24 of 27 evaluated models.
- Robust accuracy fell by 2.87 to 40.83 points across three datasets and backbones.
- On UTD-MHAD, losses were 13.49 accuracy points and 11.68 macro-F1 points.
- Code is planned for release after the paper is accepted.
Full article156 words · extracted from arxiv.org · click to collapse
Wearable human-activity recognition (HAR) models operate across sensors, subjects, and backbones, yet a smooth waveform may appear temporal while exploiting a persistent sensor offset primarily. We introduce SpectrumAudit, a label-sealed audit that fits a phase-randomized full-window stimulus on calibration windows from subjects held out from training and testing. After selection, it replays its exact DC projection and budget-constrained zero-mean residual on the same frozen victim without refitting. Across 27 victims from three datasets and three backbones, the selected waveforms cause 2.87-40.83-point three-phase robust accuracy losses. Under this replay budget, DC is more damaging than AC on 24/27 victims and recovers at least 90% of the full drop on 22/27; all 5 failures occur on WISDM. In a held-out UTD-MHAD check, the selected waveform causes 13.49-pp accuracy and 11.68-pp macro-F1 losses, versus -0.66 pp for matched random changes. The audit diagnoses offset versus zero-mean variation under a common peak-budget cap. The code will be released upon acceptance.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.29937