The Price of Peeking: Anytime-Valid Leakage Detection on ML-KEM EM Traces
Anytime-valid betting tests detect ML-KEM electromagnetic leakage while controlling false alarms under repeated looks.
The paper studies anytime-valid side-channel leakage detection for evaluators who inspect tests while acquisition is still running. It uses SKIT-type swap-pair e-processes whose false-alarm probability is controlled uniformly over time under a conditional symmetry null. On synthetic streams and a degraded open ML-KEM electromagnetic dataset, first-crossing detection needed about 1.68–2.38 times as many traces as a fixed-horizon randomization test at 80% detection. On undegraded recordings the procedure often stopped after 2–8% of a 4096-trace budget, while repeated-look Welch |t|>4.5 screening false-alarmed in 2.7–12.9% of designed-null replicates. All traces come from one device.
- Anytime-valid e-processes control false alarms uniformly while evaluators monitor tests.
- First-crossing detection needed 1.68–2.38× traces versus fixed-horizon tests at 80% detection.
- On clean recordings, the median stop was 2–8% of a 4096-trace budget.
- Repeated-look |t|>4.5 screening false-alarmed in 2.7–12.9% of null replicates.
- All recordings come from one device; natural-label results are descriptive.
Full article219 words · extracted from arxiv.org · click to collapse
Side-channel evaluators routinely inspect leakage tests while acquisition is still running, and extend or stop the campaign based on what they see. Fixed-horizon screening such as the Welch $t$-test with threshold $|t|>4.5$ gives no error guarantee for this monitored decision rule. We study anytime-valid leakage detection based on testing by betting: SKIT-type swap-pair e-processes whose false-alarm probability is controlled uniformly over time under an explicit conditional symmetry null. In matched comparisons that share the frozen witness, rows and payoff, first-crossing detection needed 1.68-2.00$\times$ the traces of a fixed-horizon randomization test at 80% detection on synthetic streams, and 1.68-2.38$\times$ on degraded recordings from an open ML-KEM electromagnetic dataset with the primary Ridge witness at $α=0.05$. With the same primary witness and level, on undegraded reference and pqm4 recordings the monitored procedure stopped early: its median stopping point was 62-72 and 146-316 evaluation traces, i.e. 2-8% of a conservative 4096-trace budget. Under exact designed nulls on the recorded backgrounds, repeated-look $|t|>4.5$ screening over all 13000-20000 samples raised a false alarm in 2.7-12.9% of replicates, against 0.0-4.7% for terminal-only screening and no rejection by a sample-wise e-Bonferroni process, which in a prespecified follow-up detected natural-label associations in 4 of 4 backgrounds after 840-3288 traces. All recordings come from one device, and natural-label results are descriptive; we state the assumptions each claim requires.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.33597