Can Pixels Alone Reveal Image Origin? Minimax Limits and Learnable Interfaces for Passive Provenance
Paper proves pixel-only image provenance has a minimax limit and shows public CLIP verifiers fail under edits.
The paper asks whether pixels alone can identify whether an image came from a human, an aggregate AI class, or a specific generator after adversarial edits. It proves that the largest robust target-acceptance gap equals the minimum total-variation distance between the target distribution and attacked source distributions, independent of verifier architecture. If a deployed verifier can be emulated to error epsilon, a surrogate black-box attack approaches the white-box optimum within 2 epsilon plus optimization error. On same-prompt real and diffusion benchmarks, public CLIP verifiers failed under targeted pixel attacks, while a ResNet-18 victim showed partial fake-to-real transfer.
- Robust acceptance gap equals a total-variation distance independent of architecture.
- Emulable public verifiers allow surrogate attacks near the white-box optimum.
- Public CLIP verifiers failed targeted pixel attacks on diffusion benchmarks.
- ResNet-18 showed partial fake-to-real transfer; abstention cut attack success.
Full article234 words · extracted from arxiv.org · click to collapse
Passive image provenance asks whether pixels alone can reveal where an image came from: a human, an aggregate AI class, or a particular generator. This becomes a robustness problem once a source image can be edited before the verifier sees it. We study the problem as source--target verification under adversarial distribution shift. Our first result gives the exact best-case limit for any image-only verifier: the largest robust target-acceptance gap equals the minimum total-variation distance between the target distribution and the set of attacked source distributions. This quantity depends on the source, target, and edit class, not on the verifier architecture. Our second result explains why deployed public verifiers can fail before this statistical limit is reached. If the verifier can be emulated on the attack region to error $\varepsilon$, then a surrogate black-box attack reaches target acceptance within $2\varepsilon$ plus optimization error of the white-box optimum; score-revealing logistic and softmax heads over public features are identifiable, and approximate score access gives stable recovery bounds. A finite-state experiment checks the minimax identity where both sides are computable. On same-prompt real/diffusion benchmarks, the evaluated public CLIP verifiers fail under targeted pixel attacks, while a ResNet-18 victim exhibits partial fake-to-real transfer. Binary feedback with abstention reduces measured attack success, but positive empirical gap upper bounds do not establish robustness. These results motivate separate evaluation of the source--target statistical ceiling and the information released by a deployed verifier.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.30997