Can Pixels Alone Reveal Image Origin? Minimax Limits and Learnable Interfaces for Passive Provenance
Paper proves pixel-only image provenance has a minimax limit and shows public CLIP verifiers fail under edits.
The paper asks whether pixels alone can identify whether an image came from a human, an aggregate AI class, or a specific generator after adversarial edits. It proves that the largest robust target-acceptance gap equals the minimum total-variation distance between the target distribution and attacked source distributions, independent of verifier architecture. If a deployed verifier can be emulated to error epsilon, a surrogate black-box attack approaches the white-box optimum within 2 epsilon plus optimization error. On same-prompt real and diffusion benchmarks, public CLIP verifiers failed under targeted pixel attacks, while a ResNet-18 victim showed partial fake-to-real transfer.