MAD-Guard: Controlled Study of Autoregressive Generation versus Direct Decision Interfaces for Closed Multimodal Forensic Tasks
MAD-Guard shows direct decision heads outperform autoregressive decoding for closed multimodal forgery forensics.
MAD-Guard compares autoregressive generation with direct decision heads for closed multimodal forensic tasks on a matched Qwen3-VL-8B backbone, using 2,400 FakeClue samples and LoRA on Huawei Ascend 910C NPUs. A binary direct head reduces latency by 2.60x-7.27x to 53.12 ms and cuts expected calibration error from 0.0845 to 0.0450, with accuracy 93.10% versus 94.90% for AR-SFT. CLM-Head reaches 96.55% binary accuracy and 98.79% seven-class attribution. On 5,000 out-of-sample images it scores 96.44% on GenImage and 97.73% on Chameleon, but FF++ compressed-face ROC-AUC is only 0.5913.
- Matched Qwen3-VL-8B study uses 2,400 FakeClue samples and LoRA on Ascend 910C.
- Binary direct head is 2.60x-7.27x faster, with ECE 0.0450 versus 0.0845.
- CLM-Head hits 96.55% binary accuracy and 98.79% seven-class attribution.
- GenImage 96.44% and Chameleon 97.73%, but FF++ ROC-AUC is 0.5913.
Full article240 words · extracted from arxiv.org · click to collapse
When should multimodal foundation models generate tokens, and when should they directly output a decision? We present MAD-Guard, a controlled study of output-decision interfaces for closed multimodal forensic tasks. Once a multimodal representation is computed, is autoregressive generation necessary for closed forensic decisions with high input complexity but low output entropy? Under a matched Qwen3-VL-8B backbone, 2,400 FakeClue training samples, and LoRA budget ($r=16, α=32$) on Huawei Ascend 910C NPUs, we evaluate a progression of decision interfaces (AR-SFT [generate] $\to$ Logit Slice $\to$ Binary Direct Head $\to$ +choice $\to$ +act $\to$ CLM-Head) and decompose latency into backbone representation (53.12 ms), 151,643-way vocabulary projection (+85.04 ms $\to$ 138.16 ms), and decoding (+248.26 ms $\to$ 386.42 ms). Under 1-to-1 binary supervision ($\mathcal{L}_{\mathrm{BCE}}$), a Binary Direct Head cuts latency by $2.60\times$-$7.27\times$ (53.12 ms) and lowers calibration error by $1.88\times$ (ECE = 0.0450 vs. 0.0845), with a -1.80% accuracy trade-off (93.10% vs. 94.90%; 0.9795 vs. 0.9871 ROC-AUC) from forfeiting token priors. Gains above AR-SFT arise either from multi-task attribution and uncertainty gating (+choice+act: 96.44% accuracy, 0.9940 ROC-AUC, 0.0187 ECE at 53.71 ms) or from a disaggregated contrastive head (CLM-Head: 96.55% binary and 96.44% multi-task accuracy, 0.0166 ECE, 98.79% 7-class attribution at 54.42 ms) retaining semantic priors without token decoding. Across 5,000 out-of-sample images from five benchmarks, our framework excels on synthetic, camouflage, and document forgeries (96.44% GenImage, 97.73% Chameleon, 91.84% Doc) while showing a clear boundary on compressed face manipulation (FF++ ROC-AUC = 0.5913).
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.33683