MarkSec: Capability-Aware Evaluation of Adversarial Attacks Against LLM Watermarks
MarkSec unifies evaluation of stealing, scrubbing, and spoofing attacks against LLM watermarks with quality-constrained success metrics under shared reporting protocols.
MarkSec is a framework unifying analysis of stealing, scrubbing, and spoofing attacks against LLM watermarks under shared detector calibration, metric definitions, and reporting protocols. It introduces a quality-constrained attack success metric that jointly assesses attack effectiveness and text quality. Experiments across representative watermark families, attacks, LLMs, and datasets show that attacks strongest by watermark removal alone can fall behind general rewriting when success requires acceptable text quality, and stealing-based scrubbers often underperform the best general-scrubbing baselines.
- Unifies stealing, scrubbing, and spoofing watermark attacks under a common protocol
- Quality-constrained success metric jointly measures attack effectiveness and text quality
- General rewriting remains a strong scrubbing baseline across watermark families
- Stealing-based scrubbers often underperform general rewriting when quality is required
Full article201 words · extracted from arxiv.org · click to collapse
LLM watermarking helps trace the origin of generated text, but faces stealing attacks that recover watermark information, scrubbing attacks that remove watermark signals, and spoofing attacks that forge text accepted as watermarked. These attacks are often studied in isolation, leaving their connections unclear. Evaluations also often lack shared detector calibration, metric definitions, and reporting protocols. Moreover, measuring attack success and text quality separately makes it difficult to identify attacks that are both effective and quality-preserving. We propose MarkSec, a general framework that unifies analyses of stealing, scrubbing, and spoofing. We evaluate attacks under a common reporting protocol and introduce a quality-constrained attack success metric to assess effectiveness and text quality jointly. Experiments across representative watermark families, attacks, LLMs, and datasets reveal three findings. First, attacks that appear strongest by watermark removal alone can fall behind general rewriting when success also requires acceptable text quality. Second, general rewriting remains a strong baseline across watermark families, while its advantage over other scrubbers varies by family. Third, in a case study of one watermark family, stealing-based scrubbers often underperform the best general-scrubbing baselines when text quality is required. These results show that apparent attack winners depend on text-quality constraints, attack generality, and capability assumptions.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.16681