RMCW: A Deletion-Robust Watermark Based on Reed--Muller Codes for Language Models
RMCW uses Reed-Muller code structure to make LLM watermarks survive token-deletion and rewriting attacks while retaining clean-text detectability.
The paper proposes Reed-Muller Code Watermarking (RMCW), which injects a secret-keyed Reed-Muller structure into generated text via a keyed vocabulary partition and detects it using local Berlekamp-Welch Reed-Solomon consistency tests. This targets deletion attacks that shift token positions and break alignment between observed tokens and original watermark positions. Experiments on C4 and ELI5 with OPT-1.3B and Llama-3.1-8B-Instruct show RMCW preserves strong clean-text detectability and outperforms or matches baselines under several deletion and rewriting attacks. Code is publicly released on GitHub.
- Injects Reed-Muller structure via secret-keyed vocabulary partition during generation
- Detection tests local subsequences for Reed-Solomon consistency with Berlekamp-Welch
- Outperforms or matches baselines under deletion and rewriting attacks
- Evaluated on C4/ELI5 with OPT-1.3B and Llama-3.1-8B-Instruct; code released
Full article158 words · extracted from arxiv.org · click to collapse
Large Language Model (LLM) watermarking provides a lightweight mechanism for identifying text generated by a specific model, but its robustness remains fragile under post-processing attacks. Deletion attacks are particularly challenging because they shift token positions and break the alignment between observed tokens and their original watermark positions. We propose Reed--Muller Code Watermarking (RMCW), an LLM watermarking method based on Reed--Muller codes. In contrast to global codeword recovery, RMCW searches for surviving local algebraic structure, leveraging the Reed--Solomon consistency induced by affine-line restrictions of Reed--Muller codewords. During generation, RMCW injects a Reed--Muller structure into the sequence via a secret-keyed vocabulary partition. During detection, it maps the given text to keyed vocabulary bins and tests local subsequences for low-degree Reed--Solomon consistency using Berlekamp--Welch tests. Experiments on C4 and ELI5 datasets with OPT-1.3B and Llama-3.1-8B-Instruct show that RMCW preserves strong clean-text detectability and outperforms or matches the baseline methods under several deletion and rewriting attacks. Our code is available at https://github.com/BaichengDanny/RMCW.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2610.02817