Tokens Change, Structure Endures: Spectral Watermarking for Generated Speech
Redwing keeps generated-speech watermarks intact through retokenization using a spectral substitution basis.
Researchers propose Redwing, a training-free token-level watermark for generated speech that models retokenization substitutions as a graph and embeds the mark in a Laplacian spectral basis. On the Moshi full-duplex system, eight consecutive Mimi resynthesis passes yield 80.7% true-positive rate at a 1% false-positive rate, versus 8.3% for KGW and at most 7.3% for WMAR. Across three other neural codecs, eight-pass TPR is 77.5-93.0%, and the gains transfer to TTS models at a speech-quality cost close to KGW.
- Builds a substitution graph whose Laplacian assigns similar values to interchangeable tokens.
- Moshi plus eight Mimi passes: 80.7% TPR at 1% FPR, versus 8.3% for KGW.
- Eight-pass TPR is 77.5-93.0% across three other neural codecs.
- TTS quality cost stays close to the KGW baseline.
Full article217 words · extracted from arxiv.org · click to collapse
Watermarking is a promising tool for establishing the provenance of AI-generated speech. While many neural audio watermarking methods rely on a separately trained watermark generator, token-level watermarking is a training-free alternative that operates directly during generation. Its main weakness is retokenization: decoding generated speech to a waveform and encoding it again can change token identities and erode the watermark. To make the watermark robust to these changes, we propose Redwing, REtokenization-Durable Watermarking IN Generation. It builds a graph from the token substitutions observed under retokenization, whose Laplacian yields a basis that assigns similar values to tokens likely to substitute for one another. Over this basis, embedding and detection functions are jointly optimized to preserve watermark signal through retokenization while limiting embedding distortion and detector variability on unwatermarked speech. On the Moshi full-duplex system, after eight consecutive passes of Mimi resynthesis, Redwing achieves 80.7% TPR at a calibrated 1% FPR, compared with 8.3% for KGW and at most 7.3% for WMAR. It also has the highest TPR after eight passes through three other neural codecs (77.5-93.0%), and the gains generalize to TTS models at a speech-quality cost close to that of KGW. These results show that retokenization is not merely a source of noise: its transition structure can be exploited as a design principle for robust token-level watermarking.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.33774