Tokens Change, Structure Endures: Spectral Watermarking for Generated Speech
Redwing keeps generated-speech watermarks intact through retokenization using a spectral substitution basis.
Researchers propose Redwing, a training-free token-level watermark for generated speech that models retokenization substitutions as a graph and embeds the mark in a Laplacian spectral basis. On the Moshi full-duplex system, eight consecutive Mimi resynthesis passes yield 80.7% true-positive rate at a 1% false-positive rate, versus 8.3% for KGW and at most 7.3% for WMAR. Across three other neural codecs, eight-pass TPR is 77.5-93.0%, and the gains transfer to TTS models at a speech-quality cost close to KGW.
45