Your Model Is Leaking: Covert Information Transfer through LLM Residual Streams
A residual-stream covert-channel attack hides sensitive information in LLM intermediate activations, recoverable with a simple linear decoder without retraining or weight modification.
This research describes a residual-stream covert-channel attack that hides sensitive information in LLM intermediate activations, enabling recovery with a simple linear decoder without model retraining or weight modification.
70