Persona Dosing: Calibrated Activation Steering for Graded Trait Control
PersonaDose calibrates activation steering so models express requested persona intensities.
PersonaDose controls persona intensity by specializing a shared, description-conditioned FLAS controller and calibrating its flow time against measured trait expression, without pairing training responses to target intensities. Across Llama-3.1-8B, Qwen3-8B, and Gemma-3-4B, it raises core-trait expression at the Persona Vectors coherence floor of 75 by 33.2, 18.3, and 17.8 points over contrastive activation addition. Calibration-selected settings retain an advantage on held-out questions, though the coherence floor does not hold for every trait. Across seven traits, mean targeting error is 4.7–6.2 points on 14–22 reachable targets out of 28 per model.
- PersonaDose calibrates a shared FLAS controller without paired target intensities.
- Core-trait gains are 33.2, 18.3, and 17.8 points on three models.
- Held-out questions keep an advantage, but coherence does not always hold.
- Mean targeting error is 4.7 to 6.2 points across seven traits.
Full article150 words · extracted from huggingface.co · click to collapse
An activation-steering coefficient sets intervention strength, but requesting a particular degree of persona expression requires a behavioral scale. We study persona dosing: controlling a language model through a trait description and a requested mean intensity. PersonaDose specializes a shared, description-conditioned FLAS controller on persona responses, then calibrates its flow time against measured trait expression. Training responses are not paired with requested target intensities. Across Llama-3.1-8B, Qwen3-8B, and Gemma-3-4B, PersonaDose raises core-trait expression at the Persona Vectors coherence floor of 75 by 33.2, 18.3, and 17.8 points over contrastive activation addition. Calibration-selected settings retain an expression advantage on held-out questions, although the coherence floor does not hold for every trait there. Across seven trained traits, calibrated requests yield mean targeting errors of 4.7-6.2 points over 14-22 calibration-reachable targets out of 28 per model. These results separate the behavioral range learned by a controller from the accuracy of requests within that range.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2609.36388