Base Models Can Reason By Taking a Cue From Training Data
Fixed starting-token cues let base models recover much of the math and coding gains from RL training.
The paper finds that particular tokens at the start of a base model's reply can elicit reasoning competitive with reinforcement-learning-trained counterparts. One cue lifts Olmo-3-7B MATH-500 pass@1 from 42% to 78%, and another lifts Qwen3-14B from 72% to 87%. Causal edits of training data can turn an arbitrary word into a reasoning cue or remove an existing cue's effect, and can make a nonsense instruction work like step-by-step prompting. A safety case study shows different cues produce distinct refusal and compliance behavior corresponding to different training-document types.
- Starting-token cues can raise base-model math accuracy close to RL models.
- Olmo-3-7B MATH-500 pass@1 rises from 42% to 78% with one cue.
- Qwen3-14B rises from 72% to 87% when prompted with Alright.
- Training-data edits can add or erase a cue, including a nonsense instruction.
- Different cues also shift refusal versus compliance in a safety case study.
Full article209 words · extracted from arxiv.org · click to collapse
In this paper, we study how training data creates associations between the tokens at the start of a base model's response and the reasoning behavior that follows. First, we demonstrate that fixing particular starting token cues makes a base model's performance competitive with that of its reinforcement learning (RL)-trained counterparts on math and coding. For instance, the cue ".\n\nOkay" raises Olmo-3-7B's MATH-500 pass@1 accuracy from 42% to 78%, while "Alright," raises Qwen3-14B's from 72% to 87%. Second, RL makes these cues more likely, while fixing them recovers much of its performance gain over the base model. Third, we trace the reasoning effects of token cues to the training data. We perform causal data interventions to turn an arbitrary word, such as "chicken", into an effective reasoning cue, or remove an existing cue's effect. A similar edit makes the prompt instruction "Think duck duck goose" as effective as "Think step by step" at eliciting reasoning. We also find that the hidden state representations induced by different cues correlate with different document types from the training set. Finally, we extend our study of token cues with a case study in language model safety, finding that different cues elicit distinct refusal and compliance behaviors that correspond to different types of training data.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2610.06851