Starting-token cues recover much of RL reasoning gains
Cues such as Okay and Alright lift base-model math scores near RL models, and data edits can add or remove those cues.
Both reports cover the same paper: particular tokens fixed at the start of a base model's reply can elicit math and coding performance competitive with reinforcement-learning-trained counterparts. The cue Okay lifts Olmo-3-7B MATH-500 pass@1 from 42% to 78%, and Alright lifts Qwen3-14B from 72% to 87%, recovering much of the RL gain, which itself increases how often those cues appear. The sources agree on those figures. The later arXiv listing adds that causal edits of training data can turn an arbitrary word into a reasoning cue, remove an existing cue's effect, or make a nonsense instruction behave like step-by-step prompting. Both reports say a safety case study finds different cues produce distinct refusal and compliance behaviors corresponding to different training-document types. No material numerical disagreement appears between the 2026-10-04 Hugging Face note and the 2026-10-05 arXiv abstract.
- Cue "Okay" raises Olmo-3-7B MATH-500 pass@1 from 42% to 78%.
- Cue "Alright" raises Qwen3-14B MATH-500 pass@1 from 72% to 87%.
- Fixing those reply-start tokens recovers much of RL's math and coding gain; RL itself raises the likelihood of the cues.
- Causal training-data edits can add or erase a cue, including making a nonsense instruction work like step-by-step prompting.
- A safety case study ties different cues to distinct refusal and compliance behaviors linked to training-document types.
- Hugging Face daily papers dated the item 2026-10-04; the arXiv cs.AI/cs.LG/cs.CL listing is 2026-10-05.
Coverage timelineoldest first · each row is one article
- · 4d agoBase Models Can Reason By Taking a Cue From Training Data
Hugging Face daily papers· 62
The cue Okay lifts Olmo-3-7B MATH-500 pass@1 from 42% to 78%, recovering much of RL's gain.
- · 3d agoBase Models Can Reason By Taking a Cue From Training Data
arXiv cs.AI / cs.LG / cs.CL· 56
Fixed starting-token cues let base models recover much of the math and coding gains from RL training.