No AI summary available for this article.
Why It Matters
In this paper, we study how training data creates associations between the tokens at the start of a base model's response and the reasoning behavior that follows.
Provenance
Discovered via ArXiv and published by ArXiv.
Key Claims
Original description
In this paper, we study how training data creates associations between the tokens at the start of a base model's response and the reasoning behavior that follows. First, we demonstrate that fixing particular starting token cues makes a base model's performance competitive with that of its reinforcement learning (RL)-trained counterparts on math and coding. For instance, the cue ".\n\nOkay" raises Olmo-3-7B's MATH-500 pass@1 accuracy from 42% to 78%, while "Alright," raises Qwen3-14B's from 72% to 87%. Second, RL makes these cues more likely, while fixing them recovers much of its performance g...
Discovered via ArXiv
Research papers and preprints from arXiv.
Publisher: arxiv.org
ID: http://arxiv.org/abs/2610.06851v1 · Indexed about 1 hour ago