No AI summary available for this article.
Why It Matters
Reasoning models often generate very long reasoning traces, making inference computationally expensive.
Provenance
Discovered via ArXiv and published by ArXiv.
Key Claims
Original description
Reasoning models often generate very long reasoning traces, making inference computationally expensive. Existing approaches typically improve efficiency either through inference-time early-stopping mechanisms or by explicitly encouraging shorter reasoning during training, for example through reinforcement learning with length penalties. We show that substantial efficiency gains can instead emerge from a different kind of supervision: \textit{confidence}. Using a self-supervised procedure, we fine-tune reasoning models to predict their confidence in the answer at intermediate points along their...
Discovered via ArXiv
Research papers and preprints from arXiv.
Publisher: arxiv.org
ID: http://arxiv.org/abs/2609.31619v1 · Indexed about 20 hours ago