No AI summary available for this article.
Why It Matters
Whether a language model has improved itself is increasingly judged not by mean accuracy but by which individual problems it gains and loses.
Provenance
Discovered via ArXiv and published by ArXiv.
Key Claims
Original description
Whether a language model has improved itself is increasingly judged not by mean accuracy but by which individual problems it gains and loses. Tracking these transitions means differencing two noisy estimates, leaving them vulnerable to measurement artifacts. Auditing three rounds of rank-$32$ LoRA self-training on Qwen3-8B against a frozen control pushed through the identical pipeline, we identify seven measurement failures, each of which inverts a reported finding when its control is absent. Several are standard practice. A ledger built on a single greedy decode manufactures capability change...
Discovered via ArXiv
Research papers and preprints from arXiv.
Publisher: arxiv.org
ID: http://arxiv.org/abs/2608.20290v1 · Indexed 8 days ago