No AI summary available for this article.
Why It Matters
Large language models (LLMs) are increasingly used in coding tasks, but their ability to reason about code execution remains unclear.
Provenance
Discovered via ArXiv and published by ArXiv.
Key Claims
Original description
Large language models (LLMs) are increasingly used in coding tasks, but their ability to reason about code execution remains unclear. Existing repository-level QA benchmarks mainly evaluate static code understanding and often rely on LLM-based evaluation, while execution-reasoning benchmarks are mostly limited to snippets or functions. We introduce SWE-Flux, a repository-level benchmark for dynamic execution reasoning containing 480 execution-grounded instances across 12 real Python repositories, with gold answers automatically harvested from instrumented test executions rather than written ma...
Discovered via ArXiv
Research papers and preprints from arXiv.
Publisher: arxiv.org
ID: http://arxiv.org/abs/2609.28449v1 · Indexed about 1 hour ago