No AI summary available for this article.
Why It Matters
We introduce SWE-Serve, a benchmark for evaluating agents on production inference engineering tasks.
Provenance
Discovered via ArXiv and published by ArXiv.
Key Claims
Original description
We introduce SWE-Serve, a benchmark for evaluating agents on production inference engineering tasks. Implementing an inference feature can require coordinating multiple changes across the serving stack, including model support, runtime execution, and public APIs. Existing benchmarks provide limited coverage of production inference engineering: repository-level software engineering benchmarks do not target inference, while general terminal-agent benchmarks include only a few inference tasks. Dedicated inference benchmarks, meanwhile, focus primarily on isolated kernel generation or performance...
Discovered via ArXiv
Research papers and preprints from arXiv.
Publisher: arxiv.org
ID: http://arxiv.org/abs/2609.26777v1 · Indexed about 2 hours ago