No AI summary available for this article.
Why It Matters
AI agents increasingly carry out long-horizon professional work, but their evaluations rarely require a finished creative deliverable.
Provenance
Discovered via ArXiv and published by ArXiv.
Key Claims
Original description
AI agents increasingly carry out long-horizon professional work, but their evaluations rarely require a finished creative deliverable. To this end, we introduce Timeline-Bench, a benchmark of 56 real video-editing tasks, each asking an agent to turn raw production material into a finished video. Tasks range from selecting dialog takes and shaping interview footage into a story to cutting commercials from product shots, voiceovers and graphics. Every task provides a brief, source assets, a container and a set of tests. A task is resolved when the output passes every test. The tests check the de...
Discovered via ArXiv
Research papers and preprints from arXiv.
Publisher: arxiv.org
ID: http://arxiv.org/abs/2609.35143v1 · Indexed 44 minutes ago