arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.35143cs.CVcs.AIcs.MM

Timeline-Bench:从原始素材到成片,评估智能体在真实视频编辑任务上的表现

Timeline-Bench: Evaluating Agents on Realistic Video-Editing Tasks, from Raw Footage to Final Cut

Gunin Gupta, Nirmit Arora, Pavan Kalyan Tankala

首次发表
浏览论文内容

中文总结 AI 辅助

提出Timeline-Bench基准,含56个真实视频编辑任务,评估AI智能体从原始素材到成片的完成能力;最佳智能体仅解决26.8%任务,多数失败源于质量测试,凸显智能体在创意工艺上的不足。

中文摘要 AI 辅助

AI智能体越来越多地承担长时程的专业工作,但对其评估很少要求产出完成的创意作品。为此,我们提出Timeline-Bench,一个包含56个真实视频编辑任务的基准,每个任务要求智能体将原始制作素材转化为成片。任务范围从选择对话片段、将采访素材塑造成故事,到从产品镜头、配音和图形中剪辑广告。每个任务提供简报、源素材、容器和一组测试。当输出通过所有测试时,任务即视为完成。测试检查交付格式、内容及简报的明确要求,并包含一项质量测试,该测试基于43位视频编辑者对2582次盲审的判断进行校准。我们评估了16个智能体,它们将前沿模型与编码智能体框架(如Codex、Claude Code和OpenCode)配对。表现最好的GPT-6 Astra(在Codex中,并配有策划的编辑指导)仅解决了56个任务中的15个(26.8%),平均智能体解决率为14.0%。人类编辑者在83.5%的判断中偏好参考剪辑。大多数未解决的运行(771次中的562次)仅未通过质量测试:智能体通过静态图像和转录文本感知素材,并检查其渲染是否有缺陷,而非工艺水平。我们在此https URL发布任务、验证器和每次运行的结果。

英文摘要

AI agents increasingly carry out long-horizon professional work, but their evaluations rarely require a finished creative deliverable. To this end, we introduce Timeline-Bench, a benchmark of 56 real video-editing tasks, each asking an agent to turn raw production material into a finished video. Tasks range from selecting dialog takes and shaping interview footage into a story to cutting commercials from product shots, voiceovers and graphics. Every task provides a brief, source assets, a container and a set of tests. A task is resolved when the output passes every test. The tests check the delivery format, the content and the brief's explicit requirements, and include a quality test calibrated on 2,582 blind judgments by 43 video editors. We evaluate 16 agents that pair frontier models with coding-agent harnesses such as Codex, Claude Code and OpenCode. The best, GPT-6 Astra in Codex with curated editorial guidance, resolves only 15 of the 56 tasks (26.8%), and the average agent resolves 14.0%. Human editors prefer the reference edit in 83.5% of judgments. Most unresolved runs (562 of 771) fail only the quality test: agents perceive footage through stills and transcripts and check their renders for defects, not craft. We release the tasks, verifier and per-run results at https://timelinebench.tensortest.com.

发表机构

  • TensorTest (Ritivel Labs Inc.)(TensorTest(Ritivel Labs Inc.))

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑