arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ArrivalBench:智能体生成的数据管道一次正确,随时间推移出错

ArrivalBench: Agent-Generated Data Pipelines Are Correct Once and Wrong Under Time

Pranay Kothari

arXiv 2610.02363首次发表:更新:

发表机构

University of Oxford(牛津大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

ArrivalBench通过对抗性可重放调度重新执行智能体管道,以批量重算为预言机,发现单次执行认证的管道在重放下有7.0-79.2%静默出错,并区分错误与崩溃以改进干预。

AI 中文摘要

针对智能体生成数据的基准测试通过针对固定快照运行一次来对管道进行评分。ArrivalBench 则相反,它在对抗性但可重放的交付调度(延迟、重复、乱序和重试的记录)下重新执行智能体留下的管道,并要求其最终状态等于对完整日志的批量重新计算。由于预言机是重新计算而非分类,因此错误的表格和崩溃是不同的判定:崩溃对团队已经运行的监控是可见的,而错误的表格则不可见。在我们构建的40个任务上,我们对单次执行评分的重新实现认证了11个模型产生的管道的86-100%;重新执行相同的工件发现其中7.0-79.2%的已认证管道静默出错。这一差距并非由修复循环产生:在同一模型和任务内,针对快照测试修复的管道在重放中失败的概率与首次通过该测试的管道大致相同。在每个模型中,幂等性风险比顺序性风险更常失败。将错误答案与崩溃区分开来也改变了干预措施的解读方式:一个风险警告将一个模型的静默失败率从48.2%降至10.5%,同时将其崩溃率从9.0%提高到37.0%,因此总失败率仅从51.0%变为44.0%。所有11个分支均独立重新运行,比率变化最多为5.9个百分点。

英文摘要

Benchmarks for agent-generated data work grade a pipeline by running it once against a fixed snapshot. ArrivalBench instead re-executes the pipeline an agent leaves behind under adversarial but replayable delivery schedules (late, duplicated, out-of-order and retried records) and requires its final state to equal a batch recomputation of the complete log. Because the oracle recomputes rather than classifies, a wrong table and a crash are distinct verdicts: a crash is visible to monitoring a team already runs, and a wrong table is not. On 40 tasks we built, our reimplementation of single-execution grading certifies 86-100% of the pipelines eleven models produce; re-executing the same artifacts finds 7.0-79.2% of the certified ones silently wrong. The gap is not produced by the repair loop: within the same model and task, pipelines repaired against the snapshot test fail replay about as often as those that passed it first time. In every model, idempotency hazards fail more often than ordering hazards. Separating a wrong answer from a crash also changes how interventions read: a hazard warning cuts one model's silent failure from 48.2% to 10.5% while raising its crash rate from 9.0% to 37.0%, so all-in failure moves only from 51.0% to 44.0%. All eleven arms were independently re-run, and rates moved by at most 5.9 points.

CommentsAccepted as a poster at the NeurIPS 2026 Workshop "Who Verifies the Agents?". 20 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑