arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CTE-Bench:有状态软件模拟器的反事实轨迹评估

CTE-Bench: Counterfactual Trace Evaluation for Stateful Software Simulators

Xinran Zhang

arXiv 2609.36647首次发表:更新:

AI 中文总结

针对编码智能体难以预测有状态软件干预后行为的问题,提出CTE-Bench基准,通过固定未来调用预测与效果步值匹配得分,评估发现当前模型仅在提供正确反馈时才能有效追踪干预效果。

AI 中文摘要

编码智能体会改变正在运行的软件:它们会修补服务的代码或覆盖其存储的状态,然后依据自己对服务后续响应的预期来采取行动。错误的预期可能要到多次调用之后才会显现。函数级代码执行基准忽略了持久化的服务状态,而智能体基准则对智能体采取的动作或最终达到的状态进行评分。我们引入了CTE-Bench,它衡量模型能否预测干预如何改变有状态服务的未来行为,而无需模型选择动作。每个场景向模型提供Python服务代码、干预前观察到的调用和响应、干预本身(源码编辑或状态覆盖)以及40个固定的未来调用;模型预测每个未来的响应,并通过执行服务来检查预测。三种记忆协议控制模型是看到正确的早期响应、看不到任何早期响应,还是看到自己之前的预测。CTE-Bench-Core-v1包含六个确定性Python服务上的255个场景,每个模型产生10,200个预测。主要得分是效果步值匹配(VM):在干预改变的2,476个未来调用上,预测响应与真实响应完全相等。当揭示正确的早期响应时,四个API托管的模型(DeepSeek V4-Flash、Kimi K2.5、Qwen3.6-35B-A3B和Claude Sonnet 4.6)达到54.3%-61.5%的效果步VM。隐藏这些响应会将效果步VM降至23.2%-28.9%;基于自生成的预测进行条件化则得到24.8%-33.2%,且最多只有1.2%的场景被端到端精确预测。因此,当前模型主要是在提供正确反馈时才能追踪干预效果,并且它们的错误在展开过程中会累积。我们发布了CTE-Bench-Core-v1及其可执行预言机、评估脚本,以及一张将每个声明映射到其协议的评估卡。

英文摘要

Coding agents change running software: they patch a service's code or overwrite its stored state, and then act on their own expectation of how the service will respond afterwards. A wrong expectation may surface only several calls later. Function-level code-execution benchmarks omit persistent service state, and agent benchmarks score the actions an agent takes or the final state it reaches. We introduce CTE-Bench, which measures whether a model can predict how an intervention changes a stateful service's future behavior, without asking it to choose actions. Each scenario gives the model Python service code, the calls and responses observed before the intervention, the intervention itself (a source edit or a state overwrite), and 40 fixed future calls; the model predicts every future response, and predictions are checked by executing the service. Three memory protocols control whether the model sees the correct earlier responses, none of them, or its own earlier predictions. CTE-Bench-Core-v1 contains 255 scenarios over six deterministic Python services, giving 10,200 predictions per model. The main score is effect-step value match (VM): exact response equality on the 2,476 future calls whose response the intervention changes. With correct earlier responses revealed, four API-hosted models (DeepSeek V4-Flash, Kimi K2.5, Qwen3.6-35B-A3B, and Claude Sonnet 4.6) reach 54.3%-61.5% effect-step VM. Hiding those responses lowers effect-step VM to 23.2%-28.9%; conditioning on self-generated predictions gives 24.8%-33.2%, and at most 1.2% of scenarios are predicted exactly end to end. Current models thus track intervention effects mainly when correct feedback is supplied, and their errors compound over a rollout. We release CTE-Bench-Core-v1 with its executable oracle, evaluation scripts, and an evaluation card mapping each claim to its protocol.

CommentsAccepted at NeurIPS 2026 (Track on Evaluations and Datasets). Dataset: https://huggingface.co/datasets/zhangxr7/cte-bench-core-v1

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑