arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CordisBench:语言模型能否推理动态智能体框架中的组件生命周期?

CordisBench: Can Language Models Reason About Component Lifecycles in Dynamic Agent Harnesses?

Damien Sileo, Dimitri Kachler

arXiv 2609.01600首次发表:更新:

发表机构

Univ. Lille; Inria; CNRS; Centrale Lille; UMR 9189 - CRIStAL(里尔大学; 法国国家信息与自动化研究所; 法国国家科学研究中心; 里尔中央理工学院; UMR 9189 - CRIStAL)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究推出1200题的CordisBench基准,评估语言模型对动态智能体框架组件生命周期的推理能力,发现模型在多交互场景下可靠性下降,额外推理可提升性能但成本较高,且存在可替代的低成本参考语义。

AI 中文摘要

动态智能体框架允许语言模型修改影响自身执行的软件,这种灵活性带来了新的推理负担:本地插件变更可通过依赖关系和清理操作传播。我们推出CordisBench,一个包含1200个问题的生命周期推理基准,它结合了受控形式设置与针对Cordis(管理组件依赖和清理的运行时)执行的程序,要求模型识别受影响组件、预测指定拆除顺序后的状态、确定哪些条件在所有或部分顺序下成立,并选择执行成功的重新配置。在这些任务中,我们以低推理 effort(2、4、8、16、24或32次相关交互)评估了三个面向效率的模型,采用确定性任务特定评分。模型通常能良好处理小型系统,但随着更多交互变得相关,可靠性下降,尤其是在预测最终状态和跨拆除顺序推理时。额外推理 effort 可为部分模型恢复显著收益,但成本不可忽视:在我们的16次交互子集上,GPT-5.6 Luna在中等 effort 下每个问题使用近3000个推理 token。对于这些受控实例,该成本可避免:独立有限参考语义在全部528个可执行问题的所有观察和动作结果上,均与Cordis执行一致。

英文摘要

Dynamic agent harnesses let language models change the software that shapes their own execution. This flexibility brings a new reasoning burden: a local plugin change can propagate through dependencies and cleanup. We introduce CordisBench, a 1,200-question benchmark of this lifecycle reasoning. It combines a controlled formal setting with programs executed against Cordis, a runtime that manages component dependencies and cleanup, and asks models to identify affected components, predict state after a specified teardown order, determine which conditions hold under all or some orders, and choose reconfigurations that succeed when executed. Across these tasks, we evaluate three efficiency-oriented models at low reasoning effort with 2, 4, 8, 16, 24, or 32 relevant interactions, using deterministic task-specific scoring. Models usually handle small systems well but grow less reliable as more interactions become relevant, especially when predicting final state and when reasoning across teardown orders. Additional inference effort recovers marked gains for some models. The cost is nontrivial: on our 16-interaction subset, GPT-5.6 Luna uses nearly 3,000 reasoning tokens per question at medium effort. For these controlled instances, that cost is avoidable: an independent finite reference semantics agrees with Cordis execution on every observation and action outcome used for scoring across all 528 executable questions.

Comments13 pages, 6 figures, 5 tables. Code: https://github.com/sileod/cordis-bench ; Data: https://huggingface.co/datasets/sileod/cordis-bench

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑