arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

校准大语言模型智能体中的准则修订:失败模式与基于轨迹锚定的协议

Calibrating Criterion Revision in LLM Agents: Failure Modes and a Trace-Anchored Protocol

Guodong Xu

arXiv 2608.20729首次发表:更新:

发表机构

Qingdao Guodongxiansheng Network Technology Co., Ltd.(青岛郭东先生网络科技有限公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对大语言模型智能体的准则修订归因问题,分析CMB-0.1测试的失败模式,提出基于轨迹锚定的CMB-0.4协议,为准则修订测试提供更具区分度的方案。

AI 中文摘要

语言模型智能体在失败后可以改进,或在多个回合中传递文本而不修订成功的判定准则。我们研究更狭义的准则修订归因问题:当准则K0接受了违反更宽泛承诺B的结果时,哪些观察结果能证明系统形成并持续使用了K1?我们要求满足五个非补偿性条件:准则失败检测、模型生成的提议、新回合传递、对所声称载体的干预敏感性以及保存。我们在12个跨域案例和4个分支(无状态推理、仅追加历史、模型生成但由框架提交的状态、评估器编写的神谕状态)上评估CMB-0.1。7个机制固定装置产生84次确定性评分器试验;4个局部量化伪影产生96次调用和192次模型-案例-分支试验。没有任何模型试验满足全部五个条件,但这一零结果并不证明普遍能力缺失。11次调用在一次重试后仍然无效;若干承诺揭示了目标差异;框架执行提交;删除操作复用了无状态调用;冲突改变了多个因素。Qwen2.5-7B在没有修订状态的情况下回答了所有传递和保存项,暴露了零状态重构问题。这些失败使CMB-0.1成为仪器校准结果而非模型排名。我们提出了一种前瞻性的、基于轨迹锚定的CMB-0.4协议,要求隐藏式传递、明确的WRITE/NO-WRITE/ESCALATE操作、单独记录的策略选择提交、匹配的干预措施、重复的隐藏项以及冻结的可执行神谕。这是一种后续设计,而非完整的验证性结果。本文贡献了一条测量链、对其首个实现的经验诊断,以及一个用于准则修订未来测试的更具区分度的协议。

英文摘要

Language-model agents can improve after failure or carry text across episodes without revising what counts as success. We study the narrower attribution problem of criterion revision: when criterion K0 accepts an outcome violating a broader commitment B, what observations justify saying that the system formed and persistently used K1? We require five non-compensatory conditions: criterion-failure detection, a model-emitted proposal, new-episode transfer, intervention sensitivity on the claimed carrier, and preservation. We evaluate CMB-0.1 on twelve cross-domain cases and four arms: stateless inference, append-only history, model-generated but harness-committed state, and evaluator-written oracle state. Seven mechanism fixtures yield 84 deterministic scorer trials; four local quantized artifacts yield 96 calls and 192 model-case-arm trials. No model trial satisfies all five conditions, but this zero does not establish general capability absence. Eleven calls remain invalid after one retry; several commitments disclose the target distinction; the harness performs commits; deletion reuses a stateless call; and conflict changes multiple factors. Qwen2.5-7B answers every transfer and preservation item without revision state, exposing zero-state reconstruction. These failures make CMB-0.1 an instrument-calibration result rather than a model ranking. We derive a prospective, trace-anchored CMB-0.4 protocol requiring concealed transfer, explicit WRITE/NO-WRITE/ESCALATE actions, a separately logged policy-selected commit, matched interventions, repeated hidden items, and a frozen executable oracle. It is a successor design, not a completed confirmatory result. The paper contributes a measurement chain, an empirical diagnosis of its first implementation, and a more discriminating protocol for future tests of criterion revision.

Comments18 pages, 8 tables, 1 figure. CMB-0.1 is an instrument-calibration study; CMB-0.4 is a prospective protocol, not an empirical result

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑