arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ContractRL:用于可审计工具调用修复的屏蔽组相对策略优化

ContractRL: Shielded Group-Relative Policy Optimization for Auditable Tool-Call Repair

Miaobo Hu, Shuhao Hu, Xiaobo Guo, Xin Wang, Bokun Wang, Yina Sa, Daren Zha, Jun Xiao

arXiv 2610.00328首次发表:更新:

发表机构

School of Artificial Intelligence, University of Chinese Academy of Sciences; Institute of Information Engineering, Chinese Academy of Sciences(中国科学院大学人工智能学院; 中国科学院信息工程研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

ContractRL提出契约约束的顺序修复协议,通过操作掩码和组相对策略优化,在减少生成令牌的同时提升结构化工具调用修复的语义成功率与可审计性。

AI 中文摘要

结构化工具调用通常仅在少量字段违反模式或执行契约后就会失败。重新生成完整对象会扩大操作面,并使重复修复难以审计。我们引入了ContractRL,一种契约约束的顺序修复协议,将验证器引导的JSON修复建模为有界决策过程。在每一步,策略观察候选对象、类型化验证器反馈、JSON指针、不可变的修复历史和剩余预算;契约派生的操作掩码在确定性验证器执行转换之前过滤格式错误或被禁止的RFC-6902操作。我们为补丁、重试和弃权(不执行)决策指定了一个契约约束的组相对目标,同时将规范目标和语义标签保持在在线状态之外,直到跟踪冻结。在相同的验证器信息下,ContractRL在每种子192个案例和五种子的情况下,以34.4个生成令牌达到0.9362的语义成功率,而Patch-SFT为0.9076和44.9个令牌,完整重新生成为0.9148和137.2个令牌。策略优化将语义成功率从监督ContractRL的0.9186提高到0.9375。与Patch-SFT的单独三种子配对评估产生+0.0396的语义差异(95%置信区间[+0.0137,+0.0662],p=0.0039)。反馈、操作掩码、预算和模式偏移分析将这些收益与局部修正联系起来,而对抗性和多轮评估则表征了剩余的失败模式。

英文摘要

Structured tool calls often fail after only a small number of fields violate a schema or an execution contract. Regenerating the complete object enlarges the action surface and makes repeated repair difficult to audit. We introduce ContractRL, a contract-constrained sequential repair protocol that models verifier-guided JSON repair as a bounded decision process. At each step the policy observes the candidate, typed verifier feedback, JSON Pointer, immutable repair history, and remaining budget; a contract-derived action mask filters malformed or prohibited RFC-6902 operations before a deterministic validator performs the transition. We specify a contract-constrained group-relative objective for patch, retry, and abstention decisions while keeping canonical targets and semantic labels outside the online state until trace freeze. Under identical verifier information, ContractRL attains 0.9362 semantic success with 34.4 generated tokens, compared with 0.9076 and 44.9 tokens for Patch-SFT and 0.9148 and 137.2 tokens for full regeneration over 192 cases per seed and five seeds. Policy optimization improves semantic success from 0.9186 for supervised ContractRL to 0.9375. A separate three-seed paired evaluation against Patch-SFT yields a semantic difference of +0.0396 (95% CI $[+0.0137,+0.0662], p=0.0039$). Feedback, action-mask, budget, and schema-shift analyses connect these gains to localized correction, while adversarial and multi-turn evaluations characterize the remaining failure modes.

Comments29 pages, 8 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑