arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.26911cs.AI

TwinCheck:基于证据的负孪生验证用于有状态工具智能体

TwinCheck: Evidence-Grounded Negative-Twin Verification for Stateful Tool Agents

  • Phillips Exeter Academy(菲利普斯埃克塞特学院)
  • Ryquo

机构由 AI 辅助整理,请以论文原文为准。

Jiaxuan Dai, Tianyi Huang

AI总结:

TwinCheck通过基于证据的负孪生验证,在推理时对工具调用进行受约束的比较替换,将GPT-5.6 Sol在BFCL V4任务上的成功率从45.3%提升至58.5%,且无成功率回退。

AI中文摘要:

单个局部看似合理的工具调用可能使原本成功的智能体轨迹偏离正轨。仅凭怀疑并不足以证明干预的合理性,因为替换本身可能引入验证本应预防的失败。我们提出TwinCheck,一种推理时验证策略,仅当轨迹满足与轨迹局部失败假设相关的证据条件时,才考虑替换。它构建一个基于轨迹的反事实替代方案,即负孪生,并且仅当孪生通过结构检查且成对验证器在两种候选顺序中都偏好它时,才替换智能体的提议。对于配对评估,精确重放保持智能体的解析响应和动作固定,直到第一次接受的替换,从而将干预效果与重采样分开。在对159个多轮BFCL V4任务(具有完整精确重放对)的主要分析中,完整策略将GPT-5.6 Sol的任务成功率从45.3%提高到58.5%(95%任务自助置信区间[8.2, 18.8]),且未观察到成功到失败的回归。总之,这些发现将执行边界修复重新定义为受约束的比较,使反事实动作本身成为验证的对象。

英文摘要:

A single locally plausible tool call can derail an otherwise successful agent trajectory. Suspicion alone does not justify intervention, because the replacement itself can introduce the very failure verification is meant to prevent. We introduce TwinCheck, an inference-time verification policy that considers replacement only when the trace satisfies an evidence condition tied to a trace-local failure hypothesis. It constructs a trace-grounded counterfactual alternative, a negative twin, and replaces the agent's proposal only if the twin passes structural checks and the pairwise verifier prefers it in both candidate orders. For paired evaluation, exact replay holds the agent's parsed responses and actions fixed until the first accepted replacement, separating intervention effects from resampling. In the primary analysis of 159 multi-turn BFCL V4 tasks with complete exact-replay pairs, the complete policy raises task success for GPT-5.6 Sol from 45.3% to 58.5% (95% task-bootstrap CI [8.2, 18.8]), with no observed success-to-failure regressions. Together, these findings recast execution-boundary repair as a constrained comparison, making the counterfactual action itself the object of verification.

补充信息

↑