发表机构
Independent Researcher(独立研究者)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究大语言模型智能体验证与调用接口问题,提出VeriHarness代码控制智能体循环。通过比较原始诊断与含失败位置等信息的反馈,发现含三个字段反馈能大幅提升成功率,多数增益源于可接受替代方案,JSON语法本身对修复无显著作用。
AI 中文摘要
大语言模型智能体在外部验证拒绝候选方案后经常重试,但验证与下一次模型调用之间的接口仍未明确规定。我们引入了VeriHarness,这是一种代码控制的智能体循环,模型生成候选方案,外部验证器控制接受、预算和跟踪。我们用它来比较原始诊断与识别失败位置、观察值和可接受替代方案的反馈。在四次调用上限下的50对TextWorld游戏中,包含所有三个字段的反馈将Qwen2.5-Coder-14B的最终成功率从14/50提高到36/50(提高44个百分点),将Llama-3.1-8B的成功率从8/50提高到29/50(提高42个百分点)。消融实验表明大部分增益来自可接受替代方案:仅包含位置和观察值的反馈仍接近原始诊断基线。以散文形式而非键控JSON记录呈现完整修复信息产生的成功率几乎相同,这表明JSON语法本身并不能改进修复。这种排序在测试的调用预算和一种采样解码设置中持续存在。
英文摘要
LLM agents often retry after external validation rejects a candidate, but the interface between validation and the next model call remains underspecified. We introduce VeriHarness, a code-controlled agent loop in which models generate candidates while external validators control acceptance, budgets, and traces. We use it to compare raw diagnostics with feedback that identifies the failure location, observed value, and admissible alternatives. Across 50 paired TextWorld games under a four-call cap, feedback containing all three fields raises terminal success from 14/50 to 36/50 for Qwen2.5-Coder-14B (+44 percentage points) and from 8/50 to 29/50 for Llama-3.1-8B (+42 points). Ablations locate most of the gain in the admissible alternatives: feedback containing only the location and observed value remains near the raw diagnostic baseline. Presenting the complete repair information in prose instead of a keyed JSON record yields nearly the same success, providing no evidence that JSON syntax itself improves repair. The ordering persists across the tested call budgets and one sampled-decoding setting.