发表机构
Google Cloud AI Research; University of Cambridge(谷歌云人工智能研究院; 剑桥大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对长时程任务中LLM智能体输出验证难题,提出VeriHarness方法,将底层LLM转化为智能体验证器,通过分歧解决与共识挑战机制,在五个基准上取得最优选择分数,并实现自我改进。
AI 中文摘要
随着LLM智能体承担越来越复杂的长时程任务,验证其输出变得越来越具有挑战性。我们研究了在固定基础模型下,如何在测试时无参考答案或评分标准的情况下增强验证能力。重复采样产生多个轨迹,其中可能包含互补的正确声明,但我们需要可靠的验证机制来确定哪些声明值得信任。我们首先发现,分歧往往暴露正确的替代方案,而共识可能隐藏错误。这些观察促使我们提出VeriHarness,它通过给生成器所用的底层LLM提供工作空间、证据工具和可复用的验证技能,将其转变为智能体验证器。一个分歧解决器对照环境证据检查竞争性声明,而一个共识挑战者测试共享声明并搜索遗漏的需求。它们的发现指导最终产物的选择和修订。在五个长时程工作空间基准和两个前沿模型上,VeriHarness在评估的基线中取得了最高的选择分数。基于证据的修订进一步提高了平均性能,与单次轨迹相比,Gemini 3.5 Flash提升了6.2分,Claude Opus 4.8提升了6.4分。我们进一步表明,验证技能可以从失败反馈中自我改进,证明VeriHarness是扩展长时程智能体验证的一种新颖且关键的方法。我们发布了所有五个基准和两个模型的约26,000个轨迹的完整池,制作成本超过100,000美元,以支持未来关于智能体验证的研究。
英文摘要
As LLM agents undertake increasingly complex, long-horizon tasks, verifying their outputs becomes increasingly challenging. We study how verification capability can be strengthened with a fixed base model, without access to reference answers or grading rubrics at test time. Repeated sampling yields multiple rollouts that can contain complementary correct claims, but we need a reliable verification mechanism to determine which claims to trust. We first find that disagreement often exposes correct alternatives, while consensus can conceal errors. These observations motivate VeriHarness, which turns the underlying LLM a generator uses into an agentic verifier by giving it a workspace, evidence tools, and reusable verification skills. A disagreement resolver checks competing claims against environmental evidence, while a consensus challenger tests shared claims and searches for omitted requirements. Their findings guide the selection and revision of the final artifact. Across five long-horizon workspace benchmarks and two frontier models, VeriHarness achieves the highest selection scores among the evaluated baselines. Evidence-backed revision further improves average performance, bringing gains over a single rollout to 6.2 points with Gemini 3.5 Flash and 6.4 points with Claude Opus 4.8. We further show that verification skills can self-improve from failure feedback, demonstrating VeriHarness as a novel and critical approach for scaling long-horizon agentic verification. We release the full pool of approximately 26,000 rollouts across all five benchmarks and both models, produced at a cost of over $100,000, to support future research on agentic verification.