arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

智能体可靠性源自何处?生产企业智能体中验证循环、专家模型和脚手架的跨基准分解

Where Does Agent Reliability Come From? A Cross-Benchmark Decomposition of Verification Loops, Specialist Models, and Scaffolding in a Production Enterprise Agent

Arunabh Dastidar

arXiv 2607.17044首次发表:更新:

发表机构

Leni Inc.(Leni公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究多步骤企业智能体任务失败问题,以生产系统Leni为研究对象,其含验证循环等。通过在三个基准测试评估,发现完整系统有提升。核心贡献是分解提升来源,指出大部分来自脚手架等,验证步骤单独贡献小但作用关键,还得出相关矩阵和模型等结论。

AI 中文摘要

多步骤企业智能体任务会以一种特定方式失败:单通道推理在决定答案和提交答案之间没有检查点。我们研究了一个生产系统(Leni),其架构安装了这样的检查点:由轻量级任务专用的后训练模型组成的验证循环(执行、观察、比较、纠正)。我们在强调不同失败模式的三个公共基准上评估未修改的生产配置:SpreadsheetBench Verified(无声计算错误)、BullshitBench v2(前提虚构)和GAIA验证分割(长工具链上的级联错误)。完整系统在SpreadsheetBench上比前沿基础模型提高了11.0个百分点(91.25%对80.25%,n = 400,p < 0.001),在BullshitBench上提高了7到10个百分点(98%对91%,n = 100),在GAIA验证上提高了约15个百分点(75.2% pass@1,n = 165;83.0% best - of - k)。我们的核心贡献是对这种提升的分解:大部分来自脚手架、路由和专家模型,而不是验证步骤本身,验证步骤的单独贡献很小(1.5个百分点),但集中在分数分布的顶部,在那里它能转换原本会失败的任务。我们对循环进行了端到端的检测,得出了一个经验验证器混淆矩阵(捕获率约为0.20,修复率为0.75,无误报回归),该矩阵为复合可靠性模型奠定了基础。专家交换消融表明,循环的价值取决于谁来观察它:用生成前沿模型替换小型训练验证器会消除大部分救援。有效前提控制在100个专家级问题中显示零过度拒绝。

英文摘要

Multi-step enterprise agent tasks fail in a characteristic way: single-pass inference has no checkpoint between deciding an answer and committing to it. We study one production system (Leni) whose architecture installs such checkpoints: verification loops (execute, observe, compare, correct) staffed by lightweight task-specialized post-trained models. We evaluate the unmodified production configuration on three public benchmarks stressing distinct failure modes: SpreadsheetBench Verified (silent computation error), BullshitBench v2 (premise confabulation), and the GAIA validation split (cascade error over long tool chains). The full system improves over its frontier base model by +11.0 percentage points on SpreadsheetBench (91.25% vs 80.25%, n=400, p<0.001), +7 to +10 percentage points on BullshitBench (98% vs 91%, n=100), and roughly +15 points on GAIA validation (75.2% pass@1, n=165; 83.0% best-of-k). Our central contribution is a decomposition of that uplift: most of it comes from scaffolding, routing, and specialist models rather than from the verification step itself, whose isolated contribution is small (+1.5 points) but concentrated at the top of the score distribution, where it converts otherwise-failing tasks. We instrument the loop end-to-end, yielding an empirical verifier confusion matrix (catch rate about 0.20, fix rate 0.75, no false-alarm regressions) that grounds a compounding-reliability model. Specialist-swap ablations suggest that the loop's value depends on who observes it: replacing the small trained verifier with the generating frontier model eliminates most rescues. A valid-premise control shows zero over-rejections in 100 expert-level questions.

Comments19 pages, 5 figures, 5 tables. Evaluations conducted March-April 2026. Run-level evaluation record and audit scripts: https://github.com/arnabdastidar/leni-agent-evals

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑