arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34428cs.CLcs.AI

AgentHop:面向智能体多跳科学问答的诊断基准

AgentHop: A Diagnostic Benchmark for Agentic Multi-Hop Scientific Question Answering

Chanhee Park, Jeongho Yoon, Sungbin Han, Hyeonseok Moon, Heuiseok Lim

首次发表
浏览论文内容

中文总结 AI 辅助

AgentHop是一个含1,011道题的诊断基准,通过四轴分解准确率,在受控沙箱中揭示19个模型的检索、综合、工具调用与资源管理失败模式,并发布完整基准与测试工具。

中文摘要 AI 辅助

智能体任务要求大型语言模型与世界交互,在资源受限的情况下,跨多个步骤导航信息并收集证据。由于这种复杂性,智能体任务的失败源于多种原因,而精确定位这些失败原因对于诊断和改进智能体系统至关重要。然而,现有基准往往只关注单一的排行榜分数,使得潜在的失败模式不透明。为了填补这一空白,我们引入了AgentHop,一个包含1,011道多项选择题的诊断基准,并在固定的令牌、轮次和工具调用约束下,配有一个受控的七工具沙箱。AgentHop通过沿智能体操作的四个轴(检索、综合、工具调用和资源管理)剖析单一准确率分数,揭示模型脆弱性。在19个模型中,我们发现行为按模型家族聚类,工具调用特征揭示了不同的家族指纹:GPT模型过早提交,Anthropic和GLM检查点在提交前验证,DeepSeek和Kimi过度搜索,而Gemini-3 Pro保持平衡。分解的轴进一步暴露了家族内部结构:Claude Opus 4.6和Sonnet 4.6的准确率相差不到一个百分点,但在检索与综合的侧重上存在分歧,Opus检索更多,而Sonnet综合更好。我们发布了完整的基准集和测试工具,以支持诊断性智能体基准测试。

英文摘要

Agentic tasks require a large language model to interact with the world, navigating information and gathering evidence across multiple steps with restricted resources. Due to this complexity, agentic task failures arise from various sources, and pinpointing these failure causes is essential to diagnose and improve agentic systems. Existing benchmarks, however, tend to focus on a single leaderboard score, leaving the underlying failure modes opaque. To fill this gap, we introduce AgentHop, a diagnostic benchmark of 1,011 multiple-choice questions paired with a controlled seven-tool sandbox under fixed token, turn, and tool-call constraints. AgentHop reveals model vulnerabilities by dissecting a single accuracy score along four axes of agent operation: retrieval, synthesis, tool-call, and resource management. Across 19 models, we find that behavior clusters by model family, with tool-call signatures revealing distinct family fingerprints: GPT models commit early, Anthropic and GLM checkpoints verify before committing, DeepSeek and Kimi over-search, and Gemini-3 Pro stays balanced. Decomposed axes further expose within-family structure: Claude Opus 4.6 and Sonnet 4.6 land within one accuracy point yet diverge on retrieval-versus-synthesis emphasis, with Opus retrieving more and Sonnet synthesizing better. We release the full benchmark set and the harness to support diagnostic agent benchmarking.

发表机构

  • Korea University(高丽大学)
  • Sookmyung Women’s University(淑明女子大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑