arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

VAKRA:评估工具使用策略下跨API与检索的多跳推理能力

VAKRA: Evaluating Multi-Hop Reasoning Across APIs and Retrieval Under Tool-Use Policies

Ankita Rajaram Naik, Anupama Murthi, Benjamin Elder, Siyu Huo, Raavi Gupta, Abhinav Jain, Praveen Venkateswaran, Abdulhamid Adebayo, Danish Contractor

arXiv 2608.12282首次发表:更新:

发表机构

IBM(国际商业机器公司(IBM))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究推出VAKRA基准评估工具使用策略下跨API与检索的多跳推理,发现现有最优模型在相关任务上性能随推理深度显著下降,失败集中于语言介导推理环节。

AI 中文摘要

部署在企业环境中的智能体必须对结构化API和文档集合进行推理,但现有基准测试会孤立地评估这些能力。我们推出VAKRA(全称Evaluating API and Knowledge Retrieval Agents,即评估API与知识检索智能体),这是一个涵盖62个领域、包含8000多个可执行API的基准,任务分为三个难度递增的场景:多样化API交互风格、结构化API上的多跳推理、受自然语言工具使用策略约束的多源推理。正确性通过对预测的工具调用重新执行实时API进行验证,允许多种有效路径。使用固定的ReAct框架以隔离模型能力与智能体架构,我们评估了前沿模型和开放权重模型,发现即使是最优模型在单跳端点式任务上也仅达到70.4%,在组合API上降至50%-51%;随着推理深度增加,性能下降超过50%,而受策略约束的问题暴露出严重缺陷(无法回答的查询上低至2.4%)。轨迹分析显示,失败集中在语言介导的推理——实体消歧、跨源 grounding,而非工具调用机制。代码和数据集可在指定URL获取。

英文摘要

Agents deployed in enterprise settings must reason across structured APIs and document collections, yet existing benchmarks evaluate these capabilities in isolation. We introduce VAKRA (e\textbf{V}aluating \textbf{A}PI and \textbf{K}nowledge \textbf{R}etrieval \textbf{A}gents), a benchmark of over $8{,}000$ executable APIs across $62$ domains with tasks spanning three settings of increasing difficulty: diverse API interaction styles, multi-hop reasoning over structured APIs, and multi-source reasoning with natural-language tool-use policy constraints. Correctness is verified by re-executing predicted tool calls against live APIs, accommodating multiple valid paths. Using a fixed ReAct harness to isolate model capabilities from agent architecture, we evaluate frontier and open-weight models and find that even the best model achieves only 70.4\% on single-hop endpoint-style tasks and drops to 50--51\% on compositional APIs; performance degrades by over 50\% as reasoning depth increases, and policy-constrained questions expose severe failures (as low as 2.4\% on unanswerable queries). Trace analysis shows failures concentrate at language-mediated reasoning - entity disambiguation, cross-source grounding, rather than tool invocation mechanics. Code is available https://github.com/IBM/VAKRA. Dataset is available https://huggingface.co/datasets/ibm-research/VAKRA

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑