arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SAAG:结构化智能体评估与基础

SAAG: Structured Agent Assessment and Grounding

Ritvik Garimella, Vedant Khandelwal, Anvi Kohli, Amit Sheth

arXiv 2607.18245首次发表:更新:

发表机构

Artificial Intelligence Institute, University of South Carolina; Indian AI Research Organization(南卡罗来纳大学人工智能研究所; 印度人工智能研究组织)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对智能体调用评估问题,提出SAAG级联诊断框架,将其分解为三个阶段进行评估,能实现迭代自我修复。经实验,结构化反馈可提高参数精度、减少值幻觉,表明阶段分解诊断评估对提升智能体调用可靠性很关键。

AI 中文摘要

智能体调用的精确匹配评估掩盖了质量上不同的失败模式:模型可能选择了正确的函数但产生了幻觉的参数值,或者以错误的理由选择智能体时满足了模式。现有基准将这些区别归结为单一的二元分数,使从业者无法诊断智能体调用失败的位置。我们提出了SAAG,一个级联诊断框架,将智能体调用评估分解为三个连续阶段:注册表一致性、结构完整性和参数基础,每个阶段都产生可解释的特定阶段诊断。这些诊断还实现了迭代自我修复:在预测失败时,特定阶段的信号指导有针对性的纠正,而不会泄露真实值。我们使用三个局部小于4B参数的模型,在从Glaive的函数调用数据集中派生的受控基准上评估这个框架,跨越5、10和15个智能体的注册表大小。与单通道推理和无信息的二元反馈相比,结构化反馈持续提高了参数精度并减少了值幻觉,而端到端F1增益适中且依赖于模型。这些结果表明,阶段分解的诊断评估是理解和提高跨模型家族和注册表规模的智能体调用可靠性的必要视角。

英文摘要

Exact-match evaluation of agent-calling obscures qualitatively different failure modes: a model may select the right function yet hallucinate argument values, or satisfy a schema while choosing a agent for the wrong reason. Existing benchmarks collapse these distinctions into a single binary score, leaving practitioners unable to diagnose where agent calls fail. We propose SAAG a cascaded diagnostic framework that decomposes agent-calling evaluation into three sequential stages: registry conformance, structural completeness, and argument grounding, each producing interpretable stage-specific diagnostics. These diagnostics additionally enable iterative self-repair: on prediction failure, the stage-specific signal guides targeted correction without leaking ground-truth values. We evaluate this framework on a controlled benchmark derived from Glaive's function-calling dataset across registry sizes of 5, 10, and 15 agents using three local sub-4B-parameter models. Structured feedback consistently improves argument precision and reduces value hallucination relative to single-pass inference and uninformative binary feedback, while end-to-end F1 gains are modest and model-dependent. These results suggest that stage-decomposed diagnostic evaluation is a necessary lens for understanding and improving agent-calling reliability across model families and registry scales.

CommentsAccepted to KnowFM @ ACL 2026 workshop (Non-Archival). 16 pages in total

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑