发表机构
Los Alamos National Laboratory(洛斯阿拉莫斯国家实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过SLM评判者反馈循环的智能体小型语言模型集成,在IFEval基准上以97.34%准确率超越独立LLM基线5.81个百分点,并分析其代币成本与延迟,证明额外测试时代币可换取更高指令遵循保真度,推动多维评估。
AI 中文摘要
随着大型语言模型(LLM)从独立助手转向智能体工作流,评估必须超越标量排行榜准确率,以考虑操作可靠性、成本、延迟和代币效率。我们以一个小型语言模型(SLM)的智能体集成作为案例研究,该集成采用SLM评判者介导的反馈循环,用于此类超越排行榜的评估。在包含541个提示的IFEval基准上,最佳集成实现了97.34%的严格提示准确率,比最强的独立LLM基线gpt-5.4高出5.81个百分点,同时运行在较低成本区间。随后,我们分析了这一增益背后的代币经济学和操作行为,包括每个样本的成本、代币组成、有用输出吞吐量、反馈循环恢复、延迟分解以及跨指令类别和约束数量的性能。我们的结果表明,智能体SLM集成可以用额外的测试时代币和编排开销换取改进的指令遵循保真度,从而为未来的智能体AI系统推动多维评估协议。
英文摘要
As large language models (LLMs) move from standalone assistants into agentic workflows, evaluation must extend beyond scalar leaderboard accuracy to account for operational reliability, cost, latency, and token efficiency. We use an agentic ensemble of small language models (SLMs) with an SLM-judge-mediated feedback loop as a case study for such beyond-leaderboard evaluation. On the 541-prompt IFEval benchmark, the best ensemble achieves 97.34% strict prompt accuracy, exceeding the strongest standalone LLM baseline, gpt-5.4, by 5.81 percentage points while operating in a lower-cost regime. We then analyze the tokenomics and operational behavior behind this gain, including cost per sample, token composition, useful-output goodput, feedback-loop recovery, latency decomposition, and performance across instruction categories and constraint counts. Our results show that agentic SLM ensembles can trade additional test-time tokens and orchestration overhead for improved instruction-following fidelity, motivating multi-dimensional evaluation protocols for future agentic AI systems.
Comments8 pages, 9 figures, Presented at ACM CAIS 2026 Workshop RLEval: Methods and Reinforcement Learning Environments for Evaluating AI Agents. Resubmission of permitted appeal, Ticket #MOD-104177