arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SQBench:用于评估面向生产工作流程中语言模型代理任务交付的基准测试

SQBench: A Benchmark for Evaluating Task Delivery by Language-Model Agents in Production-Oriented Workflows

Summer Sun

arXiv 2607.23123首次发表:更新:

发表机构

Shaqiu Community(沙丘社区)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究面向生产工作流程中语言模型代理任务交付评估问题,核心方法是引入SQBench基准测试,包含多种任务并按特定方式评估,主要贡献是揭示功能完成度不能完全体现交付质量,风险判定需单独报告。

AI 中文摘要

现有对大语言模型的评估涵盖知识、推理、编码和工具使用等方面,但很少将在受限工作流程中产生的可验证交付成果作为评估单位。我们引入了SQBench,这是一个用于评估语言模型代理面向生产任务交付的基准测试。SQBench v1.0包含220个标准化任务,分为L1原子能力、L2复合技能和L3业务场景。评估先计算功能完成度,然后从10维风险矩阵中独立证明的触发因素得出风险惩罚和性能。严格通过要求完成度=1且风险惩罚=0。我们在通用协议下评估了27种模型配置,每个配置-任务对运行一次。结果表明功能完成度不能完全表征交付质量,风险判定应单独报告。

英文摘要

Existing evaluations of large language models cover knowledge, reasoning, coding, and tool use, but they rarely treat a verifiable deliverable produced within a constrained workflow as the unit of evaluation. We introduce SQBench, a benchmark for evaluating production-oriented task delivery by language-model agents. SQBench v1.0 contains 220 standardized tasks organized into L1 atomic capabilities, L2 composite skills, and L3 business scenarios. Each task requires an agent to process input assets, use available tools, and produce an explicitly specified deliverable. The evaluation first computes functional Completion and then derives Risk Penalty and Performance from independently evidenced triggers in a 10D Risk Matrix. A Strict Pass requires Completion = 1 and Risk Penalty = 0. We evaluate 27 model configurations under a common protocol, with one run per configuration-task pair. The highest prespecified Weighted Pass@1 is 60.5%. Mean Strict Pass@1 on L3 is 18.5%, and every configuration performs worse on L3 than on both L1 and L2, indicating that delivery under domain constraints is a shared weakness within the current task set. Of 2,348 results with Completion = 1, 113 (4.8%) fail the Strict Pass criterion because of risks such as unverifiable citations, inappropriate resource use, or format violations. These results show that functional completion alone does not fully characterize delivery quality and that risk determinations should be reported separately.

Comments17 pages, 8 figures. Code and aggregate results: https://github.com/shaqiu-ai/SQBench

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑