智能体基准测试的噪声底审计
Noise Floor Audit for Agent Benchmarks
浏览论文内容
中文总结 AI 辅助
该研究针对BFCL基准的工具调用端点,审计了重复运行与提示扰动带来的测量变异性,发现提示扰动是更大噪声来源,边际准确率掩盖了稳定性与失败模式。
中文摘要 AI 辅助
我们使用匹配的AST评分,在官方BFCL多类别与并行类别上,对2家提供商的3个原生工具调用端点的测量变异性进行审计。在温度为0时,Groq端点与启用思考的Gemini设置的重复运行几乎是确定性的:翻转比例分别为0.7%、2.0%和2.7%,平均运行相关系数为0.997、0.966和0.961。保留语义的提示扰动在所有端点上造成了更大的噪声底,其中位扰动配对标准差是重复运行配对标准差的11倍至58倍。失败特征也发生了转变:格式错误输出失败占任务失败的30%、7%和<1%,因此边际准确率不仅掩盖了稳定性,还掩盖了失败模式。
英文摘要
We audit measurement variability for 3 native tool-calling endpoints across 2 providers on the official BFCL multiple and parallel categories, using matched AST grading. At temperature 0, reruns are nearly deterministic across Groq endpoints and a thinking-enabled Gemini setting: ever-flip fractions are 0.7%, 2.0%, and 2.7%, with mean run correlations of 0.997, 0.966, and 0.961. Semantics-preserving prompt perturbations create the larger floor on all endpoints, with median perturbation paired SDs 11x to 58x larger than rerun paired SDs. The failure character also shifts: malformed-output failures account for 30%, 7%, and <1% of task failures, so marginal accuracy hides not only stability but also failure mode.