SURE-EVAL:一个用于可复现评估的系统化统一智能体框架
SURE-EVAL: A Systematic and Unified Agentic Framework for Reproducible Evaluation
浏览论文内容
中文总结 AI 辅助
针对音频语音模型评估中分数不可复现的问题,提出SURE-EVAL统一智能体框架,通过工具与主智能体工作流控制推理和评分,在18个模型上完成全部评估,并揭示报告与统一结果存在0.02-0.52分差异。
中文摘要 AI 辅助
音频和语音模型发布迅速,但报告分数常常将模型能力与部署和评估选择混为一谈。同一个检查点在不同运行时、硬件、解码设置或回退策略下可能产生不同的预测。即使固定预测在不同归一化和指标实现下也可能获得不同分数。现有语音基准标准化了选定的数据集或评分程序,但很少在一个可执行工作流中连接异构模型接入、受控推理和版本化评分。我们提出SURE-EVAL,一个用于音频和语音系统可复现评估的系统化统一智能体框架。工具智能体工作流将模型发布转换为隔离的、经过验证的可调用工具。主智能体工作流提交特定任务的推理和评分协议,然后将所有涉及分数的操作委托给版本化的确定性程序。每个结果保留其运行时、协议、流水线节点、预测和审计工件。在涵盖自动语音识别、文本到语音、语音转换、说话人日志、说话人属性识别和多任务音频理解的18个公开发布中,仅使用Codex的基线一次性完成12个模型,而使用SURE-EVAL工具智能体工作流的同一智能体完成了全部18个。我们还对七个ASR测试条件和两个TTS子集进行了统一评估。对三个TTS系统的协议分析发现,论文报告结果与统一结果之间的绝对差异为0.02-0.52分,方向因模型和语言而异。这些结果表明,可复现评估需要同时控制模型执行和输出评分。
英文摘要
Audio and speech models are released rapidly, but reported scores often conflate model capability with deployment and evaluation choices. The same checkpoint can produce different predictions under different runtimes, hardware, decoding settings, or fallback policies. Even fixed predictions can receive different scores under different normalization and metric implementations. Existing speech benchmarks standardize selected datasets or scoring procedures, but rarely connect heterogeneous model onboarding, controlled inference, and versioned scoring in one executable workflow. We introduce SURE-EVAL, a Systematic and Unified Agentic framework for Reproducible Evaluation of audio and speech systems. A Tool Agent Workflow converts model releases into isolated, verified callable tools. A Main Agent Workflow commits task-specific inference and scoring protocols, then delegates all score-bearing operations to versioned deterministic programs. Each result retains its runtime, protocol, pipeline nodes, predictions, and audit artifacts. Across 18 public releases covering automatic speech recognition, text-to-speech, voice conversion, speaker diarization, speaker-attributed recognition, and multi-task audio understanding, a Codex-only baseline completes 12 models in one shot, while the same agent with the SURE-EVAL Tool Agent Workflow completes all 18. We also conduct unified evaluations over seven ASR test conditions and two TTS subsets. A protocol analysis of three TTS systems finds absolute differences of 0.02-0.52 points between paper-reported and unified results, with the direction varying by model and language. These results show that reproducible evaluation requires controlling both model execution and output scoring.
发表机构
- Shanghai Jiao Tong University(上海交通大学)
- University of Bristol(布里斯托大学)
- Imperial College London(帝国理工学院)
- Xi’an Jiaotong University(西安交通大学)
- ETH Zurich(苏黎世联邦理工学院)
- AISpeech Co., Ltd.(思必驰科技股份有限公司)
- Nanjing University(南京大学)
机构由 AI 辅助整理,请以论文原文为准。