发表机构
Northwestern Polytechnical University; yutuzhineng; Shenzhen Pimei Technology Co., Ltd.; Nanjing University(西北工业大学; 玉途智能; 深圳市品美科技有限公司; 南京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
SpeechAnnotator提出可本地部署的上下文感知多智能体框架,通过规划、标注、审查三智能体协作及有界审查循环,实现多维语音标注,并引入SA-Bench基准和SA-Eval评估体系,验证其作为商业系统的替代方案。
AI 中文摘要
近期可控语音生成需要带有说话人特质、韵律、情感、副语言线索、声学场景和上下文等细粒度标注的训练数据。现有工作流程通常依赖人工修正、付费托管的模态服务或固定处理链,这通过标注成本、外部服务依赖或跨阶段恢复能力弱而限制了大规模数据处理。我们提出SpeechAnnotator,一个完全由开源模型和工具构建的、可本地部署的上下文感知多智能体框架。支持前端模块首先获取说话人感知片段和最终片段转录,而先验证据提取器附加异质片段级线索。随后三个专家智能体通过共享状态协作:规划智能体将局部音频证据、说话人历史、相邻片段和录音级上下文转换为字段特定合约;标注智能体对可直接观察属性执行合约引导的多模态预测;审查智能体运行一个有界审查循环,检查证据支持和跨片段一致性,仅对不支持或不一致的字段触发重新标注。为解决现有评估资源在孤立任务和窄域测试集上的碎片化问题,我们引入SpeechAnnotator-Bench(SA-Bench),包含9种源格式的8.87小时人工标注音频,以及SpeechAnnotator-Eval(SA-Eval),其将Timeline-Eval用于说话人感知时间线恢复,Closed-Eval用于有限集属性,Open-Eval用于开放式属性。实验和消融研究表明,SpeechAnnotator提供了商业音频能力系统的本地可部署替代方案,而有界审查循环通过证据和上下文感知的字段级恢复改进了多维标注。
英文摘要
Recent controllable speech generation requires training data with fine-grained annotations of speaker traits, prosody, emotion, paralinguistic cues, acoustic scenes, and context. Existing workflows often rely on manual correction, paid hosted multimodal services, or fixed processing chains, which limits large-scale data processing through annotation cost, external-service dependence, or weak cross-stage recovery. We introduce SpeechAnnotator, a locally deployable, context-aware multi-agent framework built entirely from open-source models and tools. Supporting frontend modules first obtain speaker-aware segments and final segment transcripts, while prior evidence extractors attach heterogeneous segment-level cues. Three specialist agents then collaborate through shared state: the Planning Agent converts local audio evidence, speaker history, neighboring segments, and recording-level context into field-specific contracts; the Labeling Agent performs contract-guided multimodal prediction for directly observable attributes; and the Review Agent runs a bounded review loop that checks evidence support and cross-segment consistency, triggering relabeling only for unsupported or inconsistent fields. To address the fragmentation of existing evaluation resources across isolated tasks and narrow-domain test sets, we introduce SpeechAnnotator-Bench (SA-Bench), containing 8.87 hours of human-annotated audio across nine source formats, together with SpeechAnnotator-Eval (SA-Eval), which separates Timeline-Eval for speaker-aware timeline recovery, Closed-Eval for finite-set attributes, and Open-Eval for open-ended attributes. Experiments and ablations show that SpeechAnnotator provides a locally deployable alternative to commercial audio-capable systems, while the bounded review loop improves multidimensional annotation through evidence- and context-aware field-level recovery.
Comments15 pages, 3 figures, to be published in NCMMSC 2026