arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向智能体语音识别的语音记忆

Voice Memory for Agentic Speech Recognition

Chao-Han Huck Yang, Zih-Ching Chen, Piotr Zelasko, Zhehuai Chen, Jagadeesh Balam, Boris Ginsburg

arXiv 2607.26410首次发表:更新:

发表机构

NVIDIA(英伟达)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出仅推理方案Voice Memory,通过听者-思考者架构克制过度校正,在10个HyPoradise域上将加权词错误率降至7.52%,可跨校正器迁移且不增加推理参数。

AI 中文摘要

我们提出了Voice Memory,这是一种用于智能体语音识别的仅推理方案:在流式处理阶段,一个冻结的校正器读取每个域的单个this http URL(此处为原文链接,保留原格式),并针对每个话语决定是对假设进行处理还是弃权(不执行)并保留1-best结果。异步地,一个分数门控优化器通过有界编辑修改该文件,仅当编辑严格提高保留的分数时才接受该编辑。该方法从经典ASR-LM框架扩展而来,我们将这种架构称为“听者-思考者”架构;两个角色仅通过记忆耦合,因此无需更改任何权重,学习到的技能保持可审计性和可移植性。该循环发现的关键技能是克制:无约束的生成错误校正(GER)会过度校正,在财经新闻上高达64%的编辑会破坏正确的标记,而Voice Memory将该比率降至35%。在带有开放校正器的10个HyPoradise域中,Voice Memory将加权词错误率从8.36%降至7.52%(添加3个上下文示例后为7.47%),且未使任何数据集低于其1-best基线;增益集中在具有最大可恢复余量的地方,包括航空旅行命令(8.40%降至3.40%)和嘈杂远场语音(CHiME-4,12.69%降至10.46%)。该记忆可跨校正器家族迁移,且不会给推理路径增加任何参数。提供了演示和示例代码以供未来研究使用。

英文摘要

We present Voice Memory, a inference-only scheme for agentic speech recognition: at stream time, a frozen corrector reads a single per-domain memory.md and decides per utterance whether to act on the hypothesis or abstain and keep the 1-best. Asynchronously, a score-gated optimizer revises that file through bounded edits, accepting an edit only when it strictly improves a held-out score. Extended from classical ASR-LM framework, we refer this split the listener-thinker architecture; the two roles are coupled only through the memory, so no weights change and the learned skill stays auditable and portable. Restraint turns out to be the operative skill this loop discovers: unconstrained generative error correction (GER) over-corrects, breaking correct tokens on up to 64% of its edits on financial news, and Voice Memory, reduces this rate to 35%. Across ten HyPoradise domains with an open corrector, Voice Memory, lowers weighted word error rate from 8.36% to 7.52% (7.47% with three added in-context examples) without regressing any dataset below its 1-best baseline; gains concentrate where recoverable headroom is largest, including air-travel commands (8.40% to 3.40%) and noisy far-field speech (CHiME-4, 12.69% to 10.46%). The memory transfers across corrector families and adds zero parameters to the inference path. A demo and example code are provided for future studies.

CommentsPreprint. Technical report and open source: https://huggingface.co/huckiyang/voice-memory

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑