arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

无需解码的批量语音决策:单令牌监督让冻结的LLM能听到转录文本之外的信息

Batched Speech Decisions Without Decoding: Single-Token Supervision Lets a Frozen LLM Hear Beyond the Transcript

Jie Jin, Ziyin Ma, Min Yin, Jinyu Chen, Haigang Song, Zhikun Pang, Xiaowen Zhang

arXiv 2610.02638首次发表:更新:

发表机构

College of Computer Science and Technology, Zhejiang University; Future Design Lab, Innovation Center of Yangtze River Delta, Zhejiang University; University of Virginia(浙江大学计算机科学与技术学院; 浙江大学长三角创新中心未来设计实验室; 弗吉尼亚大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对全双工语音代理的封闭决策,提出DuplexJev,通过连接器将ASR隐藏状态输入冻结LLM,以单令牌分布读取选项,无需解码,实现快速批量决策并提升语音感知能力。

AI 中文摘要

全双工语音代理需要做出许多小而封闭的决策,而当前系统通过缓慢的自回归解码来响应这些决策。我们提出DuplexJev,该方法通过一个小型连接器将ASR编码器的隐藏状态输入到冻结的LLM中,并将每个问题读取为其选项上的单令牌分布。整个过程无需解码,一个8-GPU节点在约0.1秒内即可回答关于八个话语的80个决策。使用最后一层连接器时,语音问答的性能接近阅读转录文本(90%对91%)。DuplexJev还能听到说话者的信息:使用交叉注意力连接器时,性别和情绪准确率均达到90%(从55%和28%提升),其语音问答仅下降1个百分点(从83%降至82%)。我们使用读出答案令牌上的交叉熵来训练决策,而非通常的转录蒸馏——后者的教师模型从未听到语音——并保留蒸馏用于内容生成。编码器和LLM是可互换的;我们发布了权重、训练配方、用于全双工服务的批量推理管道以及一个双语语音问答数据集。

英文摘要

Full-duplex voice agents make many small, closed decisions, which current systems answer by slow autoregressive decoding. We propose DuplexJev, which feeds ASR-encoder hidden states through a small connector into a frozen LLM and reads each question as a single-token distribution over its options. Nothing is decoded, and an 8-GPU node answers 80 decisions about eight utterances in about 0.1 s. With a last-layer connector, spoken QA stays close to reading the transcript (90% vs. 91%). DuplexJev also hears the speaker: gender and emotion accuracy both reach 90% (from 55% and 28%) with a cross-attention connector, whose spoken QA drops by only 1 point (83% to 82%). We train decisions with cross-entropy on the read-out answer token, instead of the usual transcript distillation, whose teacher never hears the voice, and keep distillation for content. Encoders and LLMs are interchangeable; we release weights, training recipe, a batched-inference pipeline for full-duplex serving and a bilingual spoken-QA set.

Comments5 pages, 2 figures, 3 tables. Submitted to ICASSP 2027. Code and weights: https://github.com/adventists-ai/duplexjev

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑