arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33538cs.CLeess.AS

Jev 在语音神经假体重打分中匹敌 7B 语言模型

Jev Matches 7B Language Models for Speech-Neuroprosthesis Rescoring

Gabriele Cinà

中文总结 AI 辅助

针对语音神经假体重打分,提出用托管模型 Jev 替代 7B 语言模型,在 978 个 ALS 参与者句子上词错误率更低(7.5% vs 7.8%),无需 GPU 且成本极低。

中文摘要 AI 辅助

语音神经假体从大脑活动中解码尝试性言语,并最终通过一个数十亿参数的语言模型对解码器的候选句子进行重打分,该模型是唯一需要 GPU 的组件。用更便宜的模型替换它很困难:通用语言模型被要求从列表中选择一个句子时,其回答依据的是标签在列表中的位置,而非句子本身。我们将重打分视为一个单一的类型化决策,即一次调用返回每个候选的概率,由 Jev 提供服务,这是一个为校准决策而训练的托管模型,并将其与解码器自身的得分相结合。在来自一位肌萎缩侧索硬化症(ALS)参与者的 978 个保留句子上,已发表的解码器单独达到 8.1% 的词错误率,Jev 达到 7.5%,而 OPT-6.7b 和 Qwen2.5-7B 均为 7.8%;在重新调整解码器权重后,Jev 达到 6.9%,而 OPT-6.7b 和 Qwen2.5-7B 分别为 7.2% 和 7.4%。Jev 在所有四项比较中均领先,且在 95% 置信界限处最多落后 0.2 个百分点。其成本为每千句 0.07 美元,且无需 GPU;专用 GPU 运行 7B 模型仅在利用率超过 43% 时每句成本更低,这远超单个用户产生的负载。通过互联网的端到端延迟为 262 毫秒,其中 62 毫秒花费在提供商处,与本地 GPU 上的 7B 模型(27 毫秒)处于同一数量级,但并不更快。

英文摘要

A speech neuroprosthesis decodes attempted speech from brain activity and ends by rescoring the decoder's candidate sentences with a language model of several billion parameters, the only component that needs a GPU. Replacing that model with a cheaper one is hard: general language models asked to pick one sentence from a list answer from where a label sits in the list rather than from the sentence itself. We pose rescoring as a single typed decision, one call that returns a probability for every candidate, served by Jev, a hosted model trained for calibrated decisions, and combine it with the decoder's own score. On 978 held-out sentences from a participant with ALS, where the published decoder alone reaches 8.1% word error, Jev reaches 7.5% against 7.8% for both OPT-6.7b and Qwen2.5-7B; with the decoder's weight re-tuned, 6.9% against 7.2% and 7.4%. Jev is ahead in all four comparisons and at most 0.2 points behind at the 95% bound. It costs 0.07 USD per thousand sentences and needs no GPU; a dedicated GPU running a 7B model is cheaper per sentence only above 43% utilisation, far beyond what one user generates. End-to-end latency over the internet is 262 ms, of which 62 ms is spent at the provider, the same order as a 7B model on a local GPU (27 ms) but not faster.

补充信息

↑