arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.02887cs.LGq-bio.NC

语音脑机接口的通用通信度量

A Common Measure of Communication for Speech Brain-Computer Interfaces

Dulhan Jayalath, Benjamin Ballyk, Oiwi Parker Jones

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对语音脑机接口缺乏通用进展度量的问题,提出开放词汇互信息(OVMI)作为统一通信度量,可用于比较异构系统、优化词汇表设计并提升准确率。

中文摘要 AI 辅助

语音脑机接口(speech BCIs)可将神经活动转换为语言,为瘫痪患者恢复言语能力提供了途径,更广泛地说,还能实现新型自然人机交互。尽管前景广阔,但该领域缺乏通用的进展度量标准,因为不同系统使用不同的数据集、记录方法、语音类型和词汇表,导致其报告的分数几乎无法比较。这一测量问题背后有两个未解决的问题:(i)语音脑机接口应使用什么样的单词分布让用户进行通信;(ii)系统能传达该分布中的多少信息。我们通过推导开放词汇互信息(OVMI)来解决这两个问题,这是一种信息论量度,用于测量解码器相对于用户可能想要通信的单词参考分布所传达的信息。这使得不同条件(如不同词汇表)下测量的性能能在统一的通信尺度上进行评估。我们表明,通常报告的准确率、词错误率(WER)以及仅在系统支持的单词上计算的其他指标,可能会夸大系统能传达用户意图言语的程度。随后,我们使用OVMI比较现有系统,揭示了系统支持用户语言的程度与解码这些单词的准确性之间的权衡,表明这些比较取决于用户预期要通信的内容,并证明选择词汇表以最大化OVMI可在三个语音领域中使准确率获得高达16.3%的相对提升。因此,OVMI为语音脑机接口社区提供了一种有原则的方法,用于比较异构系统、改进词汇表设计并衡量该领域的进展。

英文摘要

Speech brain-computer interfaces (speech BCIs) translate neural activity into language, offering a path towards restoring speech for people with paralysis and, more broadly, enabling new forms of natural human-computer interaction. Despite this promise, the field lacks a common measure of progress because systems use different datasets, recording methods, types of speech, and vocabularies, so their reported scores are rarely comparable. Underlying this measurement problem are two unresolved questions: (i) what distribution of words should a speech BCI enable a user to communicate, and (ii) how much information from this distribution can a system convey. We address both by deriving open-vocabulary mutual information (OVMI), an information-theoretic quantity that measures the information conveyed by a decoder relative to a reference distribution over the words a user may wish to communicate. This allows capabilities measured under different conditions, such as distinct vocabularies, to be evaluated on a common communication scale. We show that ordinarily reported accuracy, word error rate (WER), and other metrics computed only over the words a system supports can overstate how much of a user's intended speech the system can communicate. We then use OVMI to compare existing systems, expose trade-offs between how much of the user's language a system supports and how accurately it decodes those words, show that these comparisons depend on what the user is expected to communicate, and demonstrate that selecting a vocabulary to maximise OVMI yields up to 16.3% relative improvement in accuracy across three speech domains. OVMI therefore provides the speech BCI community with a principled way to compare heterogeneous systems, improve vocabulary design, and measure progress in the field.

发表机构

  • University of Oxford(牛津大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑