AESSI:一种用于跨天在线复用且无需测试日校准的耳周无声语音接口
AESSI: An Around-Ear Silent Speech Interface for Cross-Day Online Reuse without Test-Day Calibration
AI总结:
AESSI是一种耳周无声语音接口,利用掩蔽上下文表示预训练实现跨天在线复用,无需测试日校准,在24名参与者上达到92.24%准确率,并支持端到端操作。
AI中文摘要:
无声语音接口(SSIs)能够在不发出可听见语音的情况下实现私密通信,并可能为中风后构音障碍患者提供支持。日常复用要求发音相关的表征能够跨天泛化,尽管存在传感器重新定位和生理变化。我们提出了AESSI,一种使用掩蔽上下文表示预训练(MCRP)的耳周SSI:学生模型从掩蔽的时频输入中预测教师表示,以增强对记录变异的鲁棒性。我们收集了24名参与者对25个日常普通话句子的44次电生理记录。在保留六名参与者的完整最终记录的情况下,AESSI在无需测试日校准的情况下达到了92.24%的平均准确率。AESSI比最佳自适应基线高出50.57个百分点。在每位参与者最后一次记录至少21天后,五名参与者各自完成了50次独立随机化的在线测试,无需校准,总体准确率达到98.0%。在CPU回放中,预处理和推理的中位时间为31.70毫秒。这些结果展示了端到端系统操作,并支持无需测试日校准的在线复用。演示包含在论文中。
英文摘要:
Silent speech interfaces (SSIs) enable private communication without audible speech and may support people with post-stroke dysarthria. Everyday reuse requires articulation-related representations that generalize across days despite sensor repositioning and physiological changes. We present AESSI, an around-ear SSI using masked-context representation pretraining (MCRP): a student predicts teacher representations from masked time-frequency inputs to encourage robustness to recording variability. We collected 44 electrophysiological recordings from 24 participants for 25 everyday Mandarin sentences. With six participants' complete final recordings held out, AESSI achieved 92.24 percent mean accuracy without test-day calibration. AESSI exceeded the best adapted baseline by 50.57 percentage points. At least 21 days after each participant's last recording, five participants each completed 50 independently randomized online tests without calibration, achieving 98.0 percent overall accuracy. Median preprocessing and inference time in CPU replay was 31.70 ms. These results demonstrate end-to-end system operation and support online reuse without test-day calibration. Demo is included with the paper.