个性化韩语唇读作为视觉语音识别:在OLKAVS上的迁移、普查与适配
Personalized Korean Lipreading as Visual Speech Recognition: Transfer, Census and Adaptation on OLKAVS
浏览论文内容
中文总结 AI 辅助
本文提出个性化韩语视觉语音识别系统,在OLKAVS上量化个体与群体误差差距,并采用低秩适配器微调,以少量数据显著降低高错误率用户的CER。
中文摘要 AI 辅助
我们提出一个个性化韩语视觉语音识别(VSR)系统,并在九摄像头OLKAVS语料库上量化了群体级基准分数与个体用户错误之间的差距。一个从英语训练权重初始化的仅视频Conformer,在语料库协议下达到9.95%至12.19%的字符错误率(CER),而已发表的基线为26.64%;在未见词汇上达到19.00%至21.52%。每位说话人的CER范围从1.0%到52.2%,其中已见词汇使CER降低7.0至9.0个百分点,而专业播报和自发语音分别使CER提高8.5至10.5和12.7个百分点。一个仅含4.6%参数的低秩适配器,在用户正面视频的4至29分钟数据上训练,将十二名高错误率说话人的CER降低2.13至3.58个百分点,且可无损迁移至每个摄像头,并以12%的成本保留全量微调对其他说话人收益的85%。位于口部平面以上的摄像头带来约六个CER点的恒定偏移,而在所有视角上训练可使其保持较小。
英文摘要
We present a personalized Korean visual speech recognition (VSR) system and quantify, on the nine-camera OLKAVS corpus, the gap between the population-level benchmark score and an individual user's error. A video-only Conformer initialized from English-trained weights attains 9.95 - 12.19% character error rate (CER) under the corpus protocol against the published 26.64, and 19.00 - 21.52 on unseen wording. Per speaker, CER spans 1.0 to 52.2%, with seen wording lowering CER by 7.0 - 9.0 points and professional delivery and spontaneous speech raising it by 8.5 - 10.5 and 12.7 points. A low-rank adapter with 4.6% of the parameters, trained on 4 to 29 minutes of the user's frontal video, lowers the CER of twelve high-error speakers by 2.13 to 3.58 points, transfers to every camera without loss, and keeps 85% of the full fine-tuning gain at 12% of its cost to other speakers. Cameras above the mouth plane add about six CER points as a constant offset that training on all views keeps small.
发表机构
- UX Factory, Inc.(UX工厂公司)
机构由 AI 辅助整理,请以论文原文为准。