Ruby-ASR:保留证据的联合正字法与词汇读音识别监督
Ruby-ASR: Evidence-Preserving Supervision for Joint Orthographic and Lexical-Reading Recognition
浏览论文内容
中文总结 AI 辅助
Ruby-ASR通过将识别目标细化为跨度绑定的正字法-词汇读音序列,在五个日语基准上提升了词汇读音恢复,同时保持正字法转录质量。
中文摘要 AI 辅助
传统的日语自动语音识别(ASR)以正字法转录作为监督,尽管相同的书写形式可能对应语音中不同的词汇读音。此类话语获得相同的目标,因此其读音差异在监督界面中缺失,且无法通过事后仅文本的字素到音素转换可靠地恢复。我们提出Ruby-ASR,将传统目标细化为跨度绑定的正字法-词汇读音序列。与分离的整句正字法和音系输出不同,ruby表示将每个书写跨度局部绑定到其实际读音,并允许确定性恢复两种视图。我们使用Qwen3-ASR骨干在字幕式和逐字式转录约定下实例化该目标;音素级CTC目标提供辅助的单调读音监督。在五个日语基准上的实验结果表明,细化识别目标可以在不牺牲可读正字法转录的情况下提高词汇读音恢复。我们发布了检查点和推理代码。
英文摘要
Conventional Japanese automatic speech recognition (ASR) is supervised by an orthographic transcript, although the same written form can correspond to different lexical readings realized in speech. Such utterances receive an identical target, so their reading distinction is absent from the supervision interface and cannot be recovered reliably by post-hoc text-only grapheme-to-phoneme conversion. We present Ruby-ASR, which refines the conventional target into a span-bound orthographic--lexical-reading sequence. Unlike separate full-sentence orthographic and phonological outputs, the ruby representation locally binds each written span to its realized reading and permits deterministic recovery of both views. We instantiate the target under subtitle-style and verbatim-style transcription conventions using a Qwen3-ASR backbone; a mora-level CTC objective provides auxiliary monotonic reading supervision. The experimental results across five Japanese benchmarks show that refining the recognition target can improve lexical-reading recovery without sacrificing readable orthographic transcription. We release the checkpoints and inference code.