AgenticASR:通过智能体方法优化真实场景中的语音识别
AgenticASR: Refining Speech Recognition in Real-World Scenarios via an Agentic Approach
AI总结:
该研究提出AgenticASR智能体语音识别框架,通过ASR-Refiner架构实现语音流的持续输出与修正,推出AASR-Bench双语基准,在多ASR前端上取得最优性能,可用于保留意图的实时干净转录。
AI中文摘要:
自动语音识别(ASR)在转录准确率上已取得显著提升,但逐字转录的结果未必是可直接使用的文本,它会保留填充词、重复内容、起始错误和自我修正,这些内容会增加阅读负担、模糊说话者的最终意图,并将未解决或被放弃的内容传递给下游任务。现有的口语转书面语方法会处理完整的音频或转录文本,但无法在后续语音改变对前文的解读方式时修正已输出的文本。因此,我们提出智能体语音识别(AgenticSR),这是一项音频转干净文本的任务,旨在消除不流畅内容、解决自我修正、规范书面形式,同时保留说话者的最终意图。AgenticASR通过ASR-Refiner架构实现该任务,该架构会在音频输入过程中反复转换有界的活跃上下文,并替换对应的输出片段,从而能对任意时长的语音流进行持续输出和修正。我们还推出了AASR-Bench,这是一个具有细粒度原子评分标准的双语基准。在多个ASR前端上,AgenticASR在所有被评估系统中取得了最高的AASR-Bench分数。一项人机一致性研究显示,基于评分标准的判断与独立专家评估结果一致。消融实验分析了Refiner的能力、上下文长度,以及在线和离线推理之间的质量-延迟权衡。综合这些结果,AgenticASR被确立为在语音进行过程中实现保留意图的干净转录的实用框架。代码、AASR-Bench和演示将在该httpsURL发布。
英文摘要:
Automatic speech recognition (ASR) has achieved substantial gains in transcription accuracy, yet verbatim transcription does not necessarily produce readily usable text. It retains fillers, repetitions, false starts, and self-corrections that increase reading effort, obscure the speaker's final intent, and propagate unresolved or abandoned content to downstream tasks. Existing spoken-to-written methods process completed audio or transcripts but cannot revise emitted text when later speech changes how preceding content should be interpreted. We therefore formulate Agentic Speech Recognition (AgenticSR), an audio-to-clean-text task that removes disfluencies, resolves self-corrections, and normalizes written form while preserving the speaker's final intent. AgenticASR implements this task through an ASR--Refiner architecture that repeatedly transforms a bounded active context and replaces its corresponding output span as audio arrives. This enables continual emission and revision over streams of arbitrary duration. We also introduce AASR-Bench, a bilingual benchmark with fine-grained atomic rubrics. Across multiple ASR front ends, AgenticASR attains the highest AASR-Bench scores among evaluated systems. A human--AI agreement study shows that rubric-based judgments align with independent expert assessments. Ablations characterize Refiner capacity, context length, and the quality--latency trade-off between online and offline inference. Together, these results establish AgenticASR as a practical framework for intent-preserving clean transcription during ongoing speech. Code, AASR-Bench, and a demo will be released at https://github.com/AnXMuy/AgenticASR.