arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

闭合验证循环:用于长段落详细音频字幕的自检查字幕生成

Closing the Verification Loop: Self-Check Captioning for Long-Paragraph Detailed Audio Captioning

Fengji Ma, Yan Rong, Xu Li, Chen Zhang, Pengfei Wan, Li Liu

arXiv 2608.30713首次发表:更新:

发表机构

The Hong Kong University of Science and Technology (Guangzhou); Kuaishou Technology(香港科技大学(广州); 快手科技)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对长段落详细音频字幕的未解决问题,提出SCC框架,构建LACap-50k语料库与LC-SFT方法,使系统在开源字幕生成器中达最优性能。

AI 中文摘要

长段落详细音频字幕需要对细粒度音频内容进行密集且符合原文转录的描述,目前的视听多模态语言模型仍未解决该问题。我们将此失败归因于两个结构性问题:其一为数据匮乏,尚无公开语料库同时提供长音频片段、段落级字幕及逐字转录保真度;其二为生成模式失效,表现为正确音频与打乱音频的多项选择题(MCQ)准确率存在44.8至46.4个百分点的差距。我们在自检查字幕生成(Self-Check Captioning, SCC)这一统一框架内解决上述问题,该框架在生命周期的每个阶段都将基于音频的问答作为验证基础。SCC产生三类产物:长段落音频字幕50k(Long-paragraph Audio Caption 50k, LACap-50k)是包含50222个音频片段的视听语料库,配有平均491.5词的字幕及事后自动语音识别(ASR)审计;层曲率监督微调(Layer-Curvature Supervised Fine-Tuning, LC-SFT)是首个基于中间层证据对标记进行加权的在线策略监督微调方法,其提出源于我们发现的晚层语义熵崩溃(Late-Layer Semantic-Entropy Collapse, SEC)现象;SCC验证器(SCC-Verifier)在推理阶段通过基于音频的自回答来评判字幕生成结果。在多个基准测试中,我们的系统在开源字幕生成器中达到了最优性能,且与专有基线具有竞争力。我们发布LACap-50k以填补长段落详细音频字幕研究的资源缺口。

英文摘要

Long-paragraph detailed audio captioning, which requires dense and transcript-faithful descriptions of fine-grained audio content, remains unsolved for current audio-visual multimodal language models. We attribute this failure to two structural problems. The first is data poverty, as no public corpus jointly provides long clips, paragraph captions, and verbatim-transcript fidelity. The second is generation-mode failure, evidenced by a 44.8 to 46.4 percentage-point gap between right-audio and shuffled-audio multiple-choice question (MCQ) accuracy. We address both within Self-Check Captioning (SCC), a unified framework that instantiates audio-grounded question answering as the verification primitive at every lifecycle stage. SCC yields three artifacts. Long-paragraph Audio Caption 50k (LACap-50k) is a 50,222-clip audio-visual corpus with 491.5-word captions and a post-hoc automatic speech recognition (ASR) audit. Layer-Curvature Supervised Fine-Tuning (LC-SFT) is the first on-policy supervised fine-tuning method to weight tokens by intermediate-layer evidence, motivated by our identification of Late-Layer Semantic-Entropy Collapse (SEC). SCC-Verifier arbitrates among caption rollouts via audio-grounded self-answering at inference. Across multiple benchmarks, our system attains state-of-the-art among open-source captioners and is competitive with proprietary baselines. We release LACap-50k to fill the resource gap for long-paragraph detailed audio captioning research.

CommentsEMNLP2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑