arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

听、推理与分段:使大型音频语言模型(LALMs)与编辑判断对齐以实现媒体章节化

Listen, Reason, and Segment: Aligning LALMs with Editorial Judgment for Media Chapterization

Tony Alex, Wish Suharitdamrong, Sara Atito, Armin Mustafa, Muhammad Awais, Philip J. B. Jackson, Jiankang Deng, Ismail Elezi

arXiv 2608.16539首次发表:更新:

发表机构

University of Surrey(萨里大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对LALMs在实际媒体章节化部署的不足,提出基于GRPO与CoT的AudioChaps框架,构建三类数据集,使AudioChaps-R1的平均F1较SOTA提升49个百分点,实现非结构化听觉流到结构化媒体的可靠转换。

AI 中文摘要

大型音频语言模型(LALMs)在标准化基准测试中取得了快速进展,但它们在实际媒体工作流程、内容整理、档案索引和内容分发中的部署仍基本未实现。我们将自动音频章节化——即把连续音频流分割成主题连贯的章节的任务——视为一项具有挑战性且商业意义重大的场景,该场景暴露了上述差距。章节化之所以具有挑战性,是因为章节边界并非由客观声学事件定义,而是由主观编辑判断决定,这要求模型对长声学上下文进行顺序推理,并近似创作者标注的边界决策。我们提出AudioChaps,这是一个通过思维链(CoT)推理引导的组相对策略优化(GRPO),使端到端LALMs对齐以完成该任务的后训练框架。为支持训练与评估,我们整理了三个数据集:AudioChaps-Alignment,源自YouTube上创作者标注的章节边界;AudioChaps-CoT,为格式良好、高质量且有证据支撑的边界推理提供结构化监督;以及AudioChaps-Eval,一个用于音频章节化的保留基准。不经过监督微调(SFT)冷启动,直接应用GRPO得到的AudioChaps-R1-Zero,已比最先进的LALM Audio-Flamingo-3-Think的平均F1值提升了33个百分点。AudioChaps框架生成了我们最终的对齐LALM AudioChaps-R1,其平均F1值提升了49个百分点。这些结果表明,经GRPO训练的LALMs可将非结构化听觉流可靠地转换为可导航的结构化媒体。我们的代码、模型和数据集资源将在录用后于该httpsURL发布。

英文摘要

Large Audio Language Models (LALMs) have made rapid progress on standardized benchmarks, yet their deployment in practical media workflows, curation, archival indexing, and content distribution remains largely unrealized. We identify automated audio chapterization, the task of segmenting continuous audio streams into thematically coherent chapters, as a demanding and commercially consequential setting that exposes this gap. Chapterization is challenging because boundaries are defined less by objective acoustic events than by subjective editorial judgment, requiring models to reason sequentially over long acoustic contexts and approximate creator-authored boundary decisions. We present AudioChaps, a post-training framework for aligning end-to-end LALMs for this task via Group Relative Policy Optimization (GRPO) guided by Chain-of-Thought (CoT) reasoning. To support training and evaluation, we curate three datasets: AudioChaps-Alignment, derived from creator-annotated chapter boundaries on YouTube; AudioChaps-CoT, which provides structured supervision for well-formatted, high-quality, and evidence-grounded boundary reasoning; and AudioChaps-Eval, a held-out benchmark for audio chapterization. Applying GRPO directly without a Supervised Fine-Tuning (SFT) cold start, AudioChaps-R1-Zero already improves average F1 by 33 points over the state-of-the-art LALM Audio-Flamingo-3-Think. The AudioChaps framework produces our final aligned LALM, AudioChaps-R1, which improves average F1 by 49 points. These results demonstrate that GRPO-trained LALMs can reliably transform unstructured auditory streams into navigable, structured media. Our code, models, and dataset resources will be released upon acceptance at https://github.com/ta012/AudioChaps.

Comments19 pages, 9 figures, 8 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑