发表机构
Kling Team(Kling团队)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
ACE-Cap将长段落细粒度音频字幕生成转化为自适应主动证据获取过程,通过作曲器与指令模型协同演化及LOOP-GRPO算法优化,解决了传统模型被动生成的缺陷。
AI 中文摘要
长段落细粒度音频字幕生成要求模型恢复多样的声学事实,同时避免遗漏和无依据的细节。然而,主流字幕生成器仍是被动的一次性生成模型:一旦某个细节被忽略,它们无法识别证据缺口、查询音频以获取针对性信息,也无法判定何时已收集到足够证据。我们将该任务形式化为主动证据获取,并提出面向字幕生成的智能体协同演化框架(ACE-Cap)。该框架通过作曲器(Composer)和指令模型(Instruct)之间的多轮交互形成闭环证据获取循环:字幕生成器(Captioner)先生成初始描述;仅文本输入的作曲器基于该描述和交互历史,针对未解决的声学属性提出针对性问题,而音频条件下的指令模型提供有依据的答案;随后作曲器决定终止时机,并将累积的证据合成为最终字幕。ACE-Cap通过统一的金标准到预测的奖励训练各角色,该奖励源自固定的、基于金标准的多项选择题及仅保留字幕功能的评判器。针对变长交互的信用分配,LOOP-GRPO算法将轨迹级标量优势替换为跨度对齐信号:单个问题对累积证据的留一法贡献、终止决策的质量-成本效用,以及最终合成的证据保留效用。角色级预热后交替优化作曲器和指令模型,使每次更新成为定义明确的单策略问题,同时允许各角色协同演化。因此,ACE-Cap将字幕生成从被动一次性生成转变为自适应过程,学习需获取的证据、终止时机,以及如何在长段落字幕中保留证据。
英文摘要
Long-paragraph fine-grained audio captioning requires models to recover diverse acoustic facts while avoiding omissions and unsupported details. However, prevailing captioners remain passive one-shot generators: once a detail is overlooked, they cannot identify the evidence gap, query the audio for targeted information, or decide when sufficient evidence has been collected. We formulate this task as active evidence acquisition and introduce Agentic Co-Evolution for Captioning (ACE-Cap). The framework uses multi-turn interaction between a Composer and an Instruct model to form a closed evidence-acquisition loop. A Captioner first produces an initial description. Conditioned on this description and the interaction history, a text-only Composer asks targeted questions about unresolved acoustic attributes, while an audio-conditioned Instruct model provides grounded answers. The Composer then decides when to terminate and synthesizes the accumulated evidence into a final caption. ACE-Cap trains these roles through a unified gold-to-prediction reward derived from fixed, gold-grounded multiple-choice questions and a frozen caption-only judge. For credit assignment in variable-length interactions, LOOP-GRPO replaces the trajectory-wide scalar advantage with span-aligned signals: leave-one-out contributions of individual questions to the accumulated evidence, a quality-cost utility for stopping, and an evidence-preservation utility for final synthesis. Role-wise warm-up followed by alternating Composer and Instruct optimization keeps each update a well-defined single-policy problem while allowing the roles to co-evolve. ACE-Cap thus turns captioning from passive one-shot generation into an adaptive process that learns what evidence to acquire, when to stop, and how to preserve it in a long-paragraph caption.