AI 中文总结
研究针对结构化音频字幕评估难的问题,提出多轴评估框架,结合大语言模型判断与计算指标,在五个正交轴上评估,通过受控扰动测试协议验证可靠性,能区分释义与损坏。
AI 中文摘要
近期自动音频字幕(AAC)的进展已从整体句子生成转向明确区分不同声学和语义属性的结构化格式。然而,评估这种异构数据仍是重大挑战。现有字幕指标专注于平面文本输出,无法可靠评估多模态属性。为填补这一差距,我们提出针对结构化音频描述的多轴评估框架。基于AudioCards数据集,在五个正交轴上评估输出:标签集、描述、逻辑推理、数值测量和频谱轮廓。我们的方法结合大语言模型(LLM)判断捕捉语义细微差别与确定性计算指标精确测量声学偏差。为严格验证该框架的可靠性,我们引入受控扰动测试协议,向真实注释中注入分类分级错误。结果表明该框架成功区分了保留意义的释义与真正的语义和声学损坏。
英文摘要
Recent advances in automated audio captioning (AAC) are driving a shift from monolithic sentences toward structured formats that disentangle acoustic and semantic properties, such as timestamped captions for different sound events. Such representations can support faceted sound search for creators and richer access to auditory information for Deaf and Hard of Hearing people. Yet, it remains unclear how to meaningfully evaluate these hybrid, structured captions. We propose an evaluation framework for structured audio descriptions, spanning five complementary axes: tag sets, descriptions, reasoning, numeric measurements, and spectral profiles. The framework combines large language model (LLM) judges for semantic fields with deterministic metrics for temporal and acoustic attributes. To validate these metrics, we introduce controlled perturbations that apply typed, graded changes to ground-truth annotations. Results show that the proposed metrics remain robust to meaning-preserving paraphrases while responding to genuine semantic and acoustic corruptions, enabling more reliable evaluation of structured captions.