arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33433cs.SDcs.CLeess.AS

CORA:一种用于诊断查询改写下文本到音频检索边界鲁棒性的协议

CORA: A Protocol for Diagnosing Boundary Robustness in Text-to-Audio Retrieval under Query Reformulations

Jae Min Woo, Kyongmin Kong, Bogyung Jeong, Minjeong Kim, HaeJun Yoo, Du-Seong Chang

首次发表
浏览论文内容

中文总结 AI 辅助

提出CORA协议,通过五种意图保持的查询改写诊断T2A检索的边界鲁棒性,引入RankDrop指标,发现边界退化而非文本移动是检索失败的主因。

中文摘要 AI 辅助

文本到音频(T2A)检索器通常使用字幕风格的查询进行评估,但相同的用户意图可以用多种形式表达。我们引入了CORA(字幕偏移音频检索),一种以字幕为锚定的诊断协议,它将每个源字幕改写为五种保持意图的形式(命令、疑问、间接、关键词和陈述),同时固定目标音频。通过跨查询形式跟踪同一目标,CORA定义了RankDrop,一种揭示Recall@k所隐藏失败的指标。使用皮尔逊相关系数r,我们发现RankDrop与原始文本空间移动弱相关(r=0.084),但与目标对齐损失和目标边界裕度退化强相关(r=0.508和r=0.615)。同样的模式出现在OEA检索器中,其中RankDrop更好地由边界退化(r=0.472/0.478)解释,而非查询移动(r=0.084/0.046)。总体而言,这些结果表明,鲁棒的T2A检索需要在改写下保持目标相对于竞争音频的边界优势。

英文摘要

Text-to-Audio (T2A) retrievers are typically evaluated with caption style queries, but the same user intent can be expressed in many forms. We introduce CORA (Caption-Offset Retrieval for Audio), a caption anchored diagnostic protocol that rewrites each source caption into five intent preserving forms (Command, Question, Indirect, Key phrase, and Statement) while fixing the target audio. By tracking the same target across query forms, CORA defines RankDrop, a metric revealing failures hidden by Recall@k. Using Pearson's correlation coefficient r, we find that RankDrop is weakly associated with raw text space movement (r=0.084), but strongly associated with Target Alignment Loss and Target Boundary Margin Degradation (r=0.508 and r=0.615). The same pattern appears in OEA retrievers, where RankDrop is better explained by boundary degradation (r=0.472/0.478) than by query movement (r=0.084/0.046). Overall, these results suggest that robust T2A retrieval requires preserving the target's boundary advantage over competing audio under reformulation.

发表机构

  • Sogang University(西江大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑