发表机构
Virginia Tech; Accenture(弗吉尼亚理工大学; 埃森哲公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究审计了AV-Align等4种视听同步指标,发现它们在不同任务上各有优劣且一致性差,建议以含置信区间的指标类别细分形式报告视听同步结果。
AI 中文摘要
自动视听同步(AV-sync)指标被广泛用于对视听生成器进行排名和训练,但作为测量工具却很少被审计。我们在通用可靠性协议下共同审计了AV-Align、ImageBind AV相关性、JavisScore和Synchformer/DeSync,该协议包含受控失真单调性、预处理敏感性、排名不确定性、跨指标一致性、PEAVS代理一致性及学习融合。结果显示存在轴分化,无单一最优指标:Synchformer/DeSync是最强的时间偏移跟踪器(τ=0.84),ImageBind/JavisScore更匹配PEAVS人类对齐代理(τ=0.20)及内容破坏类别,AV-Align是最弱的独立指标。各指标间一致性极低(Krippendorff α=0.066),线性或简单k-NN融合均未提升PEAVS一致性。我们建议将视听同步报告为可靠性卡(含置信区间的指标类别细分),而非单一原始同步分数。
英文摘要
Automatic AV-sync metrics are widely used to rank and train audio-visual generators, but they are rarely audited as measurement instruments. We jointly audit AV-Align, ImageBind AV-relevance, JavisScore, and Synchformer/DeSync under a common reliability protocol: controlled-distortion monotonicity, preprocessing sensitivity, rank uncertainty, cross-metric agreement, PEAVS-proxy agreement, and learned fusion. The result is an axis split, not a single winner: Synchformer/DeSync is the strongest temporal-offset tracker ($τ=0.84$), ImageBind/JavisScore better match the PEAVS human-aligned proxy ($τ=0.20$) and content-disruption families, and AV-Align is the weakest standalone metric. The metrics mutually disagree (Krippendorff $α=0.066$), and neither linear nor simple $k$-NN fusion improves PEAVS agreement over the best individual metric. We recommend reporting AV-sync as a Reliability Card (metric-family breakdowns with confidence intervals) rather than a single bare synchronization score.
CommentsAccepted at the ECCV 2026 Workshop on Generative AI for Audio-Visual Content Creation (Gen4AVC), poster presentation; non-archival workshop. 7 pages (4-page main text + references + 2-page appendix), 3 figures, 8 tables. Project page: https://jaishrm07.github.io/avsync-reliability-card/