arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AcoustiTrace:看似合理的声音为何违背物理规律

AcoustiTrace: When Plausible Sound Violates Physics

Shiyang Li, Yuewen Cao, Yihao Liu, Yuandong Pu, Baochang Zhang, Xiaofei Li, Changqing Zou

arXiv 2608.02035首次发表:更新:

发表机构

Zhejiang University; Shanghai Artificial Intelligence Laboratory; Shanghai Jiao Tong University; Beihang University; Westlake University; Zhejiang Lab(浙江大学; 上海人工智能实验室; 上海交通大学; 北京航空航天大学; 西湖大学; 之江实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出AcoustiTrace诊断基准,从八个声学维度评估T2AV和I2AV生成,构建相关数据集与评估器,发现领先模型仍难符合基础声学规律,其诊断可指导模型优化与新方向探索。

AI 中文摘要

近期的音视频生成模型能够生成语义合理、表面同步的声音,但仍可能违背可见事件与环境所隐含的声学过程。现有基准在将此类违背归因于特定声学过程并量化其严重程度方面支持有限。我们提出AcoustiTrace,这是一种将音视频生成中的声学物理真实性形式化的诊断基准。AcoustiTrace围绕声学过程组织文本到音视频(T2AV)和图像到音视频(I2AV)的评估,基于可测量的声学量覆盖声音生成、传播环境、声学接收八个维度。基于这些评估维度,我们构建了围绕声学机制组织的大规模数据集,包含真实世界音视频记录和经声学标注的RGB-D观测,并用其开发针对性提示集与经验证的评估器。实验表明,即便领先的生成模型在生成合理声音事件的同时,仍在基础声学过程上存在不足。最后,我们证明AcoustiTrace对特定声学关系的诊断可指导模型优化,以生成更符合物理规律的声音,并为将声学原理融入训练目标、奖励建模与候选样本选择开辟新方向。

英文摘要

Recent audio-video generators can produce semantically plausible and apparently synchronized sound, yet may still violate the acoustic processes implied by visible events and environments. Existing benchmarks provide limited support for attributing such violations to particular acoustic processes and quantifying their severity. We introduce AcoustiTrace, a diagnostic benchmark that formalizes acoustic physical realism in audio-video generation. AcoustiTrace organizes text-to-audio-video (T2AV) and image-to-audio-video (I2AV) evaluation around the acoustic process, covering sound generation, propagation environment, and acoustic reception through eight dimensions grounded in measurable acoustic quantities. Based on these evaluation dimensions, we construct a large-scale dataset organized around acoustic mechanisms, comprising real-world audio-video recordings and acoustically annotated RGB-D observations, and use it to develop targeted prompt suites and validated evaluators. Experiments reveal that even leading generators still struggle with fundamental acoustic processes despite producing plausible sound events. Finally, we show that the diagnostics AcoustiTrace provides for specific acoustic relations can guide model refinement toward more physically faithful audio and open new directions for incorporating acoustic principles into training objectives, reward modeling, and candidate selection.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑