arXivDaily arXiv每日学术速递 周一至周五更新

期刊&会议

European Conference on Computer Vision · 会议 · Computer Vision

2026-08-31 至 2026-08-31 共收录 2
2604.14129 2026-08-31 cs.CV 版本更新

Don't Let the Video Speak: Audio-Contrastive Preference Optimization for Audio-Visual Language Models

别让视频说话:面向音频-视觉语言模型的音频对比偏好优化

Ami Baid, Zihui Xue, Kristen Grauman

机构 * University of Texas at Austin(德克萨斯大学奥斯汀分校)

AI总结 本文提出ACPO方法,通过引入输出对比和输入对比目标,解决音频-视觉语言模型中视频驱动的音频幻觉问题,提升音频真实性和多模态能力。

Comments ECCV 2026 camera-ready version. Project page: this https URL (https://vision.cs.utexas.edu/projects/acpo/)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.16742 2026-08-31 cs.CV 版本更新

When the City Teaches the Car: Label-Free 3D Perception from Infrastructure

当城市教导汽车:从基础设施实现无标签的3D感知

Zhen Xu, Jinsu Yoo, Cristian Bautista, Zanming Huang, Tai-Yu Pan, Zhenzhen Liu, Katie Z Luo, Mark Campbell, Bharath Hariharan, Wei-Lun Chao

机构 * The Ohio State University(俄亥俄州立大学) Google(谷歌) Cornell University(康奈尔大学) Stanford University(斯坦福大学) Boston University(波士顿大学)

AI总结 本文提出一种无标签的3D感知方法,利用道路设施作为无监督教师,通过固定视角和重复观测学习局部3D检测器,并广播预测作为伪标签监督,无需基础设施即可训练独立的车辆检测器。

Comments ECCV 2026; Project Page: this https URL (https://jinsuyoo.info/civet/)

详情

展开后加载摘要…

URL PDF HTML 收藏