arXivDaily arXiv每日学术速递 周一至周五更新

大厂专区

NVIDIA(英伟达)

2026-07-07 至 2026-07-07 共收录 7
2605.27243 2026-07-07 cs.CV 版本更新

Can Retrieval Heads See Images? Multimodal Retrieval Heads in Long-Context Vision-Language Models

检索头能看见图像吗?长上下文视觉语言模型中的多模态检索头

Aaron Branson Cigres Li, Zhaowei Wang, Yu Zhao, Yiming Du, Haobo Li, Xiyu Ren, Ginny Wong, Simon See, Lishu Luo, Haodong Duan, Pasquale Minervini, Yangqiu Song

机构 * HKUST(香港科技大学) University of Edinburgh(爱丁堡大学) CUHK(香港中文大学) NVAITC, NVIDIA, Santa Clara, USA(NVIDIA Santa Clara 分公司) Tsinghua University(清华大学)

AI总结 本文提出一种多模态检索头检测方法,发现视觉语言模型中仅有4.4-10.2%的注意力头贡献了50%的正检索分数,这些头对长上下文推理至关重要,且可直接用于文档检索提升性能。

Comments Work in Progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.06358 2026-07-07 cs.CL cs.AI 版本更新

SHINE: A Scalable In-Context Hypernetwork for Mapping Context to LoRA in a Single Pass

SHINE:一种可扩展的上下文超网络,用于在单次传递中将上下文映射到LoRA

Yewei Liu, Xiyuan Wang, Yansheng Mao, Yoav Gelbery, Haggai Maron, Muhan Zhang

机构 * Institute for Artificial Intelligence, Peking University(北京大学人工智能研究院) Technion, NVIDIA(技术学院与NVIDIA) University of Oxford(牛津大学) School of Electronics Engineering(电子工程学院) Computer Science, Peking University(计算机科学,北京大学)

AI总结 本文提出SHINE,一种可扩展的上下文超网络,用于在单次传递中将多样且有意义的上下文映射到高质量的LoRA适配器,通过重用冻结LLM的自身参数和引入架构创新,克服了先前超网络的关键限制,以较少的参数实现了强大的表达能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.26946 2026-07-07 cs.CV cs.RO 版本更新

Three-Step Nav: A Hierarchical Global-Local Planner for Zero-Shot Vision-and-Language Navigation

三步导航:一种用于零样本视觉-语言导航的分层全局-局部规划器

Wanrong Zheng, Yunhao Ge, Laurent Itti

机构 * University of Southern California(南加州大学) NVIDIA Research(NVIDIA研究)

AI总结 本文提出三步导航方法,通过全局-局部分层规划解决零样本视觉-语言导航中的漂移和低成功率问题,无需梯度更新或微调,实现最先进的性能。

Comments Accepted to AISTATS 2026. Code: https://github.com/ZoeyZheng0/3-step-Nav

Journal ref Proceedings of the 29th International Conference on Artificial Intelligence and Statistics (AISTATS), 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.17248 2026-07-07 eess.AS cs.CL cs.SD 版本更新

VIBE: Voice-Induced open-ended Bias Evaluation for Large Audio-Language Models via Real-World Speech

VIBE:通过真实世界语音进行大型音频-语言模型生成偏见评估的语音诱导开放式偏见评估

Yi-Cheng Lin, Yusuke Hirota, Sung-Feng Huang, Hung-yi Lee

机构 * Graduate Institute of Communication Engineering, National Taiwan University, Taiwan(台湾大学通信工程研究所) NVIDIA, Taiwan(台湾NVIDIA) Artificial Intelligence Center of Research Excellence, National Taiwan University, Taiwan(台湾大学人工智能卓越研究中心)

AI总结 VIBE通过真实世界语音的开放式任务评估大型音频-语言模型的生成偏见,揭示了性别线索比口音线索更易引发分布偏移,表明当前模型复现了社会刻板印象。

Comments Submitted to SLT 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08828 2026-07-07 cs.CV cs.AI cs.LG cs.MM cs.RO 版本更新

Motion Attribution for Video Generation

视频生成中的运动归因

Xindi Wu, Despoina Paschalidou, Jun Gao, Antonio Torralba, Laura Leal-Taixé, Olga Russakovsky, Sanja Fidler, Jonathan Lorraine

机构 * NVIDIA Princeton University(普林斯顿大学) MIT(麻省理工学院) University of Michigan(密歇根大学) University of Toronto(多伦多大学) Vector Institute(向量研究所)

AI总结 研究视频生成模型中数据对运动影响,提出基于梯度的运动归因框架Motive,通过运动加权损失掩码分离静态外观和时间动态,用于研究微调剪辑对时间动态的影响及指导数据策划,提升视频运动质量。

Comments See the project website at https://research.nvidia.com/labs/sil/projects/MOTIVE/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24380 2026-07-07 cs.LG 版本更新

APEX: Approximate-but-exhaustive search for ultra-large combinatorial synthesis libraries

APEX:针对超大组合合成库的近似但详尽搜索

Aryan Pedawi, Jordi Silvestre-Ryan, Bradley Worley, Darren J Hsu, Kushal S Shah, Elias Stehle, Jingrong Zhang, Izhar Wallach

机构 * Numerion Labs(Numerion实验室) NVIDIA(英伟达)

AI总结 研究超大组合合成库虚拟筛选难题,提出APEX方法,利用神经网络代理,能在一分钟内在消费级GPU上实现完全枚举,可精确检索近似top-k集,性能优于其他方法。

Comments Published in the Proceedings of the 43rd International Conference on Machine Learning (ICML 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.16663 2026-07-07 cs.RO cs.CV cs.LG cs.SY eess.SY 版本更新

Mitigating Covariate Shift in Imitation Learning for Autonomous Vehicles Using Latent Space Generative World Models

使用潜在空间生成世界模型减轻自动驾驶模仿学习中的协变量转移

Alexander Popov, Alperen Degirmenci, David Wehr, Shashank Hegde, Ryan Oldja, Alexey Kamenev, Bertrand Douillard, David Nistér, Urs Muller, Ruchi Bhargava, Stan Birchfield, Nikolai Smolyanskiy

机构 * NVIDIA

AI总结 提出用潜在空间生成世界模型解决自动驾驶协变量转移问题,训练时利用世界模型减轻该问题,无需大量训练数据,还引入新感知编码器,实验显示相比之前有显著改进。

Comments 8 pages, 6 figures, original September 2024, accepted at ICRA 2025 Workshop "Robots in the Wild", for associated video file, see https://youtu.be/7m3bXzlVQvU

详情

展开后加载摘要…

URL PDF HTML 收藏