arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 7464 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7464 篇

2305.10438 2023-05-19 cs.CL cs.AI cs.CV cs.MM 76%

IMAGINATOR: Pre-Trained Image+Text Joint Embeddings using Word-Level Grounding of Images

Varuna Krishna, S Suryavardan, Shreyash Mishra, Sathyanarayanan Ramamoorthy, Parth Patwa, Megha Chakraborty, Aman Chadha, Amitava Das, Amit Sheth

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.07528 2023-05-15 cs.CV cs.AI 76%

WEDGE: A multi-weather autonomous driving dataset built from generative vision-language models

Aboli Marathe, Deva Ramanan, Rahee Walambe, Ketan Kotecha

专题命中 视觉定位与Grounding :vision-language model(title);分类 cs.CV、cs.AI

Comments Accepted in Vision Datasets Understanding at CVPR 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.10285 2022-08-10 cs.CV cs.CL cs.LG 76%

Grounding Visual Representations with Texts for Domain Generalization

Seonwoo Min, Nokyung Park, Siwon Kim, Seunghyun Park, Jinkyu Kim

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV、cs.LG

Comments ECCV 2022; 25 pages (including Supplementary Materials); Updated related works

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.03892 2022-03-28 cs.CL cs.AI cs.LG 76%

Retrieve, Caption, Generate: Visual Grounding for Enhancing Commonsense in Text Generation Models

Steven Y. Feng, Kevin Lu, Zhuofu Tao, Malihe Alikhani, Teruko Mitamura, Eduard Hovy, Varun Gangal

专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI、cs.LG

Comments Accepted to AAAI 2022. Code at https://github.com/styfeng/VisCTG

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.11805 2022-01-19 cs.AI cs.LG 76%

Neural-Symbolic Integration for Interactive Learning and Conceptual Grounding

Benedikt Wagner, Artur d'Avila Garcez

专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI、cs.LG

Comments Corrected references

Journal ref 1st Workshop on Human and Machine Decisions (WHMD 2021), NeurIPS 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.08409 2021-11-17 cs.LG cs.CV 76%

Grounding Psychological Shape Space in Convolutional Neural Networks

Lucas Bechberger, Kai-Uwe Kühnberger

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV、cs.LG

Comments accepted at CIFMA2021 (https://cifma.github.io/)

详情

展开后加载摘要…

URL PDF HTML 收藏
1910.01210 2021-06-21 cs.CV cs.LG cs.RO 76%

Embodied Language Grounding with 3D Visual Feature Representations

Mihir Prabhudesai, Hsiao-Yu Fish Tung, Syed Ashar Javed, Maximilian Sieb, Adam W. Harley, Katerina Fragkiadaki

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV、cs.LG

Journal ref Conference on Computer Vision and Pattern Recognition. 2020, pp. 2220-2229

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.05379 2020-10-13 cs.CL cs.CV cs.LG 76%

MAF: Multimodal Alignment Framework for Weakly-Supervised Phrase Grounding

Qinxin Wang, Hao Tan, Sheng Shen, Michael W. Mahoney, Zhewei Yao

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1708.00133 2018-12-07 cs.CL cs.AI cs.LG 76%

Grounding Language for Transfer in Deep Reinforcement Learning

Karthik Narasimhan, Regina Barzilay, Tommi Jaakkola

专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI、cs.LG

Comments JAIR 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
1810.01871 2018-10-05 cs.LG cs.AI cs.RO stat.ML 76%

Grounding the Experience of a Visual Field through Sensorimotor Contingencies

Alban Laflaquière

专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI、cs.LG

Comments 23 pages, 7 figures, published in Neurocomputing

Journal ref Neurocomputing, Volume 268, 13 December 2017, Pages 142-152

详情

展开后加载摘要…

URL PDF HTML 收藏
1810.01870 2018-10-05 cs.LG cs.AI cs.RO stat.ML 76%

Grounding Perception: A Developmental Approach to Sensorimotor Contingencies

Alban Laflaquière, Nikolas Hemion, Michaël Garcia Ortiz, Jean-Christophe Baillie

专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI、cs.LG

Comments 8 pages, 4 figures, workshop at IROS 2015 conference

详情

展开后加载摘要…

URL PDF HTML 收藏
1806.02739 2018-10-04 cs.RO cs.AI cs.LG 76%

Discovering space - Grounding spatial topology and metric regularity in a naive agent's sensorimotor experience

Alban Laflaquière, J. Kevin O'Regan, Bruno Gas, Alexander Terekhov

专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI、cs.LG

Comments 59 pages, 16 figures, submitted to Neural Networks

详情

展开后加载摘要…

URL PDF HTML 收藏
1805.04396 2018-05-14 cs.RO cs.AI cs.LG 76%

A Sensorimotor Perspective on Grounding the Semantic of Simple Visual Features

Alban Laflaquière

专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI、cs.LG

Comments 6 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
1511.04401 2017-12-08 cs.CV cs.CL cs.LG cs.NE 76%

Symbol Grounding Association in Multimodal Sequences with Missing Elements

Federico Raue, Andreas Dengel, Thomas M. Breuel, Marcus Liwicki

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV、cs.LG

Comments Under review on Journal of Artificial Intelligence Research (JAIR) -- Special Track on Deep Learning, Knowledge Representation, and Reasoning

详情

展开后加载摘要…

URL PDF HTML 收藏
1412.6451 2014-12-22 cs.LG cs.AI cs.RO 76%

Grounding Hierarchical Reinforcement Learning Models for Knowledge Transfer

Mark Wernsdorfer, Ute Schmid

专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI、cs.LG

Comments 14 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
1908.01211 2020-06-02 cs.CL cs.LG cs.RO 76%

Word2vec to behavior: morphology facilitates the grounding of language in machines

David Matthews, Sam Kriegman, Collin Cappelle, Josh Bongard

专题命中 视觉定位与Grounding :grounding(title,comments);分类 cs.LG

Comments D. Matthews, S. Kriegman, C. Cappelle and J. Bongard, "Word2vec to behavior: morphology facilitates the grounding of language in machines," 2019 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Macau, China, 2019. \c{opyright} 2019 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses

详情

展开后加载摘要…

URL PDF HTML 收藏
1706.00355 2017-06-02 cs.AI 76%

Grounding Symbols in Multi-Modal Instructions

Yordan Hristov, Svetlin Penkov, Alex Lascarides, Subramanian Ramamoorthy

专题命中 视觉定位与Grounding :grounding(title,comments);分类 cs.AI

Comments 9 pages, 8 figures, To appear in the Proceedings of the ACL workshop Language Grounding for Robotics, Vancouver, Canada

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.17584 2026-08-21 cs.RO 版本更新 75%

HODAgent: Towards On-Demand, Responsive Humanoids for Physical World Human Interaction

HODAgent:面向物理世界人机交互的按需、响应式人形机器人

Wang Warren Chen, Jiahao Zhang, Zhenjiang Li, Mingxu Wang, Lei Yi, Yuchen Kang, Shuo Sun, Ziping Chen, Jie Chen

机构 * Xiaopeng(小鹏)

专题命中 视觉定位与Grounding :VLM(abstract,abstract_cn);grounding(abstract)

AI总结 HODAgent作为面向服务人形机器人的System-2具身智能体,通过半双工架构实现自适应服务,在仿真与实体机器人上均显著优于基准模型。

Comments we have received a formal directive from our company requiring all company assets to undergo a mandatory internal review process before any public release. We are now required to immediately withdraw the paper to comply with this policy

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.12882 2026-07-15 cs.MM cs.IR 新提交 75%

What Would You Click? Personalized Video Thumbnail Generation with Preference-aware Highlight Retrieval

你会点击什么?基于偏好感知高光检索的个性化视频缩略图生成

Zhiyu He, Zecheng Zhao, Tong Chen, Zi Huang, Yiqun Liu, Min Zhang

专题命中 视觉定位与Grounding :VLM(summary_cn,abstract)

AI总结 研究针对视频平台个性化视频缩略图生成问题,提出两阶段框架,先通过偏好感知检索选取视觉锚点,再经VLM引导的扩散管道生成缩略图,实验和用户研究证明该方法性能优且能提升用户参与度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.03430 2026-07-14 cs.RO 版本更新 75%

ProAct: A Benchmark and Multimodal Framework for Structure-Aware Proactive Response

ProAct:用于结构感知主动响应的基准和多模态框架

Xiaomeng Zhu, Fengming Zhu, Weijie Zhou, Ye Tian, Zhenlin Hu, Yufei Huang, Yuchun Guo, Xinyu Wu, Zhengyou Zhang, Fangzhen Lin, Xuantang Xiong

机构 * Department of Computer Science and Engineering, The Hong Kong University of Science and Technology (HKUST), Hong Kong SAR, China(香港科技大学计算机科学与工程系) Tencent, Shenzhen, China(腾讯(中国深圳)) Shenzhen Institute of Advanced Technology (SIAT), Chinese Academy of Sciences, Shenzhen, China(深圳先进技术研究所(SIAT),中国科学院)

专题命中 视觉定位与Grounding :grounding(abstract);multimodal large language model(abstract);MLLM(abstract)

AI总结 针对主动智能体开发受资源缺乏阻碍的问题,引入ProAct-75基准,提出由多模态大语言模型驱动的ProAct-Helper,其利用任务图进行行动选择,实验证明该方法在触发检测等方面优于闭源模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.06402 2026-07-08 cs.CV cs.AI cs.LG 新提交 75%

What Images Cannot Say: Language-Guided Olfactory Representation Learning

图像无法表达的:语言引导的嗅觉表征学习

Eleftherios Tsonis, Xi Wang, Vicky Kalogeiton

机构 * LIX, École Polytechnique, IP Paris, CNRS(LIX,巴黎高等理工学院,IP巴黎,CNRS)

专题命中 视觉定位与Grounding :vision-language model(abstract);grounding(abstract);分类 cs.CV、cs.AI、cs.LG

AI总结 研究图像与嗅觉对齐难题,提出SCENT多模态框架,利用视觉语言模型生成场景描述符,训练气味编码器,通过语言引导潜在分解分离特定对象气味与环境贡献,提升跨模态检索性能,产生可解释表征。

Comments ECCV 2026. Project page: https://www.lix.polytechnique.fr/vista/projects/2026_scent_tsonis/

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.26923 2026-06-26 cs.CL 新提交 75%

GAVEL: Grounded Caption Error Verification and Localization

GAVEL:基于视觉定位的标题错误验证与定位

Zixian Gao, Atsushi Hashimoto, Kuniaki Saito

机构 * OMRON SINIC X Corporation(欧姆龙SINIC X公司)

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract_cn);grounding(abstract)

AI总结 提出GAVEL任务,联合解决图像-文本对的验证、解释和定位问题,构建数据集和基准,监督基线在定位和解释指标上持续改进。

Comments conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15347 2026-06-26 eess.IV cs.MM 版本更新 75%

Symmetric Entropy-Constrained Video Coding for Machines

面向机器的对称熵约束视频编码

Yuxiao Sun, Meiqin Liu, Chao Yao, Qi Tang, Jian Jin, Weisi Lin, Frederic Dufaux, Yao Zhao

专题命中 视觉定位与Grounding :MLLM(abstract,abstract_cn);grounding(abstract)

AI总结 提出SEC-VCM框架,通过双向熵约束机制对称对齐视频编解码器与视觉骨干网络,保留语义并丢弃无关信息,结合语义-像素双路径融合提升机器视觉任务性能,在多项任务上显著优于H.266/VVC。

Comments Accepted by IEEE Transactions on Image Processing. This is the author's accepted manuscript (AAM)

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.22043 2026-06-23 cs.AI cs.CV cs.LG 新提交 75%

When Does a Video-Language Model Stop Watching? Reward Strength Controls the Formation and Reversal of Visual Shortcuts in Multimodal RLVR

视频语言模型何时停止观看?多模态RLVR中视觉捷径的形成与逆转受奖励强度控制

Zekun Xu

机构 * Zekun Xu(徐泽坤)

专题命中 视觉定位与Grounding :vision-language model(abstract);grounding(abstract);分类 cs.CV、cs.AI、cs.LG

AI总结 研究多模态强化学习中视觉捷径的形成与逆转机制,发现惩罚强度可控制其出现和消除,且存在关键干预窗口。

Comments 11 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.21168 2026-06-23 cs.CL 新提交 75%

Dementia-Agents: A Multi-Modal Multi-Agent System for Dementia Staging and Phenotyping

Dementia-Agents:用于痴呆分期和表型分析的多模态多智能体系统

Yaling Shen, Maja Christensen, Yiwen Jiang, Jenna Dennison, David Darby, Amy Brodtmann, Zongyuan Ge

机构 * Monash University(莫纳什大学) Eastern Health(东部健康) Lived Experience Advisor(生活经验顾问)

专题命中 视觉定位与Grounding :MLLM(summary_cn,abstract_cn)

AI总结 提出Dementia-Agents多智能体框架,通过数据翻译、专家预测和协调聚合三步流程,在真实临床队列中实现优于单一MLLM和先前多智能体系统的痴呆综合征级分期与表型分析。

Comments 8 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.20764 2026-06-23 cs.CV cs.AI cs.GR cs.LG 新提交 75%

One Image is All You Need: Agentic One-Shot Image Generation via Text-Based World Models for Long-Tail Spatial Perception

一图足矣:基于文本世界模型的智能体单样本图像生成用于长尾空间感知

Keqin Zeng, Shuting Su, Shihao Lin, Ziyue Li, Rui Zhao

机构 * Tsinghua University(清华大学) SenseTime Research(商汤科技研究院) Sun Yat-Sen University(中山大学) Technical University of Munich(慕尼黑工业大学) Heilbronn Data Science Center(海尔布隆数据科学中心) Munich Data Institute(慕尼黑数据研究所)

专题命中 视觉定位与Grounding :vision-language model(abstract);grounding(abstract);分类 cs.CV、cs.AI、cs.LG

AI总结 提出WMGen-v1框架,利用单张参考图像通过LVLM构建结构化场景表示,LLM进行物理合理的场景扩展,再由扩散模型生成多样化的长尾训练数据,缓解空间感知中的数据稀缺问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.16898 2026-06-10 cs.RO cs.AI cs.CV cs.LG 75%

MALLVI: A Multi-Agent Framework for Integrated Generalized Robotics Manipulation

MALLVI:一种多智能体框架用于集成通用机器人操作

Mehrshad Taji, Arad Mahdinezhad Kashani, Iman Ahmadi, AmirHossein Jadidi, Saina Kashani, Babak Khalaj

机构 * Department of Electrical Engineering, Sharif University of Technology(电气工程系,谢里夫大学)

专题命中 视觉定位与Grounding :vision language model(abstract);VLM(abstract);分类 cs.CV、cs.AI、cs.LG

AI总结 MALLVI通过多智能体协作实现闭环反馈驱动的机器人操作,提升泛化能力和零样本任务成功率。

Comments Some fundemental change in text and codebase

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15408 2026-06-09 cs.CV cs.AI cs.CL cs.LG 版本更新 75%

CURE: Curriculum-guided Multi-task Training for Reliable Anatomy Grounded Report Generation

CURE:基于课程引导的多任务训练实现可靠的解剖学接地报告生成

Pablo Messina, Andrés Villa, Juan León Alcázar, Karen Sánchez, Carlos Hinojosa, Denis Parra, Álvaro Soto, Bernard Ghanem

机构 * Pontificia Universidad Católica de Chile(智利天主教大学) CENIA iHEALTH KAUST(科威特皇家科学与技术局)

专题命中 视觉定位与Grounding :vision-language model(abstract);grounding(abstract);分类 cs.CV、cs.AI、cs.LG

AI总结 提出CURE框架,通过课程学习动态调整多任务训练,提升医学报告生成的视觉接地准确性和事实一致性,无需额外数据。

Comments 31 pages, 7 figures, accepted to CVPR 2026 (oral)

Journal ref Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2026, pp. 36279-36289

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.04244 2026-06-04 cs.AI cs.CL cs.CV cs.LG 75%

VAMPS: Visual-Assisted Mathematical Problem Solving Benchmark

VAMPS: 视觉辅助数学问题求解基准

Amirhossein Dabiriaghdam, Shayan Vassef, Mohammadreza Bakhtiari, Yasamin Medghalchi, Ilker Hacihaliloglu, Mesrob Ohannessian, Lele Wang, Giuseppe Carenini

机构 * University of California, Berkeley(加州大学伯克利分校)

专题命中 视觉定位与Grounding :grounding(abstract);multimodal large language model(abstract);分类 cs.CV、cs.AI、cs.LG

AI总结 提出VAMPS基准,通过1,168道双语多选题评估多模态大模型在借助绘图工具进行数学推理时的表现,发现直接解析求解优于工具辅助视觉求解。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.31145 2026-06-01 cs.CV cs.AI cs.LG 75%

FOCUS: Forcing In-Context Object Localization through Visual Support Constraints and Policy Optimization

FOCUS: 通过视觉支持约束和策略优化强制上下文目标定位

Mohammed Asad Karim, Vinay Kumar Verma

机构 * Amazon, Seattle, USA(亚马逊(美国西雅图))

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract_cn);分类 cs.CV、cs.AI、cs.LG

AI总结 提出一种两阶段训练框架,通过优化支持框与查询图像间的上下文注意力并结合GRPO强化学习,实现无类别监督的类别无关上下文目标定位,7B模型性能超越72B模型。

Comments Accepted at ICML 2026. * Equal Contributions

详情

展开后加载摘要…

URL PDF HTML 收藏