arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 7464 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7464 篇

2501.06828 2025-03-14 cs.CV 77%

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing

Ruizhe Ou, Yuan Hu, Fan Zhang, Jiaxin Chen, Yu Liu

专题命中 视觉定位与Grounding :visual question answering(abstract);grounding(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.08625 2025-03-12 cs.CV 77%

SegAgent: Exploring Pixel Understanding Capabilities in MLLMs by Imitating Human Annotator Trajectories

Muzhi Zhu, Yuzhuo Tian, Hao Chen, Chunluan Zhou, Qingpei Guo, Yang Liu, Ming Yang, Chunhua Shen

专题命中 视觉定位与Grounding :visual reasoning(abstract);grounding(abstract);MLLM(abstract);分类 cs.CV

Comments CVPR2025;Code will be released at \url{https://github.com/aim-uofa/SegAgent}

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.07002 2025-03-11 cs.CV 77%

Taking Notes Brings Focus? Towards Multi-Turn Multimodal Dialogue Learning

Jiazheng Liu, Sipeng Zheng, Börje F. Karlsson, Zongqing Lu

专题命中 视觉定位与Grounding :grounding(abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.06430 2025-01-14 cs.CV 77%

Open Eyes, Then Reason: Fine-grained Visual Mathematical Understanding in MLLMs

Shan Zhang, Aotian Chen, Yanpeng Sun, Jindong Gu, Yi-Yu Zheng, Piotr Koniusz, Kai Zou, Anton van den Hengel, Yuan Xue

专题命中 视觉定位与Grounding :grounding(abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.09502 2025-01-09 cs.CV 77%

One missing piece in Vision and Language: A Survey on Comics Understanding

Emanuele Vivoli, Mohamed Ali Souibgui, Andrey Barsky, Artemis LLabrés, Marco Bertini, Dimosthenis Karatzas

专题命中 视觉定位与Grounding :vision-language model(abstract);visual question answering(abstract);grounding(abstract);分类 cs.CV

Comments under review. project website: https://github.com/emanuelevivoli/awesome-comics-understanding

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.04003 2025-01-08 cs.CV cs.RO 77%

Are VLMs Ready for Autonomous Driving? An Empirical Study from the Reliability, Data, and Metric Perspectives

Shaoyuan Xie, Lingdong Kong, Yuhao Dong, Chonghao Sima, Wenwei Zhang, Qi Alfred Chen, Ziwei Liu, Liang Pan

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);grounding(abstract);分类 cs.CV

Comments Preprint; 41 pages, 32 figures, 16 tables; Project Page at https://drive-bench.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.17453 2024-11-04 cs.CV 77%

VILA$^2$: VILA Augmented VILA

Yunhao Fang, Ligeng Zhu, Yao Lu, Yan Wang, Pavlo Molchanov, Jan Kautz, Jang Hyun Cho, Marco Pavone, Song Han, Hongxu Yin

专题命中 视觉定位与Grounding :VLM(abstract);visual language model(abstract);grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.16048 2024-10-29 cs.HC cs.AI 77%

GUIDE: Graphical User Interface Data for Execution

Rajat Chawla, Adarsh Jha, Muskaan Kumar, Mukunda NS, Ishaan Bhola

专题命中 视觉定位与Grounding :grounding(abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.AI

Comments 11 pages, 8 figures, 3 Tables and 1 Algorithm

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.14874 2024-09-09 cs.CV 77%

Open-Vocabulary Object Detectors: Robustness Challenges under Distribution Shifts

Prakash Chandra Chhipa, Kanjar De, Meenakshi Subhash Chippa, Rajkumar Saini, Marcus Liwicki

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);grounding(abstract);分类 cs.CV

Comments Accepted at 2024 European Conference on Computer Vision Workshops (ECCVW). Project page - https://prakashchhipa.github.io/projects/ovod_robustness

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.13811 2024-07-22 cs.CV cs.RO 77%

Which objects help me to act effectively? Reasoning about physically-grounded affordances

Anne Kemmeren, Gertjan Burghouts, Michael van Bekkum, Wouter Meijer, Jelle van Mil

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);grounding(abstract);分类 cs.CV

Comments 10 pages

Journal ref Robotics: Science and Systems. Semantic Reasoning and Goal Understanding in Robotics 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.04212 2024-07-08 cs.AI 77%

Smart Vision-Language Reasoners

Denisa Roberts, Lucas Roberts

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);grounding(abstract);分类 cs.AI

Comments Accepted in ICML 2024 MATH AI Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.12225 2024-06-19 cs.CV 77%

The Solution for CVPR2024 Foundational Few-Shot Object Detection Challenge

Hongpeng Pan, Shifeng Yi, Shouwei Yang, Lei Qi, Bing Hu, Yi Xu, Yang Yang

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);multimodal large language model(abstract);分类 cs.CV

Comments CVPR2024 Foundational Few-Shot Object Detection Challenge

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.13388 2024-06-07 cs.CV 77%

UNIMO-G: Unified Image Generation through Multimodal Conditional Diffusion

Wei Li, Xue Xu, Jiachen Liu, Xinyan Xiao

专题命中 视觉定位与Grounding :grounding(abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV

Comments Accepted by ACL 2024, Main Conference, Long Paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.01307 2024-04-02 cs.RO cs.CV 77%

SAGE: Bridging Semantic and Actionable Parts for GEneralizable Manipulation of Articulated Objects

Haoran Geng, Songlin Wei, Congyue Deng, Bokui Shen, He Wang, Leonidas Guibas

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.02969 2024-03-26 cs.CV 77%

Multi-modal Instruction Tuned LLMs with Fine-grained Visual Perception

Junwen He, Yifan Wang, Lijun Wang, Huchuan Lu, Jun-Yan He, Jin-Peng Lan, Bin Luo, Xuansong Xie

专题命中 视觉定位与Grounding :grounding(abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV

Comments CVPR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.08657 2024-02-14 cs.CV 77%

PIN: Positional Insert Unlocks Object Localisation Abilities in VLMs

Michael Dorkenwald, Nimrod Barazani, Cees G. M. Snoek, Yuki M. Asano

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.07704 2023-10-12 cs.CV cs.CL 77%

Ferret: Refer and Ground Anything Anywhere at Any Granularity

Haoxuan You, Haotian Zhang, Zhe Gan, Xianzhi Du, Bowen Zhang, Zirui Wang, Liangliang Cao, Shih-Fu Chang, Yinfei Yang

专题命中 视觉定位与Grounding :grounding(abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV

Comments 30 pages, 10 figures. Code/Project Website: https://github.com/apple/ml-ferret

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.07353 2026-08-10 cs.CL cs.AI cs.IR cs.LG 新提交 76%

Geo-Spatial Concept Probing of Large Language Models: Abstraction, Compositionality, and Grounding

大语言模型的地理空间概念探测:抽象性、组合性与接地性

Karim Radouane, Jose G Moreno, Lynda Tamine

机构 * University of Toulouse(图卢兹大学) IRIT(信息科学与技术研究院(IRIT))

专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI、cs.LG

AI总结 该研究针对LLMs的抽象性、组合性与接地性设计测试,构建空间概念基准并开展多模型实验,揭示当前LLMs的概念理解局限,为相关模型的优化提供洞见。

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.06713 2026-08-10 cs.AI cs.LG 新提交 76%

MolBioKG: Grounding Out-of-Graph Molecules in Biomedical Knowledge Graphs via Multi-Resolution Structural Anchoring

MolBioKG:通过多分辨率结构锚定将图外分子锚定到生物医学知识图谱中

Yiming Zhang, Hikaru Shindo, Shuan Chen, Kaushalya Madhawa, Jun Jin Choong, Yuna Oikawa, Takashi Fujiwara, Keisuke Ozawa

专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI、cs.LG

AI总结 MolBioKG是解决生物医学KG中图外分子冷启动问题的两层系统,通过多分辨率结构锚定实现分子与KG的关联,在多项任务中优于基线并提升关键指标。

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.01663 2026-08-04 cs.CV cs.AI 新提交 76%

Few-Shot Concept Prompt Learning for Segmentation Foundation Models via Visual Grounding

基于视觉定位的分割基础模型少样本概念提示学习

Rahul Venkataramani, Rachana Sathish

机构 * Advanced Technology Group(先进技术集团) GE HealthCare(GE医疗)

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV、cs.AI

AI总结 该研究针对分割基础模型在临床任务中表现不足的问题,提出FS-CPL方法,通过视觉定位学习概念提示,在四个医学影像基准上取得显著性能提升,且与骨干网络无关。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.21615 2026-07-27 cs.AI cs.LG 新提交 76%

FrED: External Data Influence Estimation via Domain Knowledge Graph Grounding

FrED:通过领域知识图谱基础进行外部数据影响估计

Theodoros Aivalis, Iraklis A. Klampanos, Antonis Troumpoukis, Joemon M. Jose

机构 * National Centre for Scientific Research “Demokritos”(国家科学研究中心“德谟克利特”) University of Glasgow(格拉斯哥大学)

专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI、cs.LG

AI总结 针对生成式AI训练数据归因问题,提出在黑盒设置下运行的概率框架,融合连续特征相似度与领域特定知识图谱,在艺术图像合成和天气预报等领域评估,证明其有效性,为外部数据影响分析提供高效可解释机制。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.20533 2026-07-24 cs.LG cs.AI 新提交 76%

Grounding Investor Views: Neural Predicates in the Black-Litterman Model

在布莱克-利特曼模型中锚定投资者观点:神经谓词

Marcos Florencio

专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI、cs.LG

AI总结 研究在布莱克-利特曼模型下投资组合构建问题,提出用神经谓词作为观点生成机制,将结构化金融分析数据经其处理后映射到模型相关矩阵,方法可解释且完全可微,实现端到端学习。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06001 2026-07-03 cs.RO cs.AI cs.CV 版本更新 76%

Restoring Linguistic Grounding in VLA Models via Train-Free Attention Recalibration

恢复VLA模型中的语言基础:无需训练的注意力重校准方法

Ninghao Zhang, Bin Zhu, Shijie Zhou, Jingjing Chen

机构 * Tsinghua University(清华大学) Singapore Management University(新加坡管理学院) Institute of Trustworthy Embodied AI, Fudan University(复旦大学可信具身人工智能研究院) Shanghai Key Laboratory of Multimodal Embodied AI(上海多模态具身人工智能重点实验室)

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV、cs.AI

AI总结 针对VLA模型在分布外指令下出现“语言盲视”问题,提出无需训练的注意力重校准方法IGAR,通过调整注意力分布恢复语言指令影响,在LIBERO基准和真实机器人上有效减少错误执行。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.20769 2026-06-23 cs.CL cs.AI cs.LG 新提交 76%

FirstPass: Grounding AI Scientific Judgment in Multi-Round Editorial Outcomes

FirstPass: 在多轮编辑结果中奠定AI科学判断的基础

Prabhjot Singh, Somnath Luitel, Manmeet Singh, Josh Durkee

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校) RediMinds Inc.(RediMinds公司) Disaster Science Operations Center, Western Kentucky University(西肯塔基大学灾害科学运营中心)

专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI、cs.LG

AI总结 针对同行评审AI在领域覆盖、对话建模和评估标准上的不足,提出FirstPass数据集和微调模型,利用Nature Communications多轮评审对话,通过响应损失掩码实现80.5%的编辑结果预测准确率,并生成接近人类长度的评审意见。

Comments Accepted at the AI for Science Workshop at the 43rd International Conference on Machine Learning (ICML 2026). 9 pages, 2 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.09855 2026-06-12 cs.MM cs.CV cs.LG 新提交 76%

MinhwaNet: Faithful but Insufficient Object Grounding in Korean Folk Painting

MinhwaNet: 韩国民俗画中忠实但不足的对象定位

Joonhyung Bae

机构 * Korea Advanced Institute of Science and Technology (KAIST)(韩国科学技术院)

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV、cs.LG

AI总结 提出MinhwaNet,通过部分级检测器生成对象证据图,发现韩国民俗画中符号列表不足以预测画作类型,而符号布局更重要,揭示了忠实但不足的解离现象。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03520 2026-06-11 cs.LG cs.AI cs.SY eess.SY 版本更新 76%

Certifiable Safe RLHF: Semantic Grounding and Fixed Penalty Constraint Optimization for Safer LLM Alignment

可认证安全RLHF:基于语义基础与固定惩罚约束优化的更安全大语言模型对齐

Kartik Pandit, Sourav Ganguly, Arnesh Banerjee, Shaahin Angizi, Arnob Ghosh

机构 * Department of Electrical and Computer Engineering(电气与计算机工程系) New Jersey Institute of Technology(新泽西理工学院) Department of Computer Engineering(计算机工程系) Heritage Institute of Technology(遗产理工学院)

专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI、cs.LG

AI总结 针对现有RLHF方法依赖奖励/成本函数和双变量调优导致性能敏感且缺乏可证明安全保证的问题,提出CS-RLHF,通过语义基础成本模型和固定惩罚约束优化,实现可认证安全对齐,效率提升至少5倍。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.09109 2026-06-09 cs.CV cs.IR cs.LG 新提交 76%

Driving Video Retrieval for Complex Queries with Structured Grounding

面向复杂查询的驾驶视频检索与结构化对齐

Manyi Yao, Sparsh Garg, Christian Shelton, Amit Roy-Chowdhury, Abhishek Aich

机构 * NEC Laboratories, America(美国NEC实验室) University of California, Riverside(加州大学河滨分校)

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV、cs.LG

AI总结 提出STRIVE-D框架,通过弱监督领域视频校准规则、融合视觉语言与关键词检索信号,在驾驶视频检索中实现高达84%的top-1准确率提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.08908 2026-06-09 cs.CV cs.AI 新提交 76%

Failure-Aware Refinement of Vision-Language Model for Lithography Defect Detection

面向光刻缺陷检测的视觉-语言模型失败感知精炼

Pangyun Jeong, Jiyeong Kong, Yuehua Hu, Dohee Jeong, Kyung-Tae Kang

机构 * Hanyang University(汉阳大学) Korea University(高丽大学) Korea Institute of Industrial Technology(韩国生产技术研究院)

专题命中 视觉定位与Grounding :vision-language model(title);分类 cs.CV、cs.AI

AI总结 提出两阶段视觉-语言框架,先微调Qwen3-VL检测缺陷,再通过训练精炼模块修正第一阶段错误,提升检测可靠性。

Comments 6 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.16253 2026-05-12 cs.CV cs.AI 76%

Grounding the Score: Explicit Visual Premise Verification for Reliable Vision-Language Process Reward Models

将评分接地:为可靠视觉语言处理奖励模型的显式视觉前提验证

Junxin Wang, Dai Guan, Weijie Qiu, Zhihang Li, Yongbo Gai, Zhengyi Yang, Mengyu Zhou, Erchao Zhao, Xiaoxi Jiang, Guanjun Jiang

机构 * Qwen Large Model Application Team, Alibaba(阿里巴巴大模型应用团队) Alibaba Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所阿里巴巴分所) Beijing University of Posts and Telecommunications(北京邮电大学)

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV、cs.AI

AI总结 本文提出EVPV方法,通过显式验证视觉前提提升视觉语言处理奖励模型的可靠性,实验表明其在多模态推理基准上提升了重排序准确率。

Comments 27 pages, 4 figures, 10 tables. Evaluated on VisualProcessBench and six multimodal reasoning benchmarks (LogicVista, MMMU, MathVerse-VO, MathVision, MathVista, WeMath). Includes ablations and causal analysis via controlled constraint corruption. Code: https://github.com/Qwen-Applications/EVPV-PRM

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.09625 2026-05-05 cs.CV cs.AI 76%

Grounding Synthetic Data Generation With Vision and Language Models

基于视觉和语言模型的合成数据生成基础

Ümit Mert Çağlar, Alptekin Temizel

机构 * Graduate School of Informatics(信息学院)

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV、cs.AI

AI总结 本文提出一种视觉-语言 grounded 框架,用于可解释的遥感合成数据增强与评估,引入 ARAS400k 数据集,包含 100k 真实图像和 300k 合成图像,用于语义分割和图像描述生成。

Comments Accepted for presentation at IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Synthetic Data for Computer Vision Workshop (SynData4CV) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏