arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 26465 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7494 篇

2406.00848 2024-06-04 cs.CV 80%

Eating Smart: Advancing Health Informatics with the Grounding DINO based Dietary Assistant App

Abdelilah Nossair, Hamza El Housni

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments The work presented in this paper was part of the proceedings for the First International Conference on Artificial Intelligence (ICATA 2024)

Journal ref Eating Smart: Advancing Health Informatics with the Grounding DINO-based Dietary Assistant App, International Journal of Scientific and Innovative Studies, June 2024, Volume 3, Number 3, Pages 26-34, Available online at IJSRIS

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.12198 2024-03-26 cs.CV 80%

Mask Grounding for Referring Image Segmentation

Yong Xien Chng, Henry Zheng, Yizeng Han, Xuchong Qiu, Gao Huang

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted by CVPR2024; Project page: https://yxchng.github.io/projects/mask-grounding

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.18924 2023-08-29 cs.AI cs.LO cs.PL 80%

Bottom-Up Grounding in the Probabilistic Logic Programming System Fusemate

Peter Baumgartner, Elena Tartaglia

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

Comments This is an extended version of the ICLP 2023 paper at ICLP2023:4654" target="_blank" rel="noopener">https://cgi.cse.unsw.edu.au/~eptcs/paper.cgi?ICLP2023:4654. It also includes an improvement to the grounding algorithm in Section 3

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.01821 2023-05-30 cs.CV 80%

Toward Explainable and Fine-Grained 3D Grounding through Referring Textual Phrases

Zhihao Yuan, Xu Yan, Zhuo Li, Xuhao Li, Yao Guo, Shuguang Cui, Zhen Li

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments New dataset for 3D visual grounding is available at https://yanx27.github.io/phraserefer/

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.02687 2022-07-07 cs.CV 80%

Team PKU-WICT-MIPL PIC Makeup Temporal Video Grounding Challenge 2022 Technical Report

Minghang Zheng, Dejie Yang, Zhongjie Ye, Ting Lei, Yuxin Peng, Yang Liu

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments 2st Place in PIC Makeup Temporal Video Grounding (MTVG) Challenge in ACM-MM 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.12346 2021-03-24 cs.CV 80%

Co-Grounding Networks with Semantic Attention for Referring Expression Comprehension in Videos

Sijie Song, Xudong Lin, Jiaying Liu, Zongming Guo, Shih-Fu Chang

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted to CVPR2021. The project page is at https://sijiesong.github.io/co-grounding

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.26641 2026-08-28 cs.CL 新提交 80%

Information-Guided Frontier Decoding: Contextual Utility-Driven Commitment in dMLLMs

信息引导的前沿解码:扩散多模态语言模型(dMLLMs)中上下文效用驱动的提交

Xingyou Fang, Jingxing Zhong, Xiaosong Yuan, Xiaofeng Zhang

机构 * Fuzhou University(福州大学) Jilin University(吉林大学) Shanghai Jiao Tong University(上海交通大学)

专题命中 视觉定位与Grounding :grounding(abstract,abstract_cn);MLLM(abstract,abstract_cn)

AI总结 针对扩散多模态语言模型解码时结构令牌过早提交削弱上下文传播的问题,提出无需训练的信息引导前沿解码策略,通过多维度排序优化候选选择,在多类基准上实现了优于现有方法的性能。

Comments Accepted to Findings of EMNLP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.21387 2026-08-25 cs.RO cs.SY eess.SY 新提交 80%

Multimodal-Language-Model-Driven Interaction and Companionship for Service Robots in Elderly-Care Facilities

面向养老机构服务机器人的多模态语言模型驱动的交互与陪伴

Ching-Chieh Liu, Cong-Thanh Vu, Yen-Chen Liu

机构 * National Cheng Kung University (NCKU)(成功大学)

专题命中 视觉定位与Grounding :VLM(summary_cn,abstract)

AI总结 该研究针对养老机构服务机器人缺乏综合陪伴与安全能力的问题,提出整合主动视觉跟人、LLM语音交互及VLM安全监测的智能陪伴机器人系统,实验验证其跟随交互与跌倒检测的有效性。

Comments Accepted to the 2026 IEEE/ASME International Conference on Advanced Intelligent Mechatronics (AIM)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.23733 2026-08-13 cs.CL 版本更新 80%

Multimodal QUD: Inquisitive Questions from Scientific Figures

多模态QUD:来自科学图表的探究性问题

Yating Wu, William Rudman, Venkata S Govindarajan, Alexandros G. Dimakis, Junyi Jessy Li

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校) Ithaca College(伊萨卡学院) UC Berkeley, BespokeLabs.ai(伯克利大学,BespokeLabs.ai)

专题命中 视觉定位与Grounding :grounding(summary_cn,abstract_cn);VLM(abstract_cn)

AI总结 本文提出多模态QUD数据集,通过结合图表与文本上下文生成探究性问题,提升多模态推理能力,实现更高质量的视觉 grounding。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.20913 2026-06-23 cs.CV cs.AI cs.LG 新提交 80%

PROTON: Prototype-Based Test-Time Online OOD Detection for Medical VLMs

PROTON: 基于原型的测试时在线OOD检测方法用于医学视觉语言模型

Abhijit Das, Nichula Wasalathilaka, Yifan Lu, Adinath Dukre, Dwarikanath Mahapatra, Shadab Khan, Imran Razzak

机构 * MBZUAI(穆罕默德·本·扎耶德人工智能大学) University of Peradeniya(佩拉德尼亚大学) Khalifa University(哈利法大学) ADIA Lab(阿布扎比投资局实验室) MedOS

专题命中 视觉定位与Grounding :VLM(abstract,abstract_cn);vision-language model(abstract);分类 cs.CV、cs.AI、cs.LG

AI总结 针对医学视觉语言模型在部署时难以检测分布外输入的问题,提出PROTON方法,通过在线原型库和自适应融合原型距离与最大概念匹配得分,无需修改模型或训练数据,在多个OOD场景下提升检测性能。

Journal ref 29th International Conference on Medical Image Computing and Computer Assisted Intervention 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.18738 2026-06-18 cs.SD 新提交 80%

GRIDEX: Grid-Grounded Forensic Explanations for Deepfake Spectrogram Analysis

GRIDEX:基于网格的深度伪造频谱图取证解释

Thi Ngan Ha Do, Tingmin Wu, Alsharif Abuadbba, Kristen Moore

机构 * CSIRO(澳大利亚联邦科学与工业研究组织)

专题命中 视觉定位与Grounding :VLM(abstract,abstract_cn);vision-language model(abstract);grounding(abstract)

AI总结 提出GRIDEX框架,通过两阶段学习(SFT+GRPO)定位频谱图异常区域并生成结构化取证解释,提升伪造检测的可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07343 2026-06-16 cs.CV cs.AI cs.LG cs.RO 版本更新 80%

Seeing Roads Through Words: A Language-Guided Framework for RGB-T Driving Scene Segmentation

通过文字看道路:一种语言引导的RGB-T驾驶场景分割框架

Ruturaj Reddy, Hrishav Bakul Barua, Junn Yong Loo, Thanh Thi Nguyen, Ganesh Krishnasamy

机构 * National University of Singapore(新加坡国立大学) University of Technology Sydney(悉尼科技大学)

专题命中 视觉定位与Grounding :VLM(abstract,abstract_cn);vision-language model(abstract);分类 cs.CV、cs.AI、cs.LG

AI总结 提出CLARITY框架,利用视觉语言模型先验动态调整RGB-T融合策略,并引入暗目标语义保留和层次化解码器,在MFNet数据集上达到62.3% mIoU和77.5% mAcc的新SOTA。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.13870 2026-06-15 cs.CV cs.AI cs.LG 新提交 80%

Mirage Probes: How Vision Models Fake Visual Understanding

幻象探针:视觉模型如何伪造视觉理解

Daniel Ben-Levi, Judah Goldfeder, Weiliang Zhao, Raz Lapid, Amit LeVi, Allen G. Roush, Ravid Shwartz-Ziv, Hod Lipson

机构 * Columbia University(哥伦比亚大学) Intuit Technion(以色列理工学院) Thoughtworks New York University(纽约大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract_cn);grounding(abstract);分类 cs.CV、cs.AI、cs.LG

AI总结 提出幻象探针框架,通过对比探针揭示视觉语言模型在无图像时也能回答问题的两种幻象行为:文本偏见和虚假图像,并证明后者需要表征级干预。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.07723 2026-06-09 cs.RO 新提交 80%

VoLo: A Physical Orchestrator for Open-Vocabulary Long-Horizon Manipulation

VoLo: 面向开放词汇长时程操控的物理编排器

Siyi Chen, Hugo Hadfield, Alex Zook, Mikaela Angelina Uy, Chan Hee Song, Erwin Coumans, Xuning Yang, Faisal Ladhak, Qing Qu, Stan Birchfield, Jonathan Tremblay, Valts Blukis

机构 * NVIDIA(英伟达) University of Michigan(密歇根大学)

专题命中 视觉定位与Grounding :VLM(summary_cn,abstract)

AI总结 提出VoLoAgent,利用VLM将VLA/WAM作为可中断工具进行物理编排,实现开放词汇长时程操控,并在新基准RoboVoLo上显著优于现有系统。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.06061 2026-06-05 cs.RO 80%

A Conversational Framework for Human-Robot Collaborative Manipulation with Distributed Generative AI models

基于分布式生成式AI模型的人机协作操作对话框架

Arash Ghasemzadeh Kakroudi, Roel Pieters

机构 * Automation Technology and Mechanical Engineering, Tampere University(自动化技术与机械工程,塔尔库大学)

专题命中 视觉定位与Grounding :VLM(abstract,abstract_cn);vision-language model(abstract);grounding(abstract)

AI总结 提出一个分布式对话框架,集成语言和视觉语言模型与ROS 2执行栈,实现从自由形式用户命令生成结构化操作请求,并通过视觉基础将图像空间目标转换为机器人框架目标,实验验证了端到端任务可靠性和延迟。

Comments Accepted to the 35th IEEE International Conference on Robot and Human Interactive Communication (RO-MAN 2026). The final published version will appear under the title "A Distributed Conversational Framework for Human-Robot Collaborative Manipulation Using Local LLMs and VLMs"

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.00985 2026-06-02 cs.RO 80%

Make Your VLA More Robust Without More Data By Interleaving Motion Planning

通过交错运动规划使您的VLA更鲁棒而无需更多数据

Dan BW Choe, Sundhar Vinodh Sangeetha, Samuel Coogan, Shreyas Kousik

机构 * Georgia Institute of Technology(佐治亚理工学院)

专题命中 视觉定位与Grounding :VLM(summary_cn,abstract)

AI总结 提出MPVI框架,将基于模型的运动规划与视觉-语言-动作模型交错结合,通过VLM完成检查和本体感受触发实现可靠切换,无需额外训练即可提升长时域移动操作任务的鲁棒性,在BEHAVIOR-1K基准上任务进度提升113%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.30957 2026-06-01 cs.RO 80%

RDGen: Demonstration Generation for High-Quality Robot Learning via Reinforcement Learning

RDGen: 通过强化学习生成高质量机器人学习的演示

Zijian Zhu, Menglin Zou, Zhuang Li, Yaojie Tu, Xinhai Sun

专题命中 视觉定位与Grounding :VLM(abstract,abstract_cn);grounding(abstract,abstract_cn)

AI总结 提出RDGen框架,利用从仿真到真实的强化学习策略生成高质量机器人演示轨迹,用于训练视觉-语言-动作模型,相比人工遥操作产生更平滑轨迹并提升下游性能。

Comments 13 pages, 4 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04343 2026-05-19 cs.IR 80%

The Personalization Paradox: Semantic Loss vs. Reasoning Gains in Agentic AI Q&A

个性化悖论:语义损失与推理增益在代理AI问答中的权衡

Satyajit Movidi, Stephen Russell

专题命中 视觉定位与Grounding :grounding(summary_cn,abstract)

AI总结 本文研究了个性化对系统性能的影响,通过比较不同配置发现个性化虽提升推理和 grounding 能力,但导致语义相似度下降,揭示了现有 LLM 评估方法的不足。

Journal ref Cloud Computing and Data Science 2026 May 18;7(2):290-313

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.10032 2026-05-13 cs.CL 80%

PlantMarkerBench: A Multi-Species Benchmark for Evidence-Grounded Plant Marker Reasoning

PlantMarkerBench: 一个多物种证据导向植物标记推理基准

Sajib Acharjee Dip, Song Li, Liqing Zhang

机构 * Department of Computer Science, Virginia Tech(弗吉尼亚理工学院计算机科学系) School of Plant and Environmental Sciences, Virginia Tech(弗吉尼亚理工学院植物与环境科学学院) Fralin Biomedical Research Institute, Virginia Tech(弗吉尼亚理工学院弗拉林生物医学研究学院) FBRI Cancer Research Center, Washington, DC(华盛顿特区FBRI癌症研究中心)

专题命中 视觉定位与Grounding :grounding(summary_cn,abstract)

AI总结 本文提出PlantMarkerBench,通过整合文献检索与生物 grounding,构建了包含4种植物的基准,评估文献支持的植物标记证据解释,发现模型在功能、间接和弱支持证据上表现欠佳。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.09802 2026-05-12 cs.CV cs.AI cs.LG 80%

CrossVL: Complexity-Aware Feature Routing and Paired Curriculum for Cross-View Vision-Language Detection

CrossVL: 用于跨视角视觉语言检测的复杂度感知特征路由与配对课程学习

Zhipeng Liu, Chunbo Luo

机构 * Department of Computer Science, University of Exeter(埃克塞特大学计算机科学系)

专题命中 视觉定位与Grounding :VLM(abstract,abstract_cn);vision-language model(abstract);分类 cs.CV、cs.AI、cs.LG

AI总结 本文提出CrossVL框架,结合复杂度感知路径聚合和配对课程学习,提升跨视角视觉语言模型的检测性能,实验显示其在MAVREC数据集上提升了aerial mAP并缩小了地面与空中视角的性能差距。

Comments Accepted to CVPR 2026. Code available at https://github.com/1nyourlife/Crossvl_cvpr2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.27600 2026-05-01 cs.IR 80%

Purifying Multimodal Retrieval: Fragment-Level Evidence Selection for RAG

净化多模态检索:用于RAG的片段级证据选择

Xihang Wang, Zihan Wang, Chengkai Huang, Cao Liu, Ke Zeng, Quan Z. Sheng, Lina Yao

专题命中 视觉定位与Grounding :MLLM(abstract,abstract_cn);grounding(abstract);multimodal large language model(abstract)

AI总结 本文提出FES-RAG框架,通过片段级证据选择提升多模态检索效果,减少噪声干扰,实验显示在M2RAG基准上性能提升27%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.25914 2026-04-29 cs.CL 80%

DV-World: Benchmarking Data Visualization Agents in Real-World Scenarios

DV-World:在现实场景中评估数据可视化代理的基准测试

Jinxiang Meng, Shaoping Huang, Fangyu Lei, Jingyu Guo, Haoxiang Liu, Jiahao Su, Sihan Wang, Yao Wang, Enrui Wang, Ye Yang, Hongze Chai, Jinming Lv, Anbang Yu, Huangjing Zhang, Yitong Zhang, Yiming Huang, Zeyao Ma, Shizhu He, Jun Zhao, Kang Liu

机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) University of Chinese Academy of Sciences(中国科学院大学) National University of Singapore(新加坡国立大学) Renmin University of China(中国人民大学)

专题命中 视觉定位与Grounding :grounding(abstract,abstract_cn);MLLM(abstract,abstract_cn)

AI总结 DV-World通过260个任务评估数据可视化代理在现实专业生命周期中的能力,涵盖表格操作、视觉进化和意图对齐,采用混合评估框架揭示现有模型在复杂数据可视化挑战中的不足。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.25323 2026-04-29 cs.RO 80%

ANCHOR: A Physically Grounded Closed-Loop Framework for Robust Home-Service Mobile Manipulation

ANCHOR:一种基于物理的闭环框架,用于鲁棒的家庭服务移动操作

Jinhao Jiang, Shengyu Fang, Sibo Zuo, Yujie Tang, Yirui Li

机构 * Beijing Institute of Technology(北京理工大学)

专题命中 视觉定位与Grounding :grounding(summary_cn,abstract)

AI总结 ANCHOR通过物理 grounding 和结构化故障处理,提升家庭服务机器人在动态环境中的任务成功率和扰动恢复能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.12387 2026-04-15 q-bio.GN 80%

oxo-call: Documentation-grounded Skill Augmentation for Accurate Bioinformatics Command-line Generation with Large Language Models

oxo-call:基于文档的技能增强用于准确的生物信息学命令行生成与大型语言模型

Yun Peng, Yujun Sun, Jia Ding, Bin Yan, Zhangyu Wang, Chunyang Wang, Chenyang Shu, Jian-Guo Zhou, Shixiang Wang

专题命中 视觉定位与Grounding :grounding(summary_cn,abstract)

AI总结 oxo-call通过文档优先 grounding 和精选技能增强策略,提升生物信息学命令行生成的准确性,提供150+内置技能和可扩展的流程引擎。

Comments 19 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17651 2025-10-21 cs.CV cs.AI cs.LG 80%

Frugal Federated Learning for Violence Detection: A Comparison of LoRA-Tuned VLMs and Personalized CNNs

Sébastien Thuau, Siba Haidar, Ayush Bajracharya, Rachid Chelouah

机构 * esieaLab(esiea实验室) ESIEA(ESIEA学院) ETIS Laboratory(ETIS实验室) CNRS(法国国家科学研究中心) UMR8051(UMR8051研究中心) University of CY Cergy(CY塞克大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);LLaVA(abstract);分类 cs.CV、cs.AI、cs.LG

Comments 7 pages, 1 figure, FLTA 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11616 2025-08-18 cs.CV cs.AI cs.CL cs.LG 80%

Controlling Multimodal LLMs via Reward-guided Decoding

Oscar Mañas, Pierluca D'Oro, Koustuv Sinha, Adriana Romero-Soriano, Michal Drozdzal, Aishwarya Agrawal

机构 * Mila - Quebec AI Institute(魁北克AI研究院) Université de Montréal(蒙特利尔大学) McGill University(麦吉尔大学) Meta FAIR Canada CIFAR AI Chair(加拿大CIFAR人工智能主席)

专题命中 视觉定位与Grounding :grounding(abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI、cs.LG

Comments Published at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.23573 2025-04-01 cs.CV cs.AI cs.LG 80%

DASH: Detection and Assessment of Systematic Hallucinations of VLMs

Maximilian Augustin, Yannic Neuhaus, Matthias Hein

机构 * Tübingen AI Center – University of Tübingen(蒂宾根人工智能中心——蒂宾根大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);LLaVA(abstract);分类 cs.CV、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.03151 2025-01-07 cs.AI cs.CV cs.LG 80%

Large language models for artificial general intelligence (AGI): A survey of foundational principles and approaches

Alhassan Mumuni, Fuseini Mumuni

机构 * Cape Coast Technical University(海岸角理工大学) University of Mines and Technology(矿业与技术大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);grounding(abstract);multimodal large language model(abstract);分类 cs.CV、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.23143 2026-08-25 cs.CV 新提交 79%

An end-to-end-trained vision-language model for native-language prostate pathology report generation

用于生成本土语言前列腺病理报告的端到端训练视觉-语言模型

Christian Grashei, Fabian Gülhan, Maximilian Legnar, Fabian Stögbauer, Cleo-Aron Weis, Carolin Mogler, Peter Schüffler

机构 * Technical University of Munich(慕尼黑工业大学) Munich Data Science Institute(慕尼黑数据科学研究所) Munich Center for Machine Learning(慕尼黑机器学习中心) University Hospital Heidelberg(海德堡大学医院) Heidelberg University(海德堡大学) Interdisciplinary Center for Scientific Computing (IWR)(跨学科科学计算中心(IWR))

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV

AI总结 该研究提出语言独立的切片级视觉-语言框架,用自动化流水线生成17344对图像-文本对,实现德语前列腺病理报告生成,恶性肿瘤检测F1达96.2%,可支持机构用自有档案训练本土语言报告模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.21878 2026-08-25 cs.CV 新提交 79%

ViSMoE: Visual-Aware Sparse Mixture-of-Experts for Embodied Referring Expression Grounding

ViSMoE:面向具身指代表达接地的视觉感知稀疏混合专家模型

Shuo Feng, Piji Li

机构 * College of Artificial Intelligence(人工智能学院) Nanjing University of Aeronautics and Astronautics(南京航空航天大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

AI总结 针对现有具身指代表达接地方法无法区分视角与物体导致表示模糊的问题,提出带视觉感知路由策略的ViSMoE框架,在REVERIE和SOON数据集上性能优于现有SOTA方法。

Comments Accepted by ICANN 2025

详情

展开后加载摘要…

URL PDF HTML 收藏