arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2026-02-17 至 2026-02-17 共收录 20 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 20 篇

2602.13712 2026-02-17 cs.CV cs.LG 84%

Fine-tuned Vision Language Model for Localization of Parasitic Eggs in Microscopic Images

用于微镜图像中寄生虫卵定位的微调视觉语言模型

Chan Hao Sien, Hezerul Abdul Karim, Nouar AlDahoul

专题命中 视觉定位与Grounding :vision language model(title,abstract);VLM(abstract);分类 cs.CV、cs.LG

AI总结 本文提出一种微调的视觉语言模型,用于自动定位显微图像中的寄生虫卵,实验结果显示其在mIOU指标上优于其他目标检测方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.09980 2026-02-17 cs.CV cs.RO 83%

V2V-LLM: Vehicle-to-Vehicle Cooperative Autonomous Driving with Multimodal Large Language Models

V2V-LLM:基于多模态大语言模型的车与车协同自动驾驶

Hsu-kuang Chiu, Ryo Hachiuma, Chien-Yi Wang, Stephen F. Smith, Yu-Chiang Frank Wang, Min-Hung Chen

机构 * NVIDIA Carnegie Mellon University(卡内基梅隆大学)

专题命中 视觉定位与Grounding :multimodal large language model(title,abstract);grounding(abstract);分类 cs.CV

AI总结 本文提出基于多模态大语言模型的V2V-LLM,通过车与车协同感知提升自动驾驶安全性和性能。

Comments Accepted by ICRA 2026 (IEEE International Conference on Robotics and Automation). Project: https://eddyhkchiu.github.io/v2vllm.github.io/ Code: https://github.com/eddyhkchiu/V2V-LLM Dataset: https://huggingface.co/datasets/eddyhkchiu/V2V-GoT-QA

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13313 2026-02-17 cs.CV cs.AI 81%

Agentic Spatio-Temporal Grounding via Collaborative Reasoning

基于协作推理的代理时空 grounding

Heng Zhao, Yew-Soon Ong, Joey Tianyi Zhou

机构 * CFAR, IHPC, Agency for Science, Technology and Research(A*STAR)(CFAR、IHPC、新加坡科技研究局) CCDS, Nanyang Technological University(CCDS、南洋理工大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV、cs.AI

AI总结 本文提出ASTG框架,通过协作推理实现开放世界下的时空视频grounding,提升检索效率并优于现有弱监督和零样本方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13961 2026-02-17 cs.CV astro-ph.IM cs.CL 79%

MarsRetrieval: Benchmarking Vision-Language Models for Planetary-Scale Geospatial Retrieval on Mars

MarsRetrieval: 用于火星尺度地理空间检索的视觉-语言模型基准测试

Shuoyuan Wang, Yiran Wang, Hongxin Wei

机构 * Department of Statistics and Data Science(统计与数据科学系) Department of Earth and Space Sciences(地球与空间科学系)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV

AI总结 MarsRetrieval提出了一种用于评估视觉-语言模型在火星地理空间发现中检索能力的基准测试,强调领域特定微调对可推广发现的重要性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13800 2026-02-17 cs.RO cs.HC 78%

Ontological grounding for sound and natural robot explanations via large language models

通过大语言模型实现声音和自然的机器人解释的本体论基础

Alberto Olivares-Alarcos, Muhammad Ahsan, Satrio Sanjaya, Hsien-I Lin, Guillem Alenyà

机构 * Institute of Electrical and Control Engineering, National Yang Ming Chiao Tung University(电子与控制工程研究所,国立阳明交通大学)

专题命中 视觉定位与Grounding :grounding(title,abstract)

AI总结 本文提出结合本体推理与大语言模型的框架,以生成自然且语义准确的机器人解释,提升人机交互的透明度和可解释性。

Comments An extended abstract of this article is accepted for presentation at AAMAS 2026: Olivares-Alarcos, A., Muhammad, A., Sanjaya, S., Lin, H. and Alenyà, G. (2026). Blending ontologies and language models to generate sound and natural robot explanations. In Proceedings of the International Conference on Autonomous Agents and Multiagent Systems. IFAAMAS

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18487 2026-02-17 cs.RO cs.CV cs.LG 76%

Grounding Bodily Awareness in Visual Representations for Efficient Policy Learning

在视觉表示中 grounding 身体意识以实现高效的策略学习

Junlin Wang, Zhiyun Lin

机构 * SUSTech Shenzhen(深圳科技大学)

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV、cs.LG

AI总结 本文提出ICon方法,通过对比学习提升机器人操作中策略学习的效率和跨机器人迁移能力。

Comments A preprint version

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13758 2026-02-17 cs.CV cs.AI 73%

OmniScience: A Large-scale Multi-modal Dataset for Scientific Image Understanding

OmniScience: 一个大规模多模态数据集用于科学图像理解

Haoyi Tao, Chaozheng Huang, Nan Wang, Han Lyu, Linfeng Zhang, Guolin Ke, Xi Fang

机构 * DP Technology(DP技术)

专题命中 视觉定位与Grounding :grounding(abstract);multimodal large language model(abstract);分类 cs.CV、cs.AI

AI总结 OmniScience是一个大规模多模态数据集,通过动态模型路由生成高信息密度的图像标题,提升多模态模型在科学图像理解上的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.14901 2026-02-17 cs.LG cs.AI cs.CV cs.MA 67%

Picking the Right Specialist: Attentive Neural Process-based Selection of Task-Specialized Models as Tools for Agentic Healthcare Systems

选择合适的专家:基于神经过程的注意力机制用于选择任务专用模型作为智能医疗系统工具

Pramit Saha, Joshua Strong, Mohammad Alsharid, Divyanshu Mishra, J. Alison Noble

机构 * Department of Engineering Science, University of Oxford, United Kingdom(牛津大学工程科学系) Department of Computer Science, Khalifa University, Abu Dhabi, United Arab Emirates(哈利法大学计算机科学系)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI、cs.LG

AI总结 本文提出ToolSelect,一种基于神经过程和注意力机制的模型选择方法,用于智能医疗系统中选择任务专用模型,通过实验展示其在不同任务上的优越性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11087 2026-02-17 cs.LG cs.AI cs.CL 62%

Enhancing Delta Compression in LLMs via SVD-based Quantization Error Minimization

通过基于SVD的量化误差最小化增强LLM中的delta压缩

Boya Xiong, Shuo Wang, Weifeng Ge, Guanhua Chen, Yun Chen

机构 * Shanghai University of Finance(上海财经大学) Tsinghua University(清华大学) Fudan University(复旦大学) Southern University of Science(南方科技大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG

AI总结 PrinMix通过基于SVD的量化误差最小化方法,提升LLM中delta压缩的效率与性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13419 2026-02-17 q-bio.QM cs.AI cs.CL cs.LG q-bio.BM 62%

Protect$^*$: Steerable Retrosynthesis through Neuro-Symbolic State Encoding

Protect$^*$: 通过神经符号状态编码实现可操控的逆合成

Shreyas Vinaya Sathyanarayana, Shah Rahil Kirankumar, Sharanabasava D. Hiremath, Bharath Ramsundar

机构 * Deep Forest Sciences(深林科技) Departament de Farmacologia, Toxicologia i Química Terapèutica, Universitat de Barcelona(巴塞罗那大学药理学、毒理学与治疗化学系) Institut de Nanociència i Nanotecnologia IN2UB, Universitat de Barcelona(巴塞罗那大学纳米科学与纳米技术研究所)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG

AI总结 Protect$^*$通过神经符号状态编码实现逆合成的可控生成,结合规则推理与神经模型,提升化学合成路径的可靠性与准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.14929 2026-02-17 cs.CV 57%

Wrivinder: Towards Spatial Intelligence for Geo-locating Ground Images onto Satellite Imagery

Wrivinder:迈向空间智能:将地面图像定位到卫星图像

Chandrakanth Gudavalli, Tajuddin Manhar Mohammed, Abhay Yadav, Ananth Vishnu Bhaskar, Hardik Prajapati, Cheng Peng, Rama Chellappa, Shivkumar Chandrasekaran, B. S. Manjunath

机构 * Mayachitra, Inc.(Mayachitra公司) Johns Hopkins University(约翰霍普金斯大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

AI总结 Wrivinder通过聚合地面影像与卫星图像,实现零样本的几何驱动定位,提供首个全面的跨视角对齐基准。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.14890 2026-02-17 cs.AI 57%

Lifted Relational Probabilistic Inference via Implicit Learning

通过隐式学习实现提升的关系概率推断

Luise Ge, Brendan Juba, Kris Nilsson, Alison Shao

机构 * Washington University in St. Louis(华盛顿大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

AI总结 本文提出了一种基于隐式学习的方法,实现第一-order概率逻辑的隐式学习和提升推断,解决了传统方法在模型构建上的局限性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.14833 2026-02-17 eess.SP cs.LG 57%

RF-GPT: Teaching AI to See the Wireless World

RF-GPT:教AI看见无线世界

Hang Zou, Yu Tian, Bohao Wang, Lina Bariah, Samson Lasaulce, Chongwen Huang, Mérouane Debbah

机构 * Research Institute for Digital Future, Khalifa University(数字未来研究院,哈利法大学) College of Information Science and Electronic Engineering, Zhejiang University(信息科学与电子工程学院,浙江大学) Université de Lorraine, CNRS, CRAN(洛林大学,CNRS,CRAN)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG

AI总结 RF-GPT通过多模态LLM的视觉编码器处理RF频谱图,实现射频信号的高级推理与生成,无需人工标注即可在多个无线任务中取得优异表现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06185 2026-02-17 cs.AI 57%

Dataforge: Agentic Platform for Autonomous Data Engineering

Dataforge:自主数据工程的代理平台

Xinyuan Wang, Hongyu Cao, Kunpeng Liu, Yanjie Fu

机构 * Arizona State University(亚利桑那州立大学) Clemson University(克莱姆森大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

AI总结 Dataforge是一个基于大语言模型的自动数据工程平台,通过自动数据清理和迭代优化特征操作,提升表格数据的AI准备质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.09027 2026-02-17 cs.CV 57%

Measure Twice, Cut Once: A Semantic-Oriented Approach to Video Temporal Localization with Video LLMs

两次测量,一次切割:一种面向语义的视频时间定位方法与视频LLMs

Zongshang Pang, Mayu Otani, Yuta Nakashima

机构 * The University of Osaka(大阪大学) CyberAgent, Inc.(CyberAgent公司)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

AI总结 MeCo提出一种面向语义的视频时间定位方法,通过生成和判别学习任务提升视频LLMs对事件结构的识别能力,实现更精确的时间分割。

Comments ICLR2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13320 2026-02-17 cs.AI 57%

Information Fidelity in Tool-Using LLM Agents: A Martingale Analysis of the Model Context Protocol

工具使用LLM代理中的信息保真度:模型上下文协议的鞅分析

Flint Xiaofeng Fan, Cheston Tan, Roger Wattenhofer, Yew-Soon Ong

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

AI总结 本文提出了一种分析模型上下文协议中误差累积的理论框架,通过鞅分析证明误差呈现线性增长并受$O(\sqrt{T})$约束,实验验证了语义加权和周期性再定位对误差控制的有效性。

Comments Full working version of an extended abstract accepted at the 25th International Conference on Autonomous Agents and Multiagent Systems (AAMAS 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03474 2026-02-17 cs.CV 57%

Procedural Mistake Detection via Action Effect Modeling

通过动作效果建模实现过程性错误检测

Wenliang Guo, Yujiang Pu, Yu Kong

机构 * Michigan State University(密歇根州立大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

AI总结 本文提出通过动作效果建模实现过程性错误检测,结合语义和视觉信息,在单类分类任务中取得最佳性能。

Comments Accepted by ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.14653 2026-02-17 cs.CL 50%

Is Information Density Uniform when Utterances are Grounded on Perception and Discourse?

当话语基于感知和话语进行 grounding 时,信息密度是否均匀?

Matteo Gay, Coleman Haley, Mario Giulianelli, Edoardo Ponti

专题命中 视觉定位与Grounding :grounding(abstract)

AI总结 本研究首次探讨了基于感知和话语的视觉环境对信息密度均匀性的影响,发现 grounding 能提高信息分布的均匀性,并在话语单元开始处产生最大的惊奇度降低。

Comments Accepted as main paper at EACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.11848 2026-02-17 physics.geo-ph 50%

Groundwater dynamics beneath a marine ice sheet

冰架下方的地下水动态

Gabriel Cairns, Graham Benham, Ian Hewitt

专题命中 视觉定位与Grounding :grounding(abstract)

AI总结 研究通过建立数学模型,揭示了冰架下方沉积盆地中地下水动态受海水入侵和盆地几何形状的影响,发现海水可能被困于盆地中,并影响冰流流动。

Comments 33 pages, 11 figures

Journal ref The Cryosphere, 19, 3725-3747, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13833 2026-02-17 cs.RO 50%

Semantic-Contact Fields for Category-Level Generalizable Tactile Tool Manipulation

语义接触场用于类别级可泛化的触觉工具操作

Kevin Yuchen Ma, Heng Zhang, Weisi Lin, Mike Zheng Shou, Yan Wu

专题命中 视觉定位与Grounding :grounding(abstract)

AI总结 SCFields通过融合视觉语义与密集接触估计,实现类别级可泛化的触觉工具操作,显著优于视觉和原始触觉基线。

详情

展开后加载摘要…

URL PDF HTML 收藏