arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 26465 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7494 篇

2010.02384 2020-10-07 cs.CL 78%

Fine-Grained Grounding for Multimodal Speech Recognition

Tejas Srinivasan, Ramon Sanabria, Florian Metze, Desmond Elliott

专题命中 视觉定位与Grounding :grounding(title,abstract)

Comments Accepted to Findings of EMNLP 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2002.11938 2020-02-28 eess.SY cs.SY 78%

Analysis of Attack via Grounding and Countermeasures in Discrete-Time Consensus Networks

Yamin Yan, Sonja Stuedli, Maria M. Seron, Richard H. Middleton

专题命中 视觉定位与Grounding :grounding(title,abstract)

Comments 21st IFAC World Congress, 2020, to appear

详情

展开后加载摘要…

URL PDF HTML 收藏
2001.03830 2020-01-14 cs.CL 78%

Revisiting Challenges in Data-to-Text Generation with Fact Grounding

Hongmin Wang

专题命中 视觉定位与Grounding :grounding(title,abstract)

Comments Best Paper Runner-up at INLG 2019 (12th International Conference on Natural Language Generation)

详情

展开后加载摘要…

URL PDF HTML 收藏
2001.03671 2020-01-14 cs.CV cs.AI cs.CL cs.LG 78%

Retouchdown: Adding Touchdown to StreetLearn as a Shareable Resource for Language Grounding Tasks in Street View

Harsh Mehta, Yoav Artzi, Jason Baldridge, Eugene Ie, Piotr Mirowski

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1909.07586 2019-09-18 cs.CL 78%

Grounding learning of modifier dynamics: An application to color naming

Xudong Han, Philip Schulz, Trevor Cohn

专题命中 视觉定位与Grounding :grounding(title,abstract)

Comments EMNLP 2019 (5 pages + 1 references)

详情

展开后加载摘要…

URL PDF HTML 收藏
1909.00301 2019-09-04 cs.CL 78%

Phrase Grounding by Soft-Label Chain Conditional Random Field

Jiacheng Liu, Julia Hockenmaier

专题命中 视觉定位与Grounding :grounding(title,abstract)

Comments 11 pages, 5 figures, accepted by EMNLP-IJCNLP 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1908.10751 2019-08-30 physics.comp-ph physics.geo-ph 78%

A full Stokes subgrid model for simulation of grounding line migration in ice sheets

Gong Cheng, Per Lötstedt, Lina von Sydow

专题命中 视觉定位与Grounding :grounding(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
1906.06147 2019-07-30 cs.MM eess.IV 78%

Grounding Object Detections With Transcriptions

Yasufumi Moriya, Ramon Sanabria, Florian Metze, Gareth J. F. Jones

专题命中 视觉定位与Grounding :grounding(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
1906.00347 2019-06-11 cs.CL 78%

Are You Looking? Grounding to Multiple Modalities in Vision-and-Language Navigation

Ronghang Hu, Daniel Fried, Anna Rohrbach, Dan Klein, Trevor Darrell, Kate Saenko

专题命中 视觉定位与Grounding :grounding(title,abstract)

Comments ACL 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1906.00430 2019-06-04 cs.RO physics.app-ph 78%

Effects of Different Hand-Grounding Locations on Haptic Performance With a Wearable Kinesthetic Haptic Device

Sajid Nisar, Melisa Orta Martinez, Takahiro Endo, Fumitoshi Matsuno, Allison M. Okamura

专题命中 视觉定位与Grounding :grounding(title,abstract)

Comments 8 pages, 11 figures, 1 table

Journal ref IEEE Robotics and Automation Letters, vol. 4, no. 2, pp. 351-358, April 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1811.06966 2018-11-19 cs.RO 78%

Temporal Grounding Graphs for Language Understanding with Accrued Visual-Linguistic Context

Rohan Paul, Andrei Barbu, Sue Felshin, Boris Katz, Nicholas Roy

专题命中 视觉定位与Grounding :grounding(title,abstract)

Comments Published in ICJAI 2017

详情

展开后加载摘要…

URL PDF HTML 收藏
1808.08266 2018-08-29 cs.CL 78%

A Visual Attention Grounding Neural Model for Multimodal Machine Translation

Mingyang Zhou, Runxiang Cheng, Yong Jae Lee, Zhou Yu

专题命中 视觉定位与Grounding :grounding(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
1804.11162 2018-06-01 cs.PL cs.LO 78%

Constraint Answer Set Programming without Grounding

Joaquín Arias, Manuel Carro, Elmer Salazar, Kyle Marple, Gopal Gupta

专题命中 视觉定位与Grounding :grounding(title,abstract)

Comments Paper presented at the 34nd International Conference on Logic Programming (ICLP 2018), Oxford, UK, July 14 to July 17, 2018 18 pages, LaTeX

详情

展开后加载摘要…

URL PDF HTML 收藏
1712.01097 2017-12-05 cs.CL cs.RO 78%

Generalized Grounding Graphs: A Probabilistic Framework for Understanding Grounded Commands

Thomas Kollar, Stefanie Tellex, Matthew Walter, Albert Huang, Abraham Bachrach, Sachi Hemachandra, Emma Brunskill, Ashis Banerjee, Deb Roy, Seth Teller, Nicholas Roy

专题命中 视觉定位与Grounding :grounding(title,abstract)

Comments Submitted to the Journal of Artificial Intelligence Research

详情

展开后加载摘要…

URL PDF HTML 收藏
1709.10486 2017-10-02 cs.CL 78%

Symbol, Conversational, and Societal Grounding with a Toy Robot

Casey Kennington, Sarah Plane

专题命中 视觉定位与Grounding :grounding(title,abstract)

Comments 2 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
1709.00501 2017-09-05 cs.LO 78%

Computing Stable Models of Normal Logic Programs Without Grounding

Kyle Marple, Elmer Salazar, Gopal Gupta

专题命中 视觉定位与Grounding :grounding(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
1608.01127 2016-08-04 cs.RO cs.AI cs.CV cs.LG 78%

Autonomous Grounding of Visual Field Experience through Sensorimotor Prediction

Alban Laflaquière

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV、cs.AI、cs.LG

Comments 6 pages, 4 figures, ICDL-Epirob 2016

详情

展开后加载摘要…

URL PDF HTML 收藏
1505.06289 2015-06-08 cs.CL cs.GR 78%

Text to 3D Scene Generation with Rich Lexical Grounding

Angel Chang, Will Monroe, Manolis Savva, Christopher Potts, Christopher D. Manning

专题命中 视觉定位与Grounding :grounding(title,abstract)

Comments 10 pages, 7 figures, 3 tables. To appear in ACL-IJCNLP 2015

详情

展开后加载摘要…

URL PDF HTML 收藏
1402.6889 2015-02-04 cs.LO 78%

Lazy Model Expansion: Interleaving Grounding with Search

Broes De Cat, Marc Denecker, Peter Stuckey, Maurice Bruynooghe

专题命中 视觉定位与Grounding :grounding(title,abstract)

Journal ref Journal of Artificial Intelligence Research, feb 2015, volume 52, pages 235-286

详情

展开后加载摘要…

URL PDF HTML 收藏
1311.5076 2013-11-21 physics.ins-det physics.plasm-ph 78%

Design of a mechanically actuated RF grounding system for the ITER ICRH antenna

D Hancock, M Shannon, B Beaumont, P Dumortier, F Durodie, V Kyrytsya, F Louche, R McKinley, K Nicholls, the CYCLE Team

专题命中 视觉定位与Grounding :grounding(title,abstract)

Comments 4 pages, 11 figures

Journal ref Proceedings of the 27th Symposium On Fusion Technology (SOFT-27); Liege, Belgium, September 24-28, 2012. Fusion Engineering and Design, Vol.88, Issues 9-10, October 2013, p.2100-2104

详情

展开后加载摘要…

URL PDF HTML 收藏
1111.1570 2011-11-08 cs.IR cs.SI 78%

Semantic Grounding Strategies for Tagbased Recommender Systems

Frederico Durao, Peter Dolog

专题命中 视觉定位与Grounding :grounding(title,abstract)

Comments 13 pages, 5 figures

Journal ref International Journal of Web & Semantic Technology (IJWesT) Vol.2, No.4, 2011, 67-79

详情

展开后加载摘要…

URL PDF HTML 收藏
0906.2756 2010-11-09 cs.MA cs.LO cs.SE 78%

Norms and Commitment for iOrgs(TM) Information Systems: Direct Logic(TM) and Participatory Grounding Checking

Carl Hewitt

专题命中 视觉定位与Grounding :grounding(title,abstract)

Comments expanded article

详情

展开后加载摘要…

URL PDF HTML 收藏
astro-ph/0301095 2009-12-01 astro-ph 78%

Detection of Nine M8.0-L0.5 Binaries: The Very Low Mass Binary Population and its Implications for Brown Dwarf and VLM Star Formation

Laird M. Close, Nick Siegler, Melanie Freed, Beth Biller

专题命中 视觉定位与Grounding :VLM(title,abstract)

Comments To appear in the April 10, 2003 issue of The Astrophysical Journal 30 pages, 17 figures

Journal ref Astrophys.J. 587 (2003) 407-422

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.06179 2022-08-15 cs.CV cs.AI 77%

Exploiting Feature Diversity for Make-up Temporal Video Grounding

Xiujun Shu, Wei Wen, Taian Guo, Sunan He, Chen Wu, Ruizhi Qiao

专题命中 视觉定位与Grounding :grounding(title,comments);分类 cs.CV、cs.AI

Comments 3st Place in PIC Makeup Temporal Video Grounding (MTVG) Challenge in ACM-MM 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.26722 2026-08-28 cs.CV 新提交 77%

UniGeo: A Multi-modal Large Language Model for Text-Guided Cross-View Geo-Localization

UniGeo:用于文本引导跨视角地理定位的多模态大语言模型

Jiahao Wen, Hang Yu, Zhedong Zheng

机构 * School of Computer Engineering and Science, Shanghai University(上海大学计算机工程与科学学院) Institute of Collaborative Innovation, University of Macau(澳门大学协同创新研究院)

专题命中 视觉定位与Grounding :MLLM(abstract,abstract_cn);multimodal large language model(abstract);分类 cs.CV

AI总结 UniGeo是一种统一多模态大语言模型,通过地理语义学习、跨视角生成及即插即用验证模块,在文本引导无人机地理定位任务中显著提升了检索性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.26163 2026-08-28 cs.CL cs.AI cs.HC cs.MA cs.SD 新提交 77%

From Sound to Symptom: Real-Time Respiratory Signal Understanding for Conversational Healthcare Agents

从声音到症状:面向对话式医疗智能体的实时呼吸信号理解

Tanmay Laud, Herprit Mahal, Subhabrata Mukherjee

机构 * Hippocratic AI(希波克拉底人工智能公司)

专题命中 视觉定位与Grounding :MLLM(abstract,abstract_cn);multimodal large language model(abstract);分类 cs.AI

AI总结 提出面向对话式医疗智能体的HealthCUES系统,可实时检测分析咳嗽等呼吸信号,经内部与外部数据集评估及医疗人员验证,性能优异且具临床实用性。

Comments Accepted for publication at SIGDIAL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.20492 2026-08-24 cs.CV 新提交 77%

Annotations as Rollouts: Efficient and Scalable Reinforcement Learning for Video MLLMs

以标注为回滚:面向视频多模态大语言模型的高效可扩展强化学习

Yunheng Li, Guohong Mu, Hao Li, Shengsheng Qian, Dingwen Zhang, Qibin Hou, Ming-Ming Cheng

机构 * Nankai University(南开大学) Northwestern Polytechnical University(西北工业大学) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) NKIARI

专题命中 视觉定位与Grounding :MLLM(summary_cn);multimodal large language model(abstract);分类 cs.CV

AI总结 本文提出OraRL算法,将标注作为神谕回滚解耦优势估计,实现高效可扩展的视频MLLM强化学习,在多项视频感知基准上超越现有模型,解码速度大幅提升。

Comments Project page: https://orarl.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.01460 2026-08-19 cs.CV 版本更新 77%

Reinforcing Consistency in Video MLLMs with Structured Rewards

通过结构化奖励强化视频MLLMs的一致性

Yihao Quan, Zeru Shi, Jinman Zhao, Ruixiang Tang

机构 * Rutgers University(罗格斯大学) University of Toronto(多伦多大学)

专题命中 视觉定位与Grounding :grounding(abstract,abstract_cn);multimodal large language model(abstract);分类 cs.CV

AI总结 研究通过结构化奖励提升视频MLLMs的一致性,发现传统监督不足,提出结合事实和时间单元的奖励机制,提升视频理解准确性。

Comments Accepted by COLM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19759 2026-08-19 cs.CV 版本更新 77%

Vision-Language Enhanced Foundation Model for Semi-Supervised Medical Image Segmentation

增强视觉-语言能力的半监督医学图像分割基础模型

Jiaqi Guo, Mingzhen Li, Hanyu Su, Keigo Healy, Lexiaozi Fan, Neda Tavakoli, Santiago López-Tapia, Daniel Kim, Aggelos K. Katsaggelos

机构 * ECE, Northwestern University(电气工程与计算机科学系,西北大学) Stats, Northwestern University(统计学系,西北大学) Radiology, Northwestern University(放射学系,西北大学)

专题命中 视觉定位与Grounding :VLM(abstract,abstract_cn);vision-language model(abstract);分类 cs.CV

AI总结 本文提出VESSA模型,通过增强视觉-语言能力的半监督方法提升医学图像分割精度,实验表明其在有限标注条件下表现优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.13343 2026-08-14 cs.CV 新提交 77%

AmalthAI: An Open-Source Computer Vision Platform for Cultural Heritage

AmalthAI:面向文化遗产的开源计算机视觉平台

Christos Chatzisavvas, Stelios Alvanos, Efstratios Politis, Panagiotis Rigas, Thomas Pappas, Ioannis Giannoukos, Nikolaos Mitianoudis, Agata Ulanowska, Katarzyna Żebrowska, Nazarij Buławka, Christina Margariti, George Pavlidis, Chairi Kiourt, Anestis Koutsoudis, Vassilis Katsouros, George Ioannakis

机构 * Democritus University of Thrace(德谟克利特色雷斯大学) Athena Research Center(雅典娜研究中心) National and Kapodistrian University of Athens(雅典国立卡波迪斯特里亚大学) University of Warsaw(华沙大学)

专题命中 视觉定位与Grounding :VLM(abstract,abstract_cn);vision-language model(abstract);分类 cs.CV

AI总结 AmalthAI是面向文化遗产领域专家的开源计算机视觉平台,通过集成Kubeflow、Katib等工具,支持数据集管理、模型训练与推理,可保障敏感考古数据安全,已在黏土织物印痕数据集上验证其功能。

详情

展开后加载摘要…

URL PDF HTML 收藏