VIG-RL: Learning to Search and Insert for Verified Image Grounding
VIG-RL:学习搜索与插入以实现可验证的图像定位
专题命中 视觉定位与Grounding :grounding(title,abstract)
AI总结 本文针对现有检索增强框架无法动态推理视觉证据插入时机与位置的问题,提出自主智能体框架VIG-RL,将相关工作流建模为主动决策过程,经强化学习优化后实现SOTA性能。
AI 大模型
视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。
VIG-RL:学习搜索与插入以实现可验证的图像定位
专题命中 视觉定位与Grounding :grounding(title,abstract)
AI总结 本文针对现有检索增强框架无法动态推理视觉证据插入时机与位置的问题,提出自主智能体框架VIG-RL,将相关工作流建模为主动决策过程,经强化学习优化后实现SOTA性能。
通过轻量级流匹配实现大象启发式软躯干运动的全身语义到驱动的基础
专题命中 视觉定位与Grounding :grounding(title,abstract);multimodal large language model(abstract)
AI总结 针对近距离人机交互中类似躯干机器人关联开放词汇响应难的问题,提出基于轻量级流匹配的全身语义到驱动基础框架,将多模态大语言模型响应转换为元组,参数化轨迹并采样运动,实验显示其提升关联正确性、减少推理时间,还提升了人机交互满意度。
SciLens: 多模态科学声明验证的智能蕴含与归因框架
机构 * The Hong Kong University of Science and Technology(香港科技大学) ; Hong Kong Baptist University(香港浸会大学) ; NVIDIA AI Technology Center(英伟达人工智能技术中心)
专题命中 视觉定位与Grounding :grounding(title,abstract);vision-language model(abstract)
AI总结 提出SciLens框架,通过将声明分解为原子命题并归因到表格/图表证据,实现多模态科学声明验证,在SciClaimEval上达到79.2%宏F1和63.1%配对准确率。
Comments KDD 2026 SciSoc Agents & LLMs (Oral)
通过提示实现视觉语言模型发射后能力扩展用于在轨航天器检测
机构 * Florida Institute of Technology(佛罗里达理工学院) ; University of Florida(佛罗里达大学)
专题命中 视觉定位与Grounding :vision-language model(title);grounding(abstract);分类 cs.CV、cs.AI、cs.LG
AI总结 研究利用提示驱动的视觉语言模型在轨扩展语义能力,无需修改权重即可通过自然语言提示检测新航天器部件,在129张图像上零样本实例分割达到0.385 mAP@0.5。
Comments 5 pages, 1 figure, 2 tables. Equal contribution by Nicholas A. Welsh and Lennon Shikhman. Published in the CVPR2026 Workshop on AI4Space
基于视觉线索检测手语翻译中的幻觉:是依据视觉信息还是猜测?
机构 * German Research Center for Artificial Intelligence (DFKI GmbH)(德国人工智能研究中心(DFKI GmbH)) ; Saarland Informatics Campus(萨尔兰州信息学校园) ; Barcelona Supercomputing Center (BSC-CNS)(巴塞罗那超级计算中心(BSC-CNS))
专题命中 视觉定位与Grounding :grounding(title,abstract);vision-language model(abstract)
AI总结 针对手语翻译中模型依赖语言先验而非视觉输入导致幻觉的问题,提出一种基于特征敏感性和反事实信号的令牌级可靠性度量,用于量化视觉信息利用程度,并在两个基准上验证其预测幻觉率、跨数据集泛化及与文本信号结合提升风险评估的效果。
Comments Published at ICLR2026 Code available at \url{https://github.com/yhamidullah/hallucination-slt}
动态决策学习:罕见疾病异常定位的测试时间演化
机构 * Technical University of Munich(慕尼黑技术大学) ; Munich Center for Machine Learning(慕尼黑机器学习中心) ; Imperial College London(帝国理工学院伦敦分校) ; University of Trento(特伦托大学) ; King's College London(伦敦国王学院)
专题命中 视觉定位与Grounding :grounding(title,abstract);vision-language model(abstract)
AI总结 本文提出动态决策学习框架,通过优化指令和视觉扰动下的预测整合,提升冻结大视觉语言模型在罕见疾病异常定位中的表现,实验显示其在罕见疾病案例中mAP@75提升达105%。
SAGE:基于传感器的地面引擎用于LLM驱动的睡眠护理代理
专题命中 视觉定位与Grounding :grounding(title,abstract)
AI总结 SAGE通过整合传感器数据,解决睡眠护理中数据与行动之间的鸿沟问题,提升个性化和信任度。
Comments Accepted to the Extended Abstracts of the 2026 CHI Conference on Human Factors in Computing Systems (CHI EA '26). 6 pages
Journal ref Extended Abstracts of the 2026 CHI Conference on Human Factors in Computing Systems (CHI EA '26)
人类与视觉-语言模型:叙事连贯性的一种统一衡量方法
机构 * University of Gothenburg(哥德堡大学) ; Queen Mary University of London(伦敦玛丽女王大学)
专题命中 视觉定位与Grounding :vision-language model(title,abstract);grounding(abstract)
AI总结 研究通过比较人类写作与视觉语言模型生成的叙事,评估视觉基础故事中的叙事连贯性,发现模型在连贯性方面与人类存在系统性差异。
Comments 9 pages of content, 1 page of appendices, 9 tables, 3 figures
从遮罩到像素与意义:一种新的分类、基准和度量标准用于VLM图像篡改
机构 * Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学) ; University College London(伦敦大学学院)
专题命中 视觉定位与Grounding :VLM(title,abstract);分类 cs.CV、cs.AI、cs.LG
AI总结 本文提出了一种基于像素和意义的VLM图像篡改检测新方法,引入了新的分类体系和基准,提出了基于像素的度量标准,并评估了现有模型在微编辑和非遮罩篡改上的表现。
Comments Code and data at: https://github.com/VILA-Lab/PIXAR (Accepted in CVPR 2026 Findings, but not opted in)
GoalVLM: 多智能体系统中基于视觉语言模型的目标导航
机构 * Intelligent Space Robotics Laboratory(智能空间机器人实验室) ; Center for Digital Engineering(数字工程中心) ; Skolkovo Institute of Science and Technology(斯克尔科夫科学与技术研究所)
专题命中 视觉定位与Grounding :VLM(title,abstract);vision-language model(abstract)
AI总结 本文提出GoalVLM,一种基于视觉语言模型的多智能体零样本开放词汇目标导航框架,通过整合SAM3和SpaceOM实现语义优先级探索,无需重新训练即可完成复杂目标导航任务。
Comments 8 pages, 5 figures
面向机器人操作的上下文学习:考虑混淆的视觉语言模型
机构 * Ping An Technology (Shenzhen) Co., Ltd.(平安科技(深圳)有限公司) ; Shenzhen Bao'an Middle School(深圳宝安中学)
专题命中 视觉定位与Grounding :vision-language model(title,abstract);VLM(abstract)
AI总结 本文提出CAICL方法,通过混淆定位与分析提升视觉语言模型在机器人操作中处理混淆场景的能力,实验显示其在VIMA-Bench上成功率达85.5%。
Comments Accepted by the 29th International Conference on Computer Supported Cooperative Work in Design (CSCWD 2026)
ALARM: 基于多模态大语言模型的复杂环境监控异常检测与不确定性量化
机构 * Department of Industrial and Systems Engineering, University of Washington(华盛顿大学工业与系统工程系) ; Wyze Labs, Inc.(Wyze实验室)
专题命中 视觉定位与Grounding :MLLM(title,abstract);分类 cs.CV、cs.AI、cs.LG
AI总结 ALARM通过结合不确定性量化与多模态大语言模型,实现了复杂环境中的异常检测,展现出在不同领域中的高准确性和可靠性。
I-FailSense:面向通用机器人故障检测的视觉-语言模型
机构 * ISIR, Sorbonne Université, CNRS(ISIR,索邦大学,国家科学研究中心)
专题命中 视觉定位与Grounding :vision-language model(title,abstract);VLM(abstract)
AI总结 I-FailSense通过构建专门用于检测语义错位故障的数据集,提出了一种开源视觉-语言模型框架,有效提升机器人故障检测能力。
PhysBrain: 人眼视角数据作为视觉语言模型到物理智能的桥梁
机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) ; Zhongguancun Academy(中关村学院) ; Zhongguancun Institute of Artificial Intelligence(中关村人工智能研究院) ; Harbin Institute of Technology(哈尔滨工业大学) ; Huazhong University of Science and Technology(华中科技大学)
专题命中 视觉定位与Grounding :vision language model(title,abstract);grounding(abstract)
AI总结 PhysBrain通过将人类眼动视频转化为多级具身监督,提升机器人在眼动感知和长期规划中的能力,实现从人类视角到机器人控制的有效迁移。
Comments 21 pages, 8 figures
MetaWorld: 一个用于地面指令基础的分层世界模型中的技能迁移与组合
机构 * Beijing University of Technology(北京理工大学) ; Fudan University(复旦大学) ; Tsinghua University(清华大学)
专题命中 视觉定位与Grounding :grounding(title,abstract);VLM(abstract)
AI总结 MetaWorld通过整合语义规划与物理控制,利用专家策略迁移提升人形机器人在定位-操作任务中的性能。
Comments 8 pages, 4 figures, Submitted to ICLR 2026 World Model Workshop
本体感知增强视觉语言模型在为机器人任务生成描述和子任务分割中的应用
机构 * Faculty of Science and Engineering, Waseda University(工学部,早稻田大学) ; Artificial Intelligence Laboratory, Fujitsu Limited(Fujitsu 人工智能实验室) ; National Institute of Advanced Industrial Science and Technology(国家先进工业科学与技术研究院)
专题命中 视觉定位与Grounding :vision language model(title,abstract);VLM(abstract)
AI总结 本研究通过引入本体感知数据,提升视觉语言模型在机器人任务描述和子任务分割中的性能,以增强机器人模仿学习效率。
PETAR:基于掩码感知的视觉-语言建模的局部发现生成用于PET自动报告
机构 * University of Wisconsin–Madison Department of Computer Sciences(威斯康星大学麦迪逊分校计算机科学系) ; University of Wisconsin–Madison Department Radiology(威斯康星大学麦迪逊分校放射学系) ; Microsoft(微软公司)
专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV、cs.AI、cs.LG
AI总结 PETAR通过引入PETARSeg-11K数据集和PETAR-4B模型,实现基于掩码感知的3D PET自动报告生成,提升医学影像分析的精度与实用性。
重新思考基于VLM的机器人操作中的中间表示
机构 * CUHK(香港中文大学) ; Amazon(亚马逊) ; UNC(北卡罗来纳大学教堂山分校)
专题命中 视觉定位与Grounding :VLM(title,abstract);vision-language model(abstract)
AI总结 本文提出SEAM表示方法,通过分解中间表示为词汇和语法,提升VLM在机器人操作中的可理解和通用性,结合检索增强的少样本学习策略实现高效操作,并在动作通用性和VLM可理解性上展示出优于主流方法的性能。
不要学习,而是依托:自然语言推理与视觉依托的案例
机构 * Utrecht University(乌特勒支大学)
专题命中 视觉定位与Grounding :grounding(title,abstract);visual question answering(abstract)
AI总结 本文提出了一种基于视觉依托的零样本自然语言推理方法,通过生成视觉表示并比较与假设的相似度,实现高精度推理,展示了对文本偏见的鲁棒性。
专题命中 视觉定位与Grounding :MLLM(title);grounding(abstract);multimodal large language model(abstract)
Comments We plan to revise the methodology and update the experimental analysis before resubmission
机构 * 2 Department of Electrical ; Computer Engineering University of Waterloo, Waterloo, ON, Canada N2L 3G1 Email ; 3 College of Computer ; Information Sciences Prince Sultan University Email
专题命中 视觉定位与Grounding :multimodal large language model(title,abstract);分类 cs.CV、cs.AI、cs.LG
Journal ref 2024 IEEE International Symposium on Medical Measurements and Applications (MeMeA)
机构 * Beijing Normal University(北京师范大学) ; Tencent AI Lab(腾讯AI实验室)
专题命中 视觉定位与Grounding :grounding(title,abstract);vision-language model(abstract)
Comments Accepted to EMNLP 2025 (Findings). This version corrects a redundant sentence in the Results section that appeared in the camera-ready version
机构 * Automated Manufacturing Department(自动化制造部门) ; Al-Khwarizmi College of Engineering(阿尔·卡瓦尔米工程学院) ; University of Baghdad(巴格达大学)
专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV、cs.AI、cs.LG
Comments 30 pages, 12 figures, 4 tables
机构 * Tencent(腾讯公司) ; Department of Computer Science and Technology, Tsinghua University(清华大学计算机科学与技术系) ; Zhejiang University(浙江大学) ; Singapore University of Technology and Design(新加坡科技设计大学)
专题命中 视觉定位与Grounding :grounding(title,abstract);multimodal large language model(abstract)
Comments ACL 2025 (Oral, Industry Track)
机构 * Department of Electronic Systems, Aalborg University, Denmark(电子系统系,奥胡斯大学) ; Department of Health Science and Technology, Aalborg University, Denmark(健康科学与技术系,奥胡斯大学)
专题命中 视觉定位与Grounding :visual language model(title);vision language model(abstract);VLM(abstract)
Comments ICAT 2025
专题命中 视觉定位与Grounding :vision language model(title);vision-language model(abstract);VLM(abstract)
Comments 6 pages, 6 figures, Accepted at IEEE MILCOM 2025
机构 * Hanyang University(汉阳大学) ; University of Toronto(多伦多大学) ; University Health Network(大学健康网络) ; Knovel Engineering Lab(Knovel工程实验室) ; Michigan State University(密歇根州立大学) ; KU Leuven(鲁汶大学) ; German Research Center for Artificial Intelligence (DFKI)(德国人工智能研究中心) ; Max Planck Research School for Intelligent Systems (IMPRS-IS)(马克斯·普朗克智能系统研究学校) ; University of Stuttgart(斯图加特大学) ; UC Berkeley(伯克利大学) ; Johns Hopkins University(约翰霍普金斯大学)
专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV、cs.AI、cs.LG
Comments Preprint, 51 pages
机构 * Tencent Robotics X(腾讯机器人X) ; Harbin Institute of Technology(哈尔滨工业大学)
专题命中 视觉定位与Grounding :grounding(title,abstract);visual language model(abstract)
机构 * School of Software, Tsinghua University(清华大学软件学院) ; School of computer science and engineering, Central South University(中南大学计算机科学与工程学院) ; Inspur Yunzhou Industrial Internet Co., Ltd(Inspur Yunzhou工业互联网有限公司) ; Tsinghua University(清华大学)
专题命中 视觉定位与Grounding :grounding(title,abstract);vision-language model(abstract)
Comments ACM MM '25
机构 * ECE Department(电子工程系) ; Carnegie Mellon University(卡内基梅隆大学) ; Computer Science Department(计算机科学系) ; Columbia University(哥伦比亚大学) ; Rotman School of Management(罗特曼管理学院) ; University of Toronto(多伦多大学) ; Department of Statistics(统计学系) ; George Washington University(乔治华盛顿大学) ; Department of Computer Science(计算机科学系) ; San Francisco State University(旧金山州立大学)
专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV、cs.AI、cs.LG