arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 782 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 782 篇

2607.06445 2026-07-08 cs.CV cs.AI 新提交 85%

Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders

代理分析:作为条件编码器的视觉语言模型中的定位信号

Yoav Baron, Sara Dorfman, Roni Paiss, Daniel Cohen-Or, Or Patashnik

机构 * Tel Aviv University, Tel Aviv, Israel(特拉维夫大学) Google DeepMind(谷歌DeepMind)

专题命中 视觉定位与Grounding :VLM(summary_cn,abstract);vision-language model(abstract);分类 cs.CV、cs.AI

AI总结 研究视觉语言模型(VLMs)作条件编码器时编辑管道定位性能差的问题,引入代理分析框架,通过训练代理模型分析VLM中间表示来揭示编码定位信息的表示,发现现有条件提取策略缺陷,为条件架构设计提供新思路。

Comments Accepted as a Spotlight at the ICML 2026 Mechanistic Interpretability Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.12910 2026-06-15 cs.RO cs.AI cs.CV cs.SY eess.SY 新提交 85%

Bounding Boxes as Goals: Language-Conditioned Grasping via Neuro-Symbolic Planning

边界框作为目标:通过神经符号规划实现语言条件抓取

Allison Andreyev, Landon Eum, Nestor Tiglao, Romel Gomez

机构 * University of California, Berkeley(加州大学伯克利分校)

专题命中 视觉定位与Grounding :VLM(summary_cn,abstract);vision-language model(abstract);分类 cs.CV、cs.AI

AI总结 提出GRASP框架,利用预训练VLM将自然语言查询转化为神经符号目标状态,通过边界框检测实现零样本桌面操作,无需任务特定训练。

Comments Project website: https://allisonandreyev.github.io/grasp.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.08034 2026-06-09 cs.CV cs.AI cs.CL 新提交 85%

Sci-Rho: A Multilingual Visually-Grounded Symbolic Benchmark for STEM Problems

Sci-Rho:面向STEM问题的多语言视觉基础符号基准

Muhammad Falensi Azmi, Ikhlasul Akmal Hanif, Vallerie Alexandra Putra, Adi Yeltay, Abdullah Mubarak, Fajri Koto

机构 * Independent Researcher(独立研究员) MBZUAI(穆罕默德·本·扎耶德人工智能大学) Binus University(比努斯大学) Bandung Institute of Technology(万隆理工学院)

专题命中 视觉定位与Grounding :VLM(summary_cn,abstract);grounding(abstract);分类 cs.CV、cs.AI

AI总结 提出Sci-Rho,一个多语言、视觉基础的STEM问题动态基准,包含4242个模板和42420个实例,评估17个VLM发现最差精度与平均精度存在差距,且小模型跨语言性能下降。

Comments 22 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.07613 2026-06-09 cs.CV cs.AI 新提交 85%

Can You Trust What You See? Human and AI Detection of Synthetic Legal Evidence

你能相信你所见的吗?人类与AI对合成法律证据的检测

Jinzhe Tan, Ali Ekber Cinar, Karim Benyekhlef

机构 * Faculty of Law, McGill University(麦吉尔大学法学院)

专题命中 视觉定位与Grounding :MLLM(summary_cn,abstract);multimodal large language model(abstract);分类 cs.CV、cs.AI

AI总结 研究人类和前沿多模态大模型在民事纠纷场景中区分真实照片与AI生成图像的能力,发现两者均不可靠,提出结合人工审查、MLLM筛查和来源认证的解决方案。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.25294 2026-07-29 cs.CV cs.AI cs.CL cs.LG 新提交 85%

CLBench-V: Evaluating Multimodal Context Learning from Grounding to Knowledge Acquisition

CLBench-V:评估从基础到知识获取的多模态上下文学习

Lai Wei, Chengqi Li, Jiapeng Li, Ruina Hu, Yue Wang, Weiran Huang

机构 * School of Computer Science, Shanghai Jiao Tong University(上海交通大学计算机科学与工程学院) Zhongguancun Academy(中关村科学城创新中心) Shanghai Innovation Institute(上海创新研究院)

专题命中 视觉定位与Grounding :grounding(title,abstract);visual question answering(abstract);分类 cs.CV、cs.AI、cs.LG

AI总结 研究多模态上下文学习问题,引入CLBench-V基准,围绕上下文基础、新信息应用和新知识学习三个维度组织任务,结合公共基准与新数据集,经自动化程序构建,测试六个模型,分析相关因素,揭示多模态上下文学习远未饱和及各模型表现情况。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.17710 2026-06-23 cs.CV cs.AI cs.CL cs.LG 新提交 85%

Vision-language models for chest radiography do not always need the image

胸部X光片的视觉-语言模型并不总是需要图像

Mahshad Lotfinia, Sebastian Ziegelmayer, Lisa Adams, Daniel Truhn, Andreas Maier, Soroosh Tayebi Arasteh

机构 * Pattern Recognition Lab, Friedrich-Alexander-Universität Erlangen-Nürnberg(弗里德里希-亚历山大-埃尔朗根-纽伦堡大学模式识别实验室) Department of Diagnostic and Interventional Radiology, TUM University Clinic, School of Medicine and Health, Klinikum rechts der Isar, Technical University of Munich(慕尼黑工业大学医学院与健康学院伊萨尔河右岸医院诊断与介入放射学系) Lab for AI in Medicine, RWTH Aachen University(亚琛工业大学医学人工智能实验室) Department of Diagnostic and Interventional Radiology, University Hospital RWTH Aachen(亚琛工业大学医院诊断与介入放射学系)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);grounding(abstract);分类 cs.CV、cs.AI、cs.LG

AI总结 本文通过因果审计方法,发现许多医学视觉-语言模型在胸部X光片任务中依赖文本先验而非图像,纯文本模型与多模态模型性能接近,并提出了基于图像依赖性的评估框架。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.09147 2026-08-11 cs.CV 新提交 84%

RefineAny3D: Depth Refinement as Semantic Alignment for Monocular 3D Detection

RefineAny3D:作为语义对齐的深度细化用于单目3D检测

Zhihao Zhang, Gengwei Zhang, Tianlong Chen, Xiaoming Liu

机构 * Michigan State University(密歇根州立大学) University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校)

专题命中 视觉定位与Grounding :VLM(summary_cn,abstract);vision-language model(abstract);分类 cs.CV

AI总结 RefineAny3D将单目3D检测的深度细化转化为视觉对齐问题,通过VLM实现无需数值预测的深度修正,在多类检测工具上均有性能提升且可泛化。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.01113 2026-08-04 cs.CV 新提交 84%

CoT-Edit: Let CoT Guide Instruction Video Editing

CoT-Edit:让思维链(CoT)指导指令视频编辑

Sen Liang, Fengbin Guan, Youliang Zhang, Xin Li, Zhibo Chen

机构 * University of Science and Technology of China(中国科学技术大学) Zhongguancun Academy(中关村学院) Tsinghua University(清华大学)

专题命中 视觉定位与Grounding :MLLM(summary_cn,abstract);multimodal large language model(abstract);分类 cs.CV

AI总结 本文提出CoT-Edit的plan--guide--edit框架,以CoT增强的MLLM为规划器生成空间先验,结合扩散编辑器实现高保真指令视频编辑,性能优于多个基准方法

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.00232 2026-08-04 cs.CV 新提交 84%

Real-Time Visual Obstruction Detection in Surgical Augmented Reality

手术增强现实中的实时视觉遮挡检测

Shih-Chin Yang, Yanming Xiu, Hanting Ye, Qi Chen, Elias Rotondo, Maria Gorlatova

机构 * Duke University(杜克大学)

专题命中 视觉定位与Grounding :VLM(summary_cn,abstract);vision-language model(abstract);分类 cs.CV

AI总结 针对手术AR虚拟内容遮挡手术器械的问题,提出结合VLM与分割推理的延迟感知流水线,构建伪AR基准,实现87.43%准确率、479 ms延迟,较云端基线降延迟62.90%。

Comments ISMAR 2026 Mecidal Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.25763 2026-06-25 cs.CV 新提交 84%

ShutterMuse: Capture-Time Photography Guidance with MLLMs

ShutterMuse: 基于多模态大语言模型的拍摄时刻摄影指导

Jiayu Li, Yixiao Fang, Tianyu Hu, Wei Cheng, Ping Huang, Zheheng Fan, Gang Yu, Xingjun Ma

机构 * Fudan University(复旦大学) StepFun

专题命中 视觉定位与Grounding :MLLM(summary_cn,abstract);multimodal large language model(abstract);分类 cs.CV

AI总结 提出CaptureGuide-Bench基准测试,评估MLLM在摄影师构图与主体姿态推荐方面的能力,并构建ShutterMuse统一模型,通过监督与强化微调实现最佳综合性能。

Comments Project Page:https://lijayutnt.github.io/ShutterMuse

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.18609 2026-06-18 cs.CV 新提交 84%

Hallucination Detection and Correction in Medical VLMs via Counter-Evidence Verification

基于反事实证据验证的医学视觉语言模型幻觉检测与纠正

Nan Zhou, Ke Zou, Meng Liu, Linchao He, Jiaqi Zhu, Yi Zhang, Hu Chen, Huazhu Fu

机构 * College of Computer Science, Sichuan University(四川大学计算机科学学院) Yong Loo Lin School of Medicine, National University of Singapore(新加坡国立大学杨潞龄医学院) Key Laboratory of Data Protection and Intelligent Management, Ministry of Education, Sichuan University(四川大学数据保护与智能管理教育部重点实验室) National Key Laboratory of Autonomous Intelligent Unmanned Systems, Beijing Institute of Technology(北京理工大学自主智能无人系统国家重点实验室) Institute of High Performance Computing (IHPC), Agency for Science, Technology and Research (A*STAR)(新加坡科技研究局高性能计算研究所)

专题命中 视觉定位与Grounding :VLM(summary_cn,abstract_cn);vision-language model(abstract);grounding(abstract);分类 cs.CV

AI总结 提出CoEV框架,通过文本与视觉证据的双向验证检测并纠正医学VLM幻觉,无需重新训练,在四个数据集上显著提升检测和纠正性能。

Comments MICCAI 2026 Accept. Submission Version

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.17433 2026-06-17 cs.CV 新提交 84%

LADBench: A Benchmark for Logical Fault Detection in Images

LADBench: 图像中逻辑故障检测的基准

Sahasra Kondapalli, Lara Radovanovic, Aadi Palnitkar, Mingyang Mao, Xiaomin Lin

机构 * University of South Florida(南佛罗里达大学)

专题命中 视觉定位与Grounding :VLM(summary_cn);vision language model(abstract);visual question answering(abstract);grounding(abstract)

AI总结 提出LAD-Bench基准,包含1000多张合成图像的四域逻辑异常,通过分层提示协议评估模型,揭示现有VLM在隐式逻辑故障检测上的不足。

Comments Accepted to the IEEE International Conference on Development and Learning (ICDL 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.08009 2026-08-11 cs.CV cs.AI 新提交 84%

Evidence-Grounded Forensic Reasoning for Detecting and Grounding Multi-Modal Media Manipulation

基于证据的取证推理:检测与定位多模态媒体篡改

Yichun Yeh, Yiheng Li, Xiaobo Hu, Zhen Lei, Yang Yang

专题命中 视觉定位与Grounding :grounding(title,abstract);MLLM(abstract_cn);分类 cs.CV、cs.AI

AI总结 本文针对多模态媒体篡改检测与定位问题,提出基于证据的取证推理框架,结合锚定-验证推理链、可验证奖励系统与模态解耦优势路由机制,实现最优性能与可解释性的统一。

Comments accepted by ACM MM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.00473 2026-08-04 cs.CV cs.AI cs.CL 新提交 84%

CrossProjection: Geometric Grounding Beyond Viewpoint Change in Architectural Drawings

CrossProjection:建筑图纸中超越视角变化的几何定位

Kaho Li, Pengyu Zeng, Yuqin Dai, Jun Yin, Tianjing Feng, Shuai Lu

机构 * Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院) University College London(伦敦大学学院)

专题命中 视觉定位与Grounding :grounding(title,abstract);vision-language model(abstract);分类 cs.CV、cs.AI

AI总结 该研究提出CrossProjection方法评估视觉语言模型在异构建筑视图间的几何定位能力,发现模型在自由几何定位上性能脆弱,仅封闭选择成功不代表具备可靠几何定位能力。

Comments Initial controlled diagnostic study on 23 natural drawing sets and three VLMs; broader model, building, repeated-inference, and human coverage is planned for a subsequent version

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.27558 2026-07-31 cs.CV cs.AI 新提交 84%

Drawing-Recode: Annotation Grounding for Parametric CAD Code Generation from Raster 2D CAD Drawings

Drawing-Recode:基于栅格2D CAD图纸的参数化CAD代码生成的标注定位方法

Mingi Kim, Yongjun Kim, Hyungki Kim

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV、cs.AI

AI总结 Drawing-Recode是一种从栅格2D CAD图纸生成参数化CAD代码的框架,通过图像编码器、文本识别模块、交叉注意力与AGL实现标注定位,性能优于基线且对工业扫描图纸鲁棒,可助力制造自动化。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.24570 2026-07-28 cs.CV cs.AI 新提交 84%

The Visual Bottleneck: Sparse-Frame Adaptation of MLLMs for Joint Spatial-Temporal Video Grounding

视觉瓶颈:用于联合时空视频定位的多模态大语言模型的稀疏帧适配

Jiameng Zhang, Srikanth Madikeri

机构 * University of Zurich(苏黎世大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);multimodal large language model(abstract);分类 cs.CV、cs.AI

AI总结 研究针对多模态大语言模型在视频定位中训练与部署条件不匹配问题,通过实验表明视觉特征提取是稀疏帧输入瓶颈,提出适配特定层及边界感知采样策略,证明训练策略对稀疏帧视频定位比模型规模更关键。

Comments 15 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.19857 2026-07-23 cs.CV cs.AI 新提交 84%

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos

用于流式航空视频中小目标理解的内存增强多模态大语言模型

Penglei Sun, Yehua Huang, Zhuoli Tao, Xiang Li, Runwei Guan, Yaoxian Song, Kaiyong Zhao, Henghui Ding, Bo Han, Yang Yang, Xiaowen Chu

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) University of Freiburg(弗莱堡大学) Hangzhou City University(杭州城市大学) XGRIDS(XGRIDS公司) Fudan University(复旦大学) Hong Kong Baptist University(香港浸会大学)

专题命中 视觉定位与Grounding :multimodal large language model(title,abstract);MLLM(abstract);分类 cs.CV、cs.AI

AI总结 研究针对流式航空视频中小目标理解难题,提出像素级开放词汇数据集DroneEyes,以及含语义感知令牌路由器和分层内存库的多模态大语言模型SkyAnchor,从数据和方法角度应对挑战。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.15942 2026-07-20 cs.CV cs.LG 新提交 84%

More with Less: a Large Scale Remote Sensing VLM with a Simple Recipe

以少胜多:一种简单方法构建的大规模遥感视觉语言模型

Stefan Maria Ailuro, Mario Markov, Mohammad Mahdi, Luc Van Gool, Danda Pani Paudel

机构 * INSAIT, Sofia University “St. Kliment Ohridski”(索非亚大学“圣克莱门特·奥赫里德斯基”信息与自动化研究所)

专题命中 视觉定位与Grounding :VLM(title);vision-language model(abstract);grounding(abstract);分类 cs.CV、cs.LG

AI总结 研究针对遥感视觉语言模型,质疑架构专业化必要性,提出通用模型经大规模跨数据和任务训练可获好性能。核心方法是用单一语言策略及多任务强化学习框架训练。主要贡献是在多基准测试中取得竞争力结果,证明数据规模更关键。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.25491 2026-06-25 cs.CV cs.AI 新提交 84%

HG-Bench: A Benchmark for Multi-Page Handwritten Answer-Region Grounding in Automated Homework Assessment

HG-Bench: 自动作业批改中多页手写答案区域定位的基准

Chuangxin Zhao, Boyan Shi, Yanling Wang, Yijian LU, Canran Xiao, Jiali Chen, Jun Xia, Yan Wang, Ji Qi, Juanzi Li

专题命中 视觉定位与Grounding :grounding(title,abstract);VLM(abstract_cn);分类 cs.CV、cs.AI

AI总结 针对自动作业批改中多页手写答案区域定位的缺失评估,提出HG-Bench基准,包含500个带层级标注的样本,并设计页面感知评估协议,揭示现有模型在步骤级定位上的能力差距。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.24759 2026-06-24 cs.CV cs.AI 新提交 84%

UniDrive: A Unified Vision-Language and Grounding Framework for Interpretable Risk Understanding in Autonomous Driving

UniDrive: 面向自动驾驶可解释风险理解的统一视觉-语言与定位框架

Xiaowei Gao, Pengxiang Li, Yitai Cheng, Ruihan Xu, James Haworth, Stephen Law, Yun Ye

机构 * organization= Department of Earth Science \& Engineering, Imperial College London , city= London , postcode= SW7 2AZ , country= United Kingdom organization= SpaceTimeLab, Department of Civil, Environmental Geomatic Engineering, University College London , city= London , postcode= WC1E 6BT , country= United Kingdom organization= Department of Computing, The Hong Kong Polytechnic University , city= Hong Kong , country= China organization= Trinity College, University of Oxford , city= Oxford , postcode= OX1 3BH , country= United Kingdom organization= Department of Geography, University College London , city= London , postcode= WC1E 6BT , country= United Kingdom organization= Centre for Global Infrastructure Resilience, The Bartlett School of Sustainable Construction, University College London , city= London , postcode= WC1E 7HB , country= United Kingdom

专题命中 视觉定位与Grounding :grounding(title,abstract);multimodal large language model(abstract);分类 cs.CV、cs.AI

AI总结 提出UniDrive框架,通过融合时序推理与高分辨率感知分支,联合生成风险描述和边界框定位,在DRAMA-Reasoning基准上超越现有方法,提升小目标定位和可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.11889 2026-06-11 cs.CV cs.AI cs.RO 新提交 84%

Task-Aligned Stability Analysis of Vision-Language Models for Autonomous Driving Hazard Detection

面向自动驾驶危险检测的视觉-语言模型任务对齐稳定性分析

Everett Richards

机构 * Everett Richards(埃弗里特·里奇ards)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);VLM(abstract_cn);分类 cs.CV、cs.AI

AI总结 研究视觉-语言模型在自动驾驶危险检测中,嵌入漂移与任务对齐危险分数变化的关系,发现不同腐败类型导致不同的失效模式,建议基准测试包含任务对齐稳定性指标。

Comments 8 pages (5 main body + 3 references / appendices). ICML 2026 Workshop on Combining Theory and Benchmarks (CTB)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.19752 2026-08-21 cs.SE 新提交 83%

A Fully Automated, Deployment-Aware Testing Pipeline for IoT-Based Automotive Applications

面向物联网汽车应用的全自动、感知部署的测试流水线

Denesa Zyberaj, Roman Vintonyak, Pascal Hirmer, Marco Aiello

专题命中 视觉定位与Grounding :VLM(summary_cn,abstract);vision-language model(abstract)

AI总结 针对物联网汽车应用的测试难题,提出结合LLM、VLM及人在回路机制的感知部署测试流水线,经CPDS案例验证可实现全需求覆盖与高准确率,适配OEM-供应商测试工作流。

Comments Submitted, accepted and presented at IoTBDS26

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.14724 2026-08-18 cs.CV cs.AI cs.LG 新提交 83%

Privacy-Preserving Dataset Curation for Kuala Lumpur Urban Traffic: Grounded Vision-Language Detection with Spatial Vehicle-Context Filtering

面向吉隆坡城市交通的隐私保护数据集构建:结合空间车辆上下文过滤的接地视觉-语言检测

Mohammed Abdul Al Arafat Tanzin, Rudzidatul Akmam Dziyauddin

机构 * Faculty of Artificial Intelligence, Universiti Teknologi Malaysia(马来西亚理工大学人工智能学院)

专题命中 视觉定位与Grounding :grounding(summary_cn,abstract);分类 cs.CV、cs.AI、cs.LG

AI总结 针对吉隆坡热带城市交通场景的隐私保护数据集构建难题,提出结合Grounding DINO与空间车辆ROI包含引擎的自动化匿名化框架,在1266帧图像上实现约95%的匿名化成功率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.20127 2026-08-21 cs.CV 新提交 83%

ID-VTG: Image-Disambiguated Video Temporal Grounding

ID-VTG:基于图像消歧的视频时间定位

Minghang Zheng, Jingli Wei, Hongyi Yang, Yang Liu

机构 * Wangxuan Institute of Computer Technology, Peking University(北京大学王选计算机研究所) State Key Laboratory of General Artificial Intelligence, Peking University(北京大学通用人工智能国家重点实验室)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

AI总结 针对视频时间定位中视觉相似实体的消歧难题,本文提出ID-VTG任务与VGD-Agg框架,构建两个基准数据集,实现了该任务的当前最优性能。

Comments ACM-MM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.14252 2026-08-17 cs.AI cs.CL 新提交 83%

Grounding Without Corrective Control: Truth-Tracking Profiles for Large Language Models

无校正控制的基础:大语言模型的真值追踪剖面

Brett Reynolds

机构 * Humber Polytechnic(亨伯理工学院) University of Toronto(多伦多大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

AI总结 本文研究大语言模型中无校正控制的基础问题,提出路径剖面概念以分析真值追踪,指出纯文本模型继承的模式可提供衍生可应答性,不同方法对任务的真值追踪改进可能与表面改进不一致。

Comments 24 pages, 1 figure, 1 table. A six-page methodological supplement, reproducible R script, and constructed data are included as ancillary files

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.08663 2026-08-11 cs.HC cs.CV 新提交 83%

A Dynamic-Semantics Framework for Grounding Human Referring Expressions in Visual Perceptual Data

用于在视觉感知数据中建立人类指代表达式基础的动态语义框架

Joseph Bingham

专题命中 视觉定位与Grounding :grounding(title,abstract);vision-language model(abstract);分类 cs.CV

AI总结 本文提出动态语义框架,通过显式约定状态绑定集合与感知对齐流程,解决视觉-语言模型在词汇同步任务中的不足,在斯坦福重复参考游戏语料库上达到83.56%的top-5准确率,核心是透明符号层与可检查感知通道的结合。

Comments 23 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.04510 2026-08-06 cs.RO cs.AI 新提交 83%

GUARD: Grounding Uncertainty and Ablation-Based Risk Detection for Diffusion-Based VLAs

GUARD:面向扩散型视觉-语言动作(VLA)的不确定性与基于消融的风险检测

Suhas Hegde, Jitendra Yasaswi Bharadwaj Katta

专题命中 视觉定位与Grounding :grounding(title,abstract);vision-language model(abstract);分类 cs.AI

AI总结 本文提出GUARD方法,无需修改预训练扩散型VLA策略即可检测其故障,在多基准测试中提升未见过任务的ROC-AUC,提供可跨多维度迁移的故障信号。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.00775 2026-08-04 cs.RO cs.CV cs.HC 新提交 83%

ORCESTRA: VLM-driven Visual Robot programming in Mixed Reality

ORCESTRA:基于视觉语言模型的混合现实视觉机器人编程系统

Ivan Snegirev, Elizaveta Semenyakina, Mikhail Konenkov, Artem Lykov, Miguel Altamirano Cabrera, Dzmitry Tsetserukou

机构 * Skolkovo Institute of Science and Technology(斯科尔科沃科技学院) R&D Center, MWS(MWS研发中心)

专题命中 视觉定位与Grounding :VLM(title);vision-language model(abstract);grounding(abstract);分类 cs.CV

AI总结 ORCESTRA是一款混合现实系统,通过无代码路径点示教和语言引导控制实现机器人数字孪生编程,支持多种异构机器人,其混合现实验证可作为语言引导机器人物理部署前的安全层。

Comments 4 page, 3 figures, 1 table, ISMAR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.28464 2026-07-31 cs.CV 新提交 83%

Can Vision-Language Models Reason about AI Edits in Images?

视觉-语言模型能否对图像中的AI编辑进行推理?

Darsha Udayanga, Pin-Yu Chen, Payel Das, Qiang Ji

机构 * Rensselaer Polytechnic Institute(伦斯勒理工学院) IBM Research(IBM研究院)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV

AI总结 本研究探究能否用强化学习而非显式推理监督训练视觉-语言模型推理AI图像编辑,提出基于GRPO的框架,引入eff-IoU指标,在多数据集上实现与SOTA相当的检测定位性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.24810 2026-07-29 cs.AI eess.IV 新提交 83%

RRS-10K: A Multitask Vision-Language Model Benchmark for Rare Remote Sensing Image Interpretation

RRS-10K:用于罕见遥感图像解释的多任务视觉语言模型基准测试

Yuqiao Lai, Jiancheng Qi, Fei Wang, Yuxin Liu, Kun Li, Ye Chen, Yan Gao, Yanyan Wei

机构 * Laboratory of Intelligent Language Processing, National University of Defense Technology(国防科技大学智能语言处理实验室) Hefei University of Technology(合肥工业大学) Institute of Artificial Intelligence, Hefei Comprehensive National Science Center(合肥综合性国家科学中心人工智能研究院) United Arab Emirates University(阿联酋大学)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);grounding(abstract);分类 cs.AI

AI总结 针对视觉语言模型在罕见遥感图像解释能力不足的问题,提出RRS-10K基准测试,含军事遥感图像及问答对,构建时引入干扰项过滤策略,评估多个模型,揭示其性能弱点,为开发更可靠的遥感视觉语言模型提供指导。

详情

展开后加载摘要…

URL PDF HTML 收藏