arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 26403 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7473 篇

2607.27902 2026-08-04 cs.CV 版本更新 91%

One Patch Is Enough: Reinforcement-Optimized Visual Token Grounding for MLLM-Based Scene Text Spotting

一个补丁就够:面向基于多模态大语言模型(MLLM)的场景文本定位的强化优化视觉标记接地

Rui Tang, Wentao Yang, Peirong Zhang, Yongxin Shi, Shun Zhang, Huiguo He, Lianwen Jin

机构 * South China University of Technology(华南理工大学) HiThink Research(海思思考研究院)

专题命中 视觉定位与Grounding :MLLM(title,title_cn);grounding(title,abstract);multimodal large language model(abstract);分类 cs.CV

AI总结 针对现有多补丁场景文本定位范式的冗余噪声与定位歧义问题,提出以视觉为中心的单补丁文本定位框架 SPaTS,通过强化学习优化的单补丁选择等技术实现性能提升,显著优于前沿相关模型。

Comments 15 pages, 11 figures. Accepted to ACM Multimedia 2026

Journal ref Proceedings of the 34th ACM International Conference on Multimedia (MM '26), 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.26250 2026-04-30 cs.CV 91%

Beyond Shortcuts: Mitigating Visual Illusions in Frozen VLMs via Qualitative Reasoning

超越捷径:通过定性推理缓解冻结VLM中的视觉错觉

Hao Guo, Fei Wang, Junjie Chen, Yiqi Nie, Jiaqi Zhao, Qiankun Li, Subin Huang

机构 * Institute of Artificial Intelligence, Hefei Comprehensive National Science Center(人工智能研究院,合肥国家科学中心) Anhui Polytechnic University(安徽理工大学) Hefei University of Technology(合肥工业大学) Anhui University(安徽大学) IGS, Imperial College London(帝国理工学院伦敦分校)

专题命中 视觉定位与Grounding :VLM(title_cn,summary_cn);grounding(summary_cn,abstract);vision-language model(abstract);分类 cs.CV

AI总结 本文提出SQI框架,通过定性约束提升冻结VLM的视觉 grounding,解决视觉错觉问题,实验显示在DataCV 2026挑战中表现优异,提升准确率并提供更好的可解释性。

Comments 4 pages, 2 figures, and 1 table. This is a methodology paper for the DataCV 2026 Challenge (CVPR Workshops), Task 1, where our method ranked 2nd

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.24133 2026-08-26 cs.CV cs.AI cs.IR 新提交 91%

PlaceSeek: Human-Centered Geospatial Retrieval of Urban Outdoor Places via Semantic Grounding and Affective Alignment

PlaceSeek:通过语义 grounding 与情感对齐实现以人类为中心的城市户外地点地理空间检索

Ziqi Cui, Shangyu Lou

机构 * University of British Columbia(不列颠哥伦比亚大学) University of California, Santa Barbara(加利福尼亚大学圣巴巴拉分校) San Diego State University(圣地亚哥州立大学)

专题命中 视觉定位与Grounding :grounding(title,title_cn);vision-language model(abstract);分类 cs.CV、cs.AI

AI总结 该研究针对现有地理空间检索难以满足情感/活动类需求的问题,提出 PlaceSeek 框架,通过语义 grounding 与情感对齐实现城市户外地点检索,在米兰数据集上的表现优于多种基线方法。

Comments Accepted as a Research Paper (short) at ACM SIGSPATIAL 2026. This arXiv version is the full version of the paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.15410 2026-08-18 cs.DC cs.AI cs.CV cs.RO cs.SY eess.SY 新提交 91%

FloodReasonBench: Benchmarking VLM Reasoning Segmentation for Embodied Flood Response at the Edge

FloodReasonBench:面向边缘端具身洪水响应的VLM推理分割基准

Rajat Bhattacharjya, Yoomee Jung, Minwoo Kim, Sing-Yao Wu, Eli Bozorgzadeh, Nalini Venkatasubramanian, Nikil Dutt

专题命中 视觉定位与Grounding :VLM(title,title_cn);grounding(abstract,abstract_cn);vision-language model(abstract);分类 cs.CV、cs.AI

AI总结 FloodReasonBench针对边缘端具身洪水响应,构建专用数据集并刻画推理分割流水线的性能权衡,为资源受限场景提供任务与系统层面的基准支持。

Comments Paper is currently under review. The code and dataset will be made public upon acceptance

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.07886 2026-08-11 cs.CV cs.AI cs.CL 新提交 91%

Vision-Language Grounding as Bidirectional Concept Correspondence

视觉-语言 Grounding 作为双向概念对应

Jieyu Zhang, Ziqi Gao, Luke Zettlemoyer, Ranjay Krishna

机构 * University of Washington(华盛顿大学) Allen Institute for AI(艾伦人工智能研究所) FAIR at Meta(Meta FAIR实验室)

专题命中 视觉定位与Grounding :grounding(title,title_cn);vision-language model(abstract);分类 cs.CV、cs.AI

AI总结 本研究将视觉-语言 Grounding 建模为双向概念对应,提出 ConCor-1 模型统一相关任务,在长文本数据集和零样本 LVIS 上对应 F1 分别提升 48%、29%,性能优于基线。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.06154 2026-08-07 cs.RO cs.AI cs.CV 新提交 91%

Visual Grounding in Zero-Shot Vision-Language Control

零样本视觉语言控制中的视觉 grounding

J. de Curtò, Dayani Plasencia, Diego Sánchez, I. de Zarzà

专题命中 视觉定位与Grounding :grounding(title,title_cn);VLM(abstract,abstract_cn);vision-language model(abstract);分类 cs.CV、cs.AI

AI总结 本文针对零样本视觉语言控制中VLMs决策是否基于视觉输入的问题,通过多组消融实验分析了多种VLMs的表现,提出对称共识守护者方法,验证了VLMs可作为有界的危险助手。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.06891 2026-06-08 cs.CV 新提交 91%

Stream3D-VLM: Online 3D Spatial Understanding with Incremental Geometry Priors

Stream3D-VLM:基于增量几何先验的在线3D空间理解

Hanxun Yu, Xuan Qu, Lei Ke, Boqiang Zhang, Yuxin Wang, Jianke Zhu, Dong Yu

机构 * Zhejiang University(浙江大学) Tencent Hunyuan(腾讯文汇) HKUST(香港科技大学) Shenzhen Loop Area Institute(深圳河套学院)

专题命中 视觉定位与Grounding :VLM(title,title_cn);vision-language model(abstract);grounding(abstract);分类 cs.CV

AI总结 提出在线3D视觉语言模型Stream3D-VLM,通过自回归流控制、轻量视觉-空间特征融合模块和几何自适应体素压缩,实现从流式视频中实时理解3D空间,并构建超百万在线3D问答数据集,在多项任务上超越现有模型。

Comments Project Page: https://stream3d-vlm.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.17253 2026-07-15 cs.AR 版本更新 91%

PDAGENT-BENCH: Characterizing, Grounding, and Architecting LLM/VLM Agents for VLSI Physical Design

PDAGENT-BENCH: 用于VLSI物理设计的LLM代理的特征化、基础化与架构化

Qiufeng Li, Rongqian Chen, Quan Cheng, Chengxuan Wang, Sizhe Tang, Chia-Tung Ho, David Z. Pan, Tian Lan, Weidong Cao

专题命中 视觉定位与Grounding :VLM(title,summary_cn);grounding(title);vision-language model(abstract)

AI总结 提出PDAGENT-BENCH基准,用于评估LLM/VLM代理在VLSI物理设计中的能力,涵盖任务级和工作流级评估,揭示模型在工具执行和长程推理上的局限,并验证人类技能增强工作流的有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08066 2025-08-21 cs.CV cs.AI cs.CL cs.LG 91%

ExpVG: Investigating the Design Space of Visual Grounding in Multimodal Large Language Model

Weitai Kang, Weiming Zhuang, Zhizhong Li, Yan Yan, Lingjuan Lyu

机构 * University of Illinois Chicago(伊利诺伊大学香槟分校) Sony AI(索尼人工智能)

专题命中 视觉定位与Grounding :grounding(title,abstract);multimodal large language model(title,abstract);LLaVA(abstract);MLLM(abstract)

Comments 8 pages for the main paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.21244 2026-08-24 cs.CV 新提交 91%

A VLM Answer Is Not an Anomaly Score: Rank Compression in Training-Free Video Anomaly Detection

VLM 的回答并非异常分数:无训练视频异常检测中的排名压缩

Inpyo Song, Jangwon Lee

机构 * SungKyunKwan University(成均馆大学)

专题命中 视觉定位与Grounding :VLM(title,title_cn);vision-language model(abstract);分类 cs.CV

AI总结 该研究针对基于VLM的无训练视频异常检测,发现生成式回答读出存在排名压缩问题,提出概率读出方法可提升性能,证明回答接口是该类检测器的关键组成部分。

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.20069 2026-08-21 cs.CV 新提交 91%

V-REX: Efficient Specialist VLM Training for Veterinary X-Rays

V-REX:面向兽医X光片的高效专用视觉语言模型(VLM)训练

Tim Elsner, Nicole McNally, Andre Dourson, Michael Fitzke

机构 * Vyyo AI Mars Petcare(玛氏宠物护理)

专题命中 视觉定位与Grounding :VLM(title,title_cn);grounding(abstract);分类 cs.CV

AI总结 该研究针对兽医X光片领域,提出高效专用VLM训练方案,无需依赖额外数据,用更少资源开发出能生成兽医X光诊断报告的模型,性能远超同类开源基础模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.18577 2026-08-12 cs.CV 版本更新 91%

Attention Without Grounding: Causal Evaluation of Visual Explanations in Medical VLMs

无基础的注意力:医学视觉语言模型中视觉解释的因果评估

Binesh Sadanandan, Vahid Behzadan

机构 * SAIL Lab, University of New Haven(萨伊勒实验室,纽黑文大学)

专题命中 视觉定位与Grounding :VLM(summary_cn,abstract);grounding(title);LLaVA(abstract,abstract_cn);vision-language model(abstract)

AI总结 研究医学视觉语言模型中视觉解释的因果评估,通过多种方式审核注意力热图忠实度,发现没有评估的VLM满足忠实标准,热图虽视觉上安心但不忠实,临床解释需可控指标和因果扰动而非仅视觉检查。

Comments iMIMIC Workshop 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.02791 2026-08-05 cs.CV 新提交 91%

Better, Stronger, Faster, and Broader: Structured All-Mask Prediction for MLLM-Based Segmentation

更好、更强、更快、更广泛:基于多模态大型语言模型(MLLM)的结构化全掩码预测分割

Jiazhen Liu, Mingkuan Feng, Long Chen

机构 * The Hong Kong University of Science and Technology (HKUST)(香港科技大学)

专题命中 视觉定位与Grounding :MLLM(title,title_cn);grounding(abstract);分类 cs.CV

AI总结 STAMPlus通过结构化全掩码预测解耦自回归对话与非自回归掩码预测,解决了MLLM分割的三难问题,在提升性能的同时降低延迟,实现多类开放词汇等分割任务的SOTA表现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.01638 2026-08-04 cs.CV 新提交 91%

Dynamic Resolution Routing for Efficient Egocentric Grounding

面向高效自我中心 grounding 的动态分辨率路由

Huixin Sun, Wangbo Zhao, Fanyue Wei, Qiuxia Lin, Pengzhan Sun, Angela Yao

机构 * National University of Singapore(新加坡国立大学) The Hong Kong University of Science and Technology(香港科技大学) Nanyang Technological University(南洋理工大学)

专题命中 视觉定位与Grounding :grounding(title,title_cn);multimodal large language model(abstract);分类 cs.CV

AI总结 该研究针对自我中心视觉 grounding 中视觉 token 处理成本过高的问题,提出 SmartRes 框架,通过动态分辨率路由减少视觉 token,在 Ego4D 等数据集上实现 token 缩减与性能提升,代码将公开。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.11025 2026-08-04 cs.CV 版本更新 91%

Test-time Scaling over Perception: Resolving the Grounding Paradox in Thinking with Images

图像推理中的测试时感知扩展:解决图像推理中的 grounding 困境

Zheng Jiang, Yiming Chen, Nan He, Jiahui Chen, Chaoyang Li, Houde Qian, Lifeng Sun

机构 * Tsinghua University(清华大学) Beijing University of Technology(北京工业大学)

专题命中 视觉定位与Grounding :grounding(title,title_cn);multimodal large language model(abstract);分类 cs.CV

AI总结 本文提出 TTSP 框架,通过测试时扩展感知解决图像推理中的 grounding 困境,通过生成多个感知轨迹并过滤不可靠轨迹来提升细粒度视觉推理能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.27823 2026-07-31 cs.CV 新提交 91%

Hallucinations Leave a Grounding Signature:Verifier-Guided Decoding for Selective Object Correction

幻觉留下 grounding 特征:用于选择性对象修正的验证器引导解码

Lei Yang, Xinze Liu, Dayan Wu, Ding Wang, Hengjie Zhu, Zihao Zhang, Tianzhu Hu, Hanqi Wu, Peng Fu, Zheng Lin

专题命中 视觉定位与Grounding :grounding(title,title_cn);vision-language model(abstract);分类 cs.CV

AI总结 该研究针对大型视觉语言模型的对象幻觉问题,提出基于内在 grounding 特征(IGS)的验证器引导解码(VGD)框架,在保留视觉理解和对象覆盖率的同时,实现了最先进的对象幻觉缓解效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.26913 2026-07-30 cs.CV 新提交 91%

Prior Directions: Why GUI Grounding Gets Locked in the Past

先验方向:为什么GUI grounding会被锁定在过去

Weile Gong, Zijian Lu, Mingcai Chen, Yiping Zuo, Xin He, Weibei Fan

机构 * Nanjing University of Posts and Telecommunications(南京邮电大学)

专题命中 视觉定位与Grounding :grounding(title,title_cn);vision-language model(abstract);分类 cs.CV

AI总结 该研究探究视觉-语言模型的GUI grounding锁定问题,提出先验方向概念,发现其是锁定的关键,移除该方向分量可恢复视觉grounding。

Comments 13 pages, 8 figures. Code: https://github.com/phare111/prior-directions

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.11957 2026-07-27 cs.CV 版本更新 91%

Anomalous Frame Detection Using VLM-Based Description Comparison for Extracting Expert-Specific Actions and Contextual Decision-Making Scenes with Intra-Video Self-Similarity

基于VLM描述比较的异常帧检测,用于提取特定专家操作和具有视频内自相似性的上下文决策场景

Ryo Sakai, Kaname Yokoyama

专题命中 视觉定位与Grounding :VLM(title,title_cn);vision-language model(abstract);分类 cs.CV

AI总结 本文针对关键基础设施维护中专家知识传承问题,提出基于VLM描述比较检测异常帧的方法,可提取特定专家操作及上下文决策场景,在模拟实验中该方法的提取率高于传统方法,有效发现含专家知识的候选场景。

Comments 17 pages, 11 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17455 2026-06-29 cs.CV cs.RO 版本更新 91%

VLM-Guided Visual Place Recognition for Planet-Scale Geo-Localization

VLM引导的视觉地点识别用于行星级地理定位

Sania Waheed, Na Min An, Michael Milford, Sarvapali D. Ramchurn, Shoaib Ehsan

机构 * Univ. of Southampton, UK(英国南安普顿大学) KAIST, South Korea(韩国科学技术院) QUT, Australia(昆士兰科技大学) Univ. of Essex, UK(英国埃塞克斯大学)

专题命中 视觉定位与Grounding :VLM(title,title_cn);vision-language model(abstract);分类 cs.CV

AI总结 提出VLM与检索式VPR结合的混合框架,利用VLM生成先验约束搜索空间,再通过检索和重排序实现行星级地理定位,在街道和城市级别分别提升4.51%和13.52%。

Journal ref Proceedings of the Australasian Conference on Robotics and Automation (ACRA 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.22537 2026-06-26 cs.CV 新提交 91%

NegAS: Negative Label Guided Attention and Scoring for Out-of-Distribution Object Detection with Vision-Language Models

NegAS: 基于负标签引导的注意力与评分用于视觉语言模型的分布外目标检测

Yingjie Zhang, Shuai Li, Peng Wang

机构 * Northwestern Polytechnical University(西北工业大学)

专题命中 视觉定位与Grounding :VLM(summary_cn,abstract);vision-language model(title,abstract);grounding(abstract,abstract_cn);分类 cs.CV

AI总结 提出NegAS框架,利用负标签引导注意力(NegA)和基于sigmoid的OOD评分函数(NegS),解决VLM检测器中OOD区域利用不足和评分不兼容问题,显著提升OOD检测性能。

Comments Accept to ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.14803 2026-06-16 cs.CV 新提交 91%

HSQ-VLM: A Novel Spatially-Constrained Quadrant Segmentation VLM Model for Explainability in Diabetic Retinopathy

HSQ-VLM: 一种用于糖尿病视网膜病变可解释性的新型空间约束象限分割VLM模型

Shivum Telang

机构 * Pittsburgh, Pennsylvania(宾夕法尼亚州匹兹堡)

专题命中 视觉定位与Grounding :VLM(title,title_cn);vision-language model(abstract);分类 cs.CV

AI总结 提出HSQ-VLM,利用地标锚定笛卡尔交叉注意力机制和四象限拓扑潜在分割,实现眼底图像中病变的解剖精确量化与自然语言报告生成,在出血和微动脉瘤检测上达到99.6%和96.4%的灵敏度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.22558 2026-05-22 cs.CV 91%

GeoWeaver: Grounding Visual Tokens with Geometric Evidence before Scene Reasoning

GeoWeaver: 在场景推理前通过几何证据 grounding 视觉 token

Deshui Miao, Xingsen Huang, Yameng Gu, Xin Li, Haijun Zhang, Ming-Hsuan Yang

机构 * Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳)) Pengcheng Laboratory(鹏城实验室) Pazhou Lab (Huangpu)(琶洲实验室(黄埔)) Hainan University(海南大学) University of California at Merced(加州大学默塞德分校)

专题命中 视觉定位与Grounding :grounding(title,title_cn);vision-language model(abstract);分类 cs.CV

AI总结 本文提出 GeoWeaver,一种在场景推理前通过几何证据对视觉 token 进行 grounding 的框架,以提升空间推理能力并保持多模态能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06828 2026-08-03 cs.CV cs.AI 版本更新 91%

Step-Level Visual Grounding Faithfulness Predicts Out-of-Distribution Generalization in Long-Horizon Vision-Language Models

步级视觉 grounding 信念预测长视界视觉语言模型的分布外泛化

Md Ashikur Rahman, Md Arifur Rahman, Niamul Hassan Samin, Abdullah Ibne Hanif Arean, Juena Ahmed Noshin

机构 * The KOW Company(KOW公司) University of Dhaka(达卡大学) American International University - Bangladesh (AIUB)(孟加拉国美国国际大学)

专题命中 视觉定位与Grounding :grounding(title,title_cn);vision-language model(title,abstract);分类 cs.CV、cs.AI

AI总结 研究发现长视界视觉语言模型中,步级视觉接地信念预测分布外泛化能力,揭示了模型中间推理与视觉状态的一致性对鲁棒性的影响。

Comments Following the initial submission, we conducted additional experiments that materially changed our understanding of the problem. These new results do not support the central claim of the current manuscript. To avoid disseminating conclusions that we no longer consider adequately supported, we are withdrawing this version while we reassess the findings and prepare a substantially revised manuscript

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.13156 2026-07-07 cs.CV cs.AI 新提交 91%

Iterative Visual Thinking and the Self-Correction Mirage in VLM Grounding

迭代视觉思维:通过视觉反馈教会视觉语言模型空间自我修正

Animesh Tripathy, Aswanth Krishnan

机构 * QpiAI India Pvt. Ltd(QpiAI印度私人有限公司)

专题命中 视觉定位与Grounding :VLM(title,abstract);grounding(title,abstract);vision-language model(abstract);分类 cs.CV、cs.AI

AI总结 提出迭代视觉思维(IVT)框架,通过视觉反馈闭环和两阶段训练(SFT+GRPO),使视觉语言模型具备空间自我修正能力,在三个基准上提升指标2.4-3.2个百分点。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.10978 2026-03-12 cs.CV cs.AI 91%

GroundCount: Grounding Vision-Language Models with Object Detection for Mitigating Counting Hallucinations

GroundCount: 通过目标检测对视觉语言模型进行接地以缓解计数幻觉

Boyuan Chen, Minghao Shao, Siddharth Garg, Ramesh Karri, Muhammad Shafique

专题命中 视觉定位与Grounding :grounding(title,abstract);vision-language model(title);vision language model(abstract);VLM(abstract)

AI总结 GroundCount通过目标检测提供空间接地,缓解视觉语言模型计数幻觉,提升准确率并减少推理时间。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.15426 2025-07-17 cs.CV cs.AI 91%

Visual Position Prompt for MLLM based Visual Grounding

Wei Tang, Yanpeng Sun, Qinying Gu, Zechao Li

机构 * School of Computer Science and Engineering, Nanjing University of Science and Technology(南京理工大学计算机科学与工程学院) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

专题命中 视觉定位与Grounding :grounding(title,abstract);MLLM(title,abstract);LLaVA(abstract);multimodal large language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.10412 2026-08-12 cs.HC cs.CY 新提交 90%

When the Interviewer Is a Bot: Behavior, Breakdowns, and Trust in MLLM-Led Interviews

当面试官是机器人时:多模态大语言模型(MLLM)主导的面试中的行为、故障与信任

He Zhang, Kambinachi Chukwuma, ChanMin Kim, John M. Carroll

专题命中 视觉定位与Grounding :MLLM(title,title_cn);grounding(abstract)

AI总结 该研究通过构建InterviewBot系统开展实证研究,分析了MLLM主导面试的行为、故障与社会动态,为以人为中心的面试自动化提供设计启示。

Comments Accepted to ACM HCOMP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.04574 2026-08-06 cs.CL 新提交 90%

When Memory Lies: An Empirical Study of Spatial Memory Staleness in VLM Agents

当记忆“说谎”:VLM智能体中空间记忆过时的实证研究

Yushi Sun, Yanjie Zhang

机构 * Tencent LIGHTSPEED(腾讯光速工作室)

专题命中 视觉定位与Grounding :VLM(title,title_cn);grounding(abstract)

AI总结 该研究通过动态FrozenLake测试平台,探究了VLM智能体的空间记忆过时问题,发现文本可解性不代表视觉接地、信任过时记忆存在安全隐患、审核无法完全消除差距,明确了相关核心挑战。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.29649 2026-06-30 cs.CL 90%

Resolution Thresholds in VLM Detection of Harmful ASCII Art Across Construction Modes and Languages

VLM检测有害ASCII艺术的分辨率阈值:跨构建模式与语言的研究

Yikai Hua, Peter West

机构 * Department of Computer Science(计算机科学系) The University of British Columbia(不列颠哥伦比亚大学)

专题命中 视觉定位与Grounding :VLM(title,title_cn);vision-language model(abstract)

AI总结 研究图像分辨率如何影响VLM检测有害ASCII艺术,发现检测率在特定分辨率阈值以上急剧下降,且基于单词的模式最难检测,揭示了VLM内容审核系统的系统性漏洞。

Comments 13 pages, 9 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.17362 2026-06-17 cs.CV cs.AI cs.LG cs.RO 新提交 90%

DriveJudge: Rethinking Autonomous Driving Evaluation with Vision-Language Models

DriveJudge: 用视觉-语言模型重新思考自动驾驶评估

Xinglong Sun, Kevin Xie, Jenny Schmalfuss, Despoina Paschalidou, Xiuming Zhang, Sanja Fidler, Kashyap Chitta, Jose M. Alvarez

机构 * NVIDIA(英伟达)

专题命中 视觉定位与Grounding :VLM(summary_cn,abstract);vision-language model(title,abstract);grounding(abstract);分类 cs.CV、cs.AI、cs.LG

AI总结 提出DriveJudge,结合规则评估与VLM推理,通过选择性调用物理规则函数实现可解释且上下文感知的驾驶评估,在驾驶质量分类和轨迹偏好选择任务上超越现有方法。

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏