arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 26403 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7473 篇

2508.05008 2026-05-15 cs.CV 84%

Multimodal Causal-Driven Representation Learning for Generalizable Medical Image Segmentation

多模态因果驱动表示学习用于通用化医学图像分割

Xusheng Liang, Lihua Zhou, Nianxin Li, Miao Xu, Ziyang Song, Dong Yi, Jinlin Wu, Jiawei Ma, Hongbin Liu, Zhen Lei, Jiebo Luo

机构 * City University of Hong Kong(香港城市大学) Shenzhen Loop Area Institute(深圳河套学院) CAIR, HKISI, Chinese Academy of Sciences(中国科学院计算智能研究所) UESTC(电子科技大学) MAIS, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)

专题命中 视觉定位与Grounding :VLM(summary_cn,abstract);vision-language model(abstract);分类 cs.CV

AI总结 本文提出MCDRL框架,结合因果推理与VLM解决医学图像分割的领域泛化问题,通过文本提示识别病变区域并消除领域特定影响,提升分割准确性与泛化能力。

Comments Accepted by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.12237 2026-05-13 cs.CV 84%

UHR-Micro: Diagnosing and Mitigating the Resolution Illusion in Earth Observation VLMs

UHR-Micro: 诊断和缓解地球观测VLMs中的分辨率错觉

Shuo Ni, Tong Wang, Jing Zhang, He Chen, Haonan Guo, Ning Zhang, Bo Du

机构 * National Key Laboratory of Science and Technology on Space-Born Intelligent Information Processing(国家空间智能信息处理科技重点实验室) Beijing Institute of Technology(北京理工大学) School of Computer Science(计算机学院) Wuhan University(武汉大学) Zhongguancun Academy(中关村学院) State Key Laboratory of Information Engineering in Surveying, Mapping and Remote Sensing(测绘遥感信息工程国家重点实验室) Hong Kong Polytechnic University(香港理工大学)

专题命中 视觉定位与Grounding :VLM(summary_cn,abstract_cn);vision-language model(abstract);grounding(abstract);分类 cs.CV

AI总结 UHR-Micro通过11253条指令和1212张超高清图像,评估VLM在高分辨率地球观测图像中的空间极限表现,揭示高分辨率输入下空间定位和证据解析的失败,并提出MAP代理以证据为中心提升微尺度感知。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.09598 2026-05-13 cs.CV 84%

SoccerLens: Grounded Soccer Video Understanding Beyond Accuracy

SoccerLens: 超越准确性的 grounded 球赛视频理解

Ismael Elsharkawi, Ahmed Sait, Silvio Giancola, Bernard Ghanem, Hossam Sharara, Abdelrahman Eldesokey

机构 * Department of Computer Science and Engineering, The American University in Cairo(美国亚历山大大学计算机科学与工程系) Image And Visual Understanding Lab (IVUL), KAUST(卡塔尔大学图像与视觉理解实验室)

专题命中 视觉定位与Grounding :grounding(summary_cn,abstract);vision-language model(abstract);分类 cs.CV

AI总结 本文提出 SoccerLens 基准测试,评估球赛视频理解中视觉 grounding 的有效性,发现现有模型在准确率高但 grounding 表现差,揭示了时空复杂领域中 grounded 评估的重要性。

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.22367 2026-04-27 cs.CL cs.AI 84%

CNSL-bench: Benchmarking the Sign Language Understanding Capabilities of MLLMs on Chinese National Sign Language

CNSL-bench:用于评估多模态大语言模型在中文国家手语理解能力的基准

Rui Zhao, Xuewen Zhong, Xiaoyun Zheng, Jinsong Su, Yidong Chen

机构 * School of Informatics, Xiamen University, China(厦门大学信息学院) Key Lab of Digital Protection and Intelligent Processing of Intangible Cultural Heritage of Fujian-Taiwan (XMU), Ministry of Culture and Tourism, China(福建省-台湾非物质文化遗产数字化保护与智能处理重点实验室(XMU),文化和旅游部,中国) National Language Resources Monitoring and Research Center for Education and Teaching Media, Xiamen University, China(教育与教学媒体语言资源监测与研究中心,厦门大学,中国)

专题命中 视觉定位与Grounding :grounding(summary_cn,abstract);multimodal large language model(abstract);分类 cs.AI

AI总结 本文提出CNSL-bench,首个评估多模态大语言模型在中文国家手语理解能力的基准,通过权威 grounding、多模态覆盖和手部动作多样性,评估21个模型,发现当前模型在手语理解上仍显著劣于人类表现。

Comments Accepted as the Main Conference at ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.19326 2024-05-30 cs.CV cs.GR cs.HC 84%

Reasoning3D -- Grounding and Reasoning in 3D: Fine-Grained Zero-Shot Open-Vocabulary 3D Reasoning Part Segmentation via Large Vision-Language Models

Tianrun Chen, Chunan Yu, Jing Li, Jianqi Zhang, Lanyun Zhu, Deyi Ji, Yong Zhang, Ying Zang, Zejian Li, Lingyun Sun

专题命中 视觉定位与Grounding :vision-language model(title);grounding(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.23978 2026-08-26 cs.AI cs.CV 新提交 84%

When Seeing Is Not Enough: Benchmarking Interactive Visual Grounding in LVLMs

当视觉不再足够:评估大视觉语言模型(LVLMs)中的交互式视觉定位

Zhengxiang Wang, Owen Rambow

机构 * Stony Brook University(石溪大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);vision-language model(abstract);分类 cs.CV、cs.AI

AI总结 本文针对大视觉语言模型(LVLMs)提出交互式视觉定位的受控评估框架,发现当前LVLMs表现低于人类基线,主动提问式定位困难且校准不佳,该任务仍具挑战性。

Comments EMNLP 2026 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.23928 2026-08-26 cs.CV cs.AI 新提交 84%

RefineRank: Joint Box Refinement and Ranking for Surgical Spatio-Temporal Grounding

RefineRank:用于外科时空定位的联合框优化与排序

Linzhe Jiang, Jiayuan Huang, Changhao Zhang, Chunyang Jiang, Zhehua Mao, Mobarak I. Hoque

机构 * UCL Hawkes Institute, University College London(伦敦大学学院霍克斯研究所) King’s College London(伦敦国王学院) School of Medicine, Nankai University(南开大学医学院) University of Manchester(曼彻斯特大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);vision language model(abstract);分类 cs.CV、cs.AI

AI总结 RefineRank通过RefineNet模块结合冻结医学视觉语言模型与开放集检测器,实现外科时空定位的联合框优化与排序,在MedVidBench上取得较高STG mIoU,提升了定位性能且无需重训主干。

Comments 17 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.23474 2026-08-26 cs.CL cs.AI cs.CV 版本更新 84%

What's the Catch? Evaluating Temporal Consistency in Vision-Language Models

有何玄机?评估视觉-语言模型的时间一致性

Marek Hradil, Danae Sánchez Villegas

机构 * University of Copenhagen(哥本哈根大学)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);grounding(abstract);分类 cs.CV、cs.AI

AI总结 本研究提出TimeCatch基准,发现视觉-语言模型可检测单帧异常但难整合跨帧信息,在时间异常检测上表现接近随机水平。

Comments 17 pages, ACL format

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.22232 2026-08-26 cs.AI cs.CL cs.CV cs.MM 版本更新 84%

Beyond What Meets the Eye: Unveiling Situational Illusions for Multimodal Large Language Models

超越表象:揭示多模态大语言模型的情境错觉

Zhiming Yang, Zhuoxi Xiong, Donglin Zhou, Wenjun Wei, Shiyao Cui, Jinqiao Shi

专题命中 视觉定位与Grounding :multimodal large language model(title,abstract);grounding(abstract);分类 cs.CV、cs.AI

AI总结 本文针对多模态大语言模型(MLLMs)面临的情境错觉问题,构建分类体系并推出MSIBench基准,发现27种模型配置均易受6类错觉影响,通过提示工程和监督微调最多提升性能20%。

Comments 9 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.15821 2026-08-12 cs.CL cs.AI cs.LG 版本更新 84%

The Truth Stays in the Family: Enhancing Contextual Grounding via Inherited Truthful Heads in Model Lineages

真相留在家族中:通过模型谱系中继承的真相头增强上下文基础

Miso Choi, Seonga Choi, Mincheol Kwon, Woosung Joung, Jinkyu Kim, Jungbeom Lee

机构 * University of Science and Technology of China(中国科学技术大学)

专题命中 视觉定位与Grounding :grounding(title);MLLM(abstract,abstract_cn);分类 cs.AI、cs.LG

AI总结 研究发现基础LLM与下游变体间存在上下文真相分数的强继承性,提出TruthProbe软门控策略放大真相头以提升上下文真实性并减少多模态幻觉。

Comments Accepted at ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.08009 2026-08-11 cs.CV cs.AI 新提交 84%

Evidence-Grounded Forensic Reasoning for Detecting and Grounding Multi-Modal Media Manipulation

基于证据的取证推理:检测与定位多模态媒体篡改

Yichun Yeh, Yiheng Li, Xiaobo Hu, Zhen Lei, Yang Yang

专题命中 视觉定位与Grounding :grounding(title,abstract);MLLM(abstract_cn);分类 cs.CV、cs.AI

AI总结 本文针对多模态媒体篡改检测与定位问题,提出基于证据的取证推理框架,结合锚定-验证推理链、可验证奖励系统与模态解耦优势路由机制,实现最优性能与可解释性的统一。

Comments accepted by ACM MM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.00473 2026-08-04 cs.CV cs.AI cs.CL 新提交 84%

CrossProjection: Geometric Grounding Beyond Viewpoint Change in Architectural Drawings

CrossProjection:建筑图纸中超越视角变化的几何定位

Kaho Li, Pengyu Zeng, Yuqin Dai, Jun Yin, Tianjing Feng, Shuai Lu

机构 * Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院) University College London(伦敦大学学院)

专题命中 视觉定位与Grounding :grounding(title,abstract);vision-language model(abstract);分类 cs.CV、cs.AI

AI总结 该研究提出CrossProjection方法评估视觉语言模型在异构建筑视图间的几何定位能力,发现模型在自由几何定位上性能脆弱,仅封闭选择成功不代表具备可靠几何定位能力。

Comments Initial controlled diagnostic study on 23 natural drawing sets and three VLMs; broader model, building, repeated-inference, and human coverage is planned for a subsequent version

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.01982 2026-08-03 cs.CV cs.AI q-bio.BM 版本更新 84%

MolSight: A Graph-Aware Vision-Language Model for Unified Chemical Image Understanding

MolSight: 一种用于统一化学图像理解的图感知视觉语言模型

Wenda Wang, Yihan Tong, Yuwei Hu, Xuchen Pan, Zhewei Wei, Yaliang Li, Bolin Ding

机构 * Renmin University of China(中国人民大学)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);grounding(abstract);分类 cs.CV、cs.AI

AI总结 提出MolSight框架,通过分子拓扑模块和分子接地模块增强视觉语言模型对分子图像的结构理解,在多项化学视觉理解任务中显著优于现有模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.27558 2026-07-31 cs.CV cs.AI 新提交 84%

Drawing-Recode: Annotation Grounding for Parametric CAD Code Generation from Raster 2D CAD Drawings

Drawing-Recode:基于栅格2D CAD图纸的参数化CAD代码生成的标注定位方法

Mingi Kim, Yongjun Kim, Hyungki Kim

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV、cs.AI

AI总结 Drawing-Recode是一种从栅格2D CAD图纸生成参数化CAD代码的框架,通过图像编码器、文本识别模块、交叉注意力与AGL实现标注定位,性能优于基线且对工业扫描图纸鲁棒,可助力制造自动化。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09611 2026-07-29 cs.CV cs.AI cs.CR 版本更新 84%

AGMark: Attention-Guided Dynamic Watermarking for Large Vision-Language Models

AGMark:基于注意力的动态水印技术用于大视觉-语言模型

Yue Li, Xin Yi, Dongsheng Shi, Yongyi Cui, Gerard de Melo, Linlin Wang

机构 * East China Normal University(华东师范大学) Hasso Plattner Institute(哈索普拉特纳研究所) University of Potsdam(波茨坦大学)

专题命中 视觉定位与Grounding :vision-language model(title);vision language model(abstract);grounding(abstract);分类 cs.CV、cs.AI

AI总结 AGMark通过动态注意力机制提升大视觉-语言模型的水印可靠性,实现高质量生成和强视觉语义保真度。

Comments KDD 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.24570 2026-07-28 cs.CV cs.AI 新提交 84%

The Visual Bottleneck: Sparse-Frame Adaptation of MLLMs for Joint Spatial-Temporal Video Grounding

视觉瓶颈:用于联合时空视频定位的多模态大语言模型的稀疏帧适配

Jiameng Zhang, Srikanth Madikeri

机构 * University of Zurich(苏黎世大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);multimodal large language model(abstract);分类 cs.CV、cs.AI

AI总结 研究针对多模态大语言模型在视频定位中训练与部署条件不匹配问题,通过实验表明视觉特征提取是稀疏帧输入瓶颈,提出适配特定层及边界感知采样策略,证明训练策略对稀疏帧视频定位比模型规模更关键。

Comments 15 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.19857 2026-07-23 cs.CV cs.AI 新提交 84%

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos

用于流式航空视频中小目标理解的内存增强多模态大语言模型

Penglei Sun, Yehua Huang, Zhuoli Tao, Xiang Li, Runwei Guan, Yaoxian Song, Kaiyong Zhao, Henghui Ding, Bo Han, Yang Yang, Xiaowen Chu

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) University of Freiburg(弗莱堡大学) Hangzhou City University(杭州城市大学) XGRIDS(XGRIDS公司) Fudan University(复旦大学) Hong Kong Baptist University(香港浸会大学)

专题命中 视觉定位与Grounding :multimodal large language model(title,abstract);MLLM(abstract);分类 cs.CV、cs.AI

AI总结 研究针对流式航空视频中小目标理解难题,提出像素级开放词汇数据集DroneEyes,以及含语义感知令牌路由器和分层内存库的多模态大语言模型SkyAnchor,从数据和方法角度应对挑战。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.15942 2026-07-20 cs.CV cs.LG 新提交 84%

More with Less: a Large Scale Remote Sensing VLM with a Simple Recipe

以少胜多:一种简单方法构建的大规模遥感视觉语言模型

Stefan Maria Ailuro, Mario Markov, Mohammad Mahdi, Luc Van Gool, Danda Pani Paudel

机构 * INSAIT, Sofia University “St. Kliment Ohridski”(索非亚大学“圣克莱门特·奥赫里德斯基”信息与自动化研究所)

专题命中 视觉定位与Grounding :VLM(title);vision-language model(abstract);grounding(abstract);分类 cs.CV、cs.LG

AI总结 研究针对遥感视觉语言模型,质疑架构专业化必要性,提出通用模型经大规模跨数据和任务训练可获好性能。核心方法是用单一语言策略及多任务强化学习框架训练。主要贡献是在多基准测试中取得竞争力结果,证明数据规模更关键。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23020 2026-07-08 cs.CV cs.AI 版本更新 84%

OpenGround: Planning-based Online Perception for Open-World 3D Visual Grounding

OpenGround:基于规划的开放世界3D视觉定位在线感知

Wenyuan Huang, Zhenyu Zhang, Zhao Wang, Zhou Wei, Ting Huang, Fang Zhao, Jian Yang

机构 * Nanjing University, School of Intelligent Science and Technology(南京大学智能科学与技术学院) China Mobile Zijin Innovation Institute(中国移动紫金创新院)

专题命中 视觉定位与Grounding :grounding(title,abstract);visual language model(abstract);分类 cs.CV、cs.AI

AI总结 针对开放世界3D视觉定位中现有方法的局限,提出OpenGround框架,集成任务链规划与上下文引导感知,还构建OpenTarget数据集。实验表明其在多个数据集上性能优异,在开放世界评估中有显著提升。

Comments ECCV2026, 46 pages, 13 figures, 15 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.09945 2026-07-03 cs.CV cs.AI cs.CL 版本更新 84%

Cross-Cultural Value Attribution in Large Vision-Language Models

跨文化价值观意识在大型视觉-语言模型中的体现

Phillip Howard, Xin Su, Kathleen C. Fraser

机构 * Thoughtworks University of Ottawa(渥太华大学)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);grounding(abstract);分类 cs.CV、cs.AI

AI总结 本文研究了图像中文化背景如何影响大型视觉-语言模型对人物道德、伦理和政治价值观的判断,通过多维分析揭示模型对文化价值观差异的认知。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.18481 2026-07-01 cs.CV cs.LG 版本更新 84%

T-QPM: Enabling Temporal Out-Of-Distribution Detection and Domain Generalization for Vision-Language Models in Open-World

T-QPM:面向开放世界视觉语言模型的时间分布外检测与域泛化

Aditi Naiknaware, Salimeh Sekeh

机构 * San Diego State University(圣地亚哥州立大学)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);VLM(abstract_cn);分类 cs.CV、cs.LG

AI总结 提出T-QPM框架,通过时间四元模式匹配和轻量级融合权重学习,增强视觉语言模型在动态环境下的分布外检测和协变量偏移鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07683 2026-06-30 cs.CV cs.AI 84%

TAR: Temporal Anchor-Constrained Reasoning for Video Temporal Grounding

TAR:基于时间锚的视频时间定位推理

Chaohong Guo, Xun Mo, Yongwei Nie, Fei Ma, Xuemiao Xu, Chengjiang Long

机构 * South China University of Technology(华南理工大学) Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ)(广东人工智能与数字经济实验室(深圳)) Bytedance Inc.(字节跳动有限公司)

专题命中 视觉定位与Grounding :grounding(title,abstract);vision-language model(abstract);分类 cs.CV、cs.AI

AI总结 TAR通过引入时间锚机制,提升视频时间定位任务中推理过程的可信度和自主性,采用自举方法生成高质量推理轨迹,实现更精确的最终预测。

Comments Accepted by ECCV2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.25491 2026-06-25 cs.CV cs.AI 新提交 84%

HG-Bench: A Benchmark for Multi-Page Handwritten Answer-Region Grounding in Automated Homework Assessment

HG-Bench: 自动作业批改中多页手写答案区域定位的基准

Chuangxin Zhao, Boyan Shi, Yanling Wang, Yijian LU, Canran Xiao, Jiali Chen, Jun Xia, Yan Wang, Ji Qi, Juanzi Li

专题命中 视觉定位与Grounding :grounding(title,abstract);VLM(abstract_cn);分类 cs.CV、cs.AI

AI总结 针对自动作业批改中多页手写答案区域定位的缺失评估,提出HG-Bench基准,包含500个带层级标注的样本,并设计页面感知评估协议,揭示现有模型在步骤级定位上的能力差距。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.24759 2026-06-24 cs.CV cs.AI 新提交 84%

UniDrive: A Unified Vision-Language and Grounding Framework for Interpretable Risk Understanding in Autonomous Driving

UniDrive: 面向自动驾驶可解释风险理解的统一视觉-语言与定位框架

Xiaowei Gao, Pengxiang Li, Yitai Cheng, Ruihan Xu, James Haworth, Stephen Law, Yun Ye

机构 * organization= Department of Earth Science \& Engineering, Imperial College London , city= London , postcode= SW7 2AZ , country= United Kingdom organization= SpaceTimeLab, Department of Civil, Environmental Geomatic Engineering, University College London , city= London , postcode= WC1E 6BT , country= United Kingdom organization= Department of Computing, The Hong Kong Polytechnic University , city= Hong Kong , country= China organization= Trinity College, University of Oxford , city= Oxford , postcode= OX1 3BH , country= United Kingdom organization= Department of Geography, University College London , city= London , postcode= WC1E 6BT , country= United Kingdom organization= Centre for Global Infrastructure Resilience, The Bartlett School of Sustainable Construction, University College London , city= London , postcode= WC1E 7HB , country= United Kingdom

专题命中 视觉定位与Grounding :grounding(title,abstract);multimodal large language model(abstract);分类 cs.CV、cs.AI

AI总结 提出UniDrive框架,通过融合时序推理与高分辨率感知分支,联合生成风险描述和边界框定位,在DRAMA-Reasoning基准上超越现有方法,提升小目标定位和可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.01733 2026-06-16 cs.CV cs.AI 版本更新 84%

GEASS: Gated Evidence-Adaptive Selective Caption Trust for Vision-Language Models

GEASS: 基于证据适应的门控选择性描述信任机制用于视觉-语言模型

Zeshang Li, Shuoyang Zhang

机构 * University of International Relations(国际关系大学)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);grounding(abstract);分类 cs.CV、cs.AI

AI总结 本文提出GEASS,一种无需训练的模块,通过门控、加权和证据标准来决定模型在每个查询中消耗多少描述信息,从而提升视觉-语言模型的准确性。

Comments 18 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.11889 2026-06-11 cs.CV cs.AI cs.RO 新提交 84%

Task-Aligned Stability Analysis of Vision-Language Models for Autonomous Driving Hazard Detection

面向自动驾驶危险检测的视觉-语言模型任务对齐稳定性分析

Everett Richards

机构 * Everett Richards(埃弗里特·里奇ards)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);VLM(abstract_cn);分类 cs.CV、cs.AI

AI总结 研究视觉-语言模型在自动驾驶危险检测中,嵌入漂移与任务对齐危险分数变化的关系,发现不同腐败类型导致不同的失效模式,建议基准测试包含任务对齐稳定性指标。

Comments 8 pages (5 main body + 3 references / appendices). ICML 2026 Workshop on Combining Theory and Benchmarks (CTB)

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.01612 2026-06-02 cs.CV cs.LG 84%

Self-Improving Small Object Grounding in LVLMs

LVLMs中的自改进小目标定位

Tianze Yang, Yucheng Shi, Ruitong Sun, Ninghao Liu, Jin Sun

机构 * University of Georgia(佐治亚大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);vision language model(abstract);分类 cs.CV、cs.LG

AI总结 利用LVLMs内部注意力模式,通过轻量级IoU回归器或无需训练的注意力熵选择器,从多个候选框中选出最佳框,实现小目标定位的自改进。

Comments 29 Pages, 15 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.00275 2026-06-02 cs.CV cs.AI 84%

Hyperbolic and Evidence-Prioritized Experts for Large Vision-Language Models

超几何与证据优先专家用于大型视觉-语言模型

Zijie Zhou, Dandan Zhu, Hangxiangpan Wang, Heng Zhang, Huishen Jiao, Yi Zhao

机构 * China University of Petroleum (Beijing)(中国石油大学(北京)) Hainan Institute of China University of Petroleum (Beijing)(中国石油大学(北京)海南学院) South China Normal University(华南师范大学)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);grounding(abstract);分类 cs.CV、cs.AI

AI总结 针对大型视觉-语言模型中视觉与语言模态的不对称性,提出AsyMoE架构,通过超几何跨模态专家和证据优先语言专家分别建模层级关系与保持上下文基础,在减少参数的同时提升性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.26661 2026-05-27 cs.CV cs.AI 84%

Respecting Modality Gap in Post-hoc Out-of-distribution Detection with Pre-trained Vision-Language Models

在预训练视觉语言模型的后验分布外检测中尊重模态差距

Yuanwei Hu, Bo Peng, Yadan Luo, Zhen Fang, Ling Chen, Jie Lu

机构 * The University of Queensland(昆士兰大学) University of Technology Sydney(悉尼科技大学)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);VLM(abstract_cn);分类 cs.CV、cs.AI

AI总结 针对预训练视觉语言模型在后验分布外检测中文本原型与视觉原型存在模态差距的问题,提出在线伪监督框架直接在视觉特征空间学习类原型,实现新最优性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.18803 2026-04-28 cs.CV cs.AI 84%

LLM-as-Judge Framework for Evaluating Tone-Induced Hallucination in Vision-Language Models

用于评估视觉-语言模型中语气诱导幻觉的LLM-as-Judge框架

Zhiyuan Jiang, Weihao Hong, Xinlei Guan, Tejaswi Dhandu, Miles Q. Li, Meng Xu, Kuan Huang, Umamaheswara Rao Tida, Bingyu Shen, Daehan Kwak, Boyang Li

机构 * Kean University(凯恩大学) North Dakota State University(北达科他州立大学) McGill University(麦吉尔大学) University of Notre Dame(圣母大学)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);grounding(abstract);分类 cs.CV、cs.AI

AI总结 本文提出Ghost-100基准,通过结构化提示强度框架评估不同模型在提示压力下的表现,发现H-Rate和H-Score在模型家族间显著分化,且某些模型对语气敏感度非单调。

Comments 23 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏