LLaVA-Grounding: Grounded Visual Chat with Large Multimodal Models
专题命中 视觉定位与Grounding :LLaVA(title,abstract);grounding(title,abstract);分类 cs.CV
AI 大模型
视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。
专题命中 视觉定位与Grounding :LLaVA(title,abstract);grounding(title,abstract);分类 cs.CV
基于专用分割模型实现具身智能视觉语言模型(Agentic VLMs)的细粒度车辆损伤评估
机构 * Northeastern University(东北大学)
专题命中 视觉定位与Grounding :grounding(title,abstract);VLM(abstract,abstract_cn);vision-language model(abstract);分类 cs.CV、cs.AI
AI总结 针对VLMs空间定位不可靠问题,本文提出TinyDamage架构,将空间定位委托给专用多任务分割模型,集成至7节点LangGraph智能体流程,在车辆损伤评估中大幅降低报告虚构率。
Comments 8 pages, 2 figures
PruneGround: 用于3D视觉定位的即插即用空间剪枝框架
机构 * Knovel Engineering Lab(Knovel工程实验室) ; University of Toronto(多伦多大学) ; Vector Institute(向量研究所) ; ELLIS Institute(ELLIS研究所) ; Max Planck Institute(马克斯·普朗克研究所)
专题命中 视觉定位与Grounding :grounding(title,abstract);VLM(abstract,abstract_cn);vision language model(abstract);分类 cs.CV、cs.AI
AI总结 提出PruneGround框架,通过语言引导的空间剪枝、多视角描述重构和LLM定位器,在剪枝区域高效定位目标,在多个基准上取得最优结果。
Comments Preprint
基于语义采样的医学图像空间定位
机构 * Case Western Reserve University(凯斯西储大学) ; Cleveland Clinic(克利夫兰诊所)
专题命中 视觉定位与Grounding :grounding(title,abstract);VLM(abstract,abstract_cn);vision language model(abstract);分类 cs.CV、cs.LG
AI总结 研究视觉语言模型在医学图像三维空间定位中的挑战,提出MIS-Ground基准和MIS-SemSam优化方法,通过语义采样提升定位精度13.06%。
Comments 10 pages, 2 figures. Accepted at MICCAI 2026
多模态生成式引擎优化:针对视觉-语言模型排序器的排名操纵
机构 * Georgetown University(乔治城大学) ; University of Southern California(南加州大学) ; University of Maryland, College Park(马里兰大学学院公园分校) ; Arizona State University(亚利桑那州立大学)
专题命中 视觉定位与Grounding :vision-language model(title,abstract);VLM(abstract,abstract_cn);grounding(abstract);分类 cs.AI、cs.LG
AI总结 提出多模态生成式引擎优化(MGEO)方法,通过联合优化图像扰动和文本后缀,利用视觉-语言模型内部跨模态知识耦合,实现对产品排名的有效操纵,揭示了多模态基础模型知识基础的脆弱性。
Comments Proceedings of the 4th Workshop on Towards Knowledgeable Foundation Models (KnowFM) at ACL 2026
Qwen3-VL-Seg: 解锁基于视觉-语言接地的开放世界指代分割
机构 * Tongyi Lab, Alibaba Group(通义实验室,阿里巴巴集团)
专题命中 视觉定位与Grounding :grounding(title,abstract);MLLM(abstract,abstract_cn);multimodal large language model(abstract);分类 cs.CV、cs.AI
AI总结 Qwen3-VL-Seg通过视觉-语言接地框架实现开放世界指代分割,采用轻量级的框引导掩码解码器,结合多尺度空间特征注入和迭代掩码感知查询优化,实现参数高效且精确的像素级分割。
MARVL:通过视觉-语言模型实现机器人操作的多阶段引导
机构 * School of Intelligent Science and Technology, Nanjing University, China(南京大学智能科学与技术学院) ; National Key Laboratory for Novel Software Technology, School of Artificial Intelligence, Nanjing University, China(南京大学新型软件技术国家实验室,人工智能学院) ; School of Artificial Intelligence, Nanjing University, China(南京大学人工智能学院) ; MACS Lab, University of Washington(华盛顿大学MACS实验室)
专题命中 视觉定位与Grounding :vision-language model(title,abstract);VLM(abstract,abstract_cn);grounding(abstract);分类 cs.CV、cs.LG
AI总结 本文提出MARVL,通过视觉-语言模型实现机器人操作的多阶段引导,解决传统密集奖励函数设计中的问题,提升样本效率和鲁棒性。
构图的艺术:用于组合视觉定位的注意力正则化训练
机构 * The University of British Columbia(不列颠哥伦比亚大学) ; Nanyang Technological University(南洋理工大学)
专题命中 视觉定位与Grounding :grounding(title,abstract);VLM(abstract,abstract_cn);vision-language model(abstract);分类 cs.CV、cs.LG
AI总结 本文提出CompART,通过分解标题为以对象为中心的短语并构建复合短语,提升多对象定位能力,验证了其在多种视觉语言模型上的有效性。
专题命中 视觉定位与Grounding :grounding(title,abstract);vision-language model(abstract);VLM(abstract);LLaVA(abstract)
Comments The paper is not yet mature and needs further improvement
机构 * U.S. Bank(美国银行)
专题命中 视觉定位与Grounding :vision language model(title,abstract);VLM(abstract);multimodal large language model(abstract);MLLM(abstract)
Comments 12 pages, 5 figures, 2 tables
机构 * Mohamed bin Zayed University of AI(莫扎德人工智能大学) ; University of Surrey(萨里大学) ; Jio Institute(乔研究所) ; UNSW(新南威尔士大学)
专题命中 视觉定位与Grounding :VLM(title,abstract);vision-language model(abstract);LLaVA(abstract);grounding(abstract)
Grounding-IQA: 多模态语言模型的接地图像质量评估
机构 * Shanghai Jiao Tong University(上海交通大学) ; Joy Future Academy(京东探索研究院) ; Huawei Noah’s Ark Lab(华为诺亚实验室) ; Westlake University(西湖大学) ; Huawei(华为)
专题命中 视觉定位与Grounding :grounding(title,abstract);visual question answering(abstract);multimodal large language model(abstract);MLLM(abstract)
AI总结 本文提出grounding-IQA任务范式,通过结合多模态指称与图像质量评估,实现更细粒度的图像质量感知。
Comments Accepted to ICLR 2026. Code is available at: https://github.com/zhengchen1999/Grounding-IQA
具有隐式和显式几何的3D感知视觉语言模型
机构 * Nanyang Technological University(南洋理工大学) ; DAMO Academy, Alibaba Group(阿里巴巴达摩院) ; HuPan Lab(湖畔实验室)
专题命中 视觉定位与Grounding :VLM(summary_cn,abstract);vision-language model(abstract);grounding(abstract);分类 cs.CV、cs.AI、cs.LG
AI总结 研究针对多数2D视觉输入的VLM处理3D任务困难的问题,提出VLM-IE3D框架,通过引入隐式和显式几何令牌及3D感知适配器,融合几何表示与视觉线索,在多种3D任务中表现优异。
Comments Accepted by ECCV 2026, Open Sourced
迈向人工通用教师:基于视觉-语言模型的程序几何数据生成与视觉语义
机构 * Freya
专题命中 视觉定位与Grounding :vision-language model(title,abstract);grounding(title);分类 cs.CV、cs.AI、cs.LG
AI总结 本文研究几何教育中的视觉解释问题,提出通过程序生成几何图示数据,结合视觉-语言模型进行领域特定微调,提升几何元素分割性能,并引入几何感知的评估指标。
Comments 12 pages, 7 figures
专题命中 视觉定位与Grounding :vision-language model(title,abstract);grounding(title);分类 cs.CV、cs.AI、cs.LG
Comments Accepted at EMNLP 2023 main conference
Gen2Physics:通过多视图材料分解将生成的三维网格与物理特性关联
机构 * University of Bristol(布里斯托大学) ; Google DeepMind(谷歌DeepMind) ; Google Research(谷歌研究院)
专题命中 视觉定位与Grounding :grounding(title,abstract);VLM(abstract,abstract_cn);vision-language model(abstract);分类 cs.CV
AI总结 Gen2Physics是一个将生成的三维网格与物理特性关联的自动化框架,通过多视图材料分解实现,其材料分割精度较现有方法提升超一倍,还能输出水密性各材料子网格。
HaReCAP:面向递归大语言模型智能体的习惯性动作 grounding 方法
机构 * North China Institute of Computer System Engineering(华北计算机系统工程研究所) ; University of Science and Technology of China(中国科学技术大学) ; China Information Security Research Institute Co., Ltd.(中国信息安全研究院有限公司)
专题命中 视觉定位与Grounding :grounding(title,title_cn);分类 cs.AI
AI总结 HaReCAP 是针对 ReCAP 的低侵入性扩展,通过编译叶子反射规则减少长视距具身任务中 LLM 的重复调用,在 Robotouille 和 ALFWorld 上显著降低了 token 消耗。
Comments 15 pages, 3 figures
Grounded-Exo2Ego:用于鲁棒外部视角到自我视角视频生成的结构化语义 grounding
机构 * NVIDIA(英伟达)
专题命中 视觉定位与Grounding :grounding(title,title_cn);分类 cs.CV
AI总结 本文提出Grounded-Exo2Ego框架,通过双分支视频扩散模型、相机重定位算法和自动合成数据引擎,在EgoExo4D数据集上大幅提升了外部视角到自我视角视频生成的性能。
Comments website url: https://research.nvidia.com/labs/amri/projects/grounded-exo2ego/
视觉定位解码器是否需要前馈网络?基于冻结的视觉-语言特征的受控研究
专题命中 视觉定位与Grounding :grounding(title,abstract);VLM(abstract,abstract_cn);vision-language model(abstract);分类 cs.CV
AI总结 本研究通过受控实验对比不同视觉定位解码器,发现仅注意力的4块解码器(A4)性能与含FFN的4块解码器相当,8块仅注意力解码器(A8)可弥补A4的微小差距且更优,同时A4能减少参数量与延迟。
Comments 14 pages, 8 figures, 5 tables. Code and project page: https://github.com/TarunTomar122/attention-is-all-you-need-for-vlms
ADOPD:面向多模态大语言模型(MLLM)的工业异常检测的参考特权在线策略蒸馏
机构 * Zhejiang University(浙江大学)
专题命中 视觉定位与Grounding :MLLM(title,title_cn);multimodal large language model(abstract);分类 cs.CV
AI总结 本研究针对工业异常检测问题,提出ADOPD参考特权在线策略蒸馏框架,通过参考感知教师监督仅查询的学生模型,在MMAD基准零样本推理下获77.31%平均准确率,优于Qwen3-VL-4B主干及其一样本设置。
一图胜千词:视觉语言模型如何在提升准确率的同时降低AI能源成本
专题命中 视觉定位与Grounding :vision language model(title);VLM(abstract,abstract_cn);vision-language model(abstract);grounding(abstract)
AI总结 该研究提出将时间序列编码为二维图的视觉语言模型,可大幅减少输入词元、降低推理能耗,同时在电信异常检测等任务中提升准确率,解决了LLM处理多维度KPI数据的低效问题。
Comments Accepted at the 14th European Conference on Renewable Energy Systems (ECRES), July 7--9, 2026, London, UK
EgoAfford:基于自我中心指称分割的面向任务的可供性定位
专题命中 视觉定位与Grounding :grounding(title,abstract);VLM(abstract,abstract_cn);multimodal large language model(abstract);分类 cs.CV
AI总结 针对部件级可供性定位难以适配复杂多步桌面任务的问题,提出基准EgoAfford及参考模型EgoLens,验证了下一步推理与动作角色条件部件定位的互补挑战,为感知与规划联合研究提供基础。
ReGround:通过自诊断与视觉重检验恢复多步推理中的视觉接地
机构 * University of Science and Technology of China(中国科学技术大学) ; School of Artificial Intelligence and Data Science(人工智能与数据科学学院) ; State Key Laboratory of Precision and Intelligent Chemistry(精准与智能化学国家重点实验室)
专题命中 视觉定位与Grounding :grounding(title,abstract);VLM(abstract,abstract_cn);vision-language model(abstract);分类 cs.CV
AI总结 ReGround是无需架构修改或外部工具的两阶段框架,通过自诊断与视觉重检验解决VLMs多步推理中的视觉接地丢失问题,在八个基准上获一致增益且推理开销适度。
Comments Accepted to ACM Multimedia 2026 (MM '26). 8 pages main text, 4 figures, plus appendix
CROSS:用于遥感指称分割的级联蒸馏与双约束 grounding
专题命中 视觉定位与Grounding :grounding(title,title_cn);VLM(abstract,abstract_cn);分类 cs.CV
AI总结 针对遥感指称分割中架构弱耦合与语义偏差问题,本文提出CROSS范式,通过LGCD与PSCL实现性能突破,成为RRSIS的稳健新方案。
Comments Accepted at the European Conference on Computer Vision (ECCV) 2026. 20 pages, 6 figures, and 5 tables. Tingzhang Luo and Ruizhong Liu contributed equally. Jianyuan Guo is the corresponding author. Project page: https://clarence-cv.github.io/CROSS/
GeoArbiter:面向遥感多模态大语言模型的可验证性导向 grounding 方法
机构 * University of Minnesota, Twin Cities(明尼苏达大学双城分校)
专题命中 视觉定位与Grounding :grounding(title,title_cn);multimodal large language model(abstract);分类 cs.LG
AI总结 该研究针对遥感多模态大语言模型的事实断言问题,提出无需训练的 GeoArbiter 流程,通过仅注入图像无法验证的地理事实,有效降低了模型的断言级幻觉并提升了对冲突记录的鲁棒性。
用于手语翻译的注意力引导视觉-语言模型
机构 * Rochester Institute of Technology(罗切斯特理工学院) ; University of Virginia(弗吉尼亚大学) ; National Technical Institute for the Deaf(国家聋人技术学院)
专题命中 视觉定位与Grounding :vision-language model(title,abstract);VLM(abstract,abstract_cn);grounding(abstract);分类 cs.CV
AI总结 针对视觉-语言模型在手语翻译中时空视觉定位差的问题,提出AttnSign框架,通过空间注意力监督和RL运动节奏引导提升性能,在How2Sign和OpenASL基准上表现优于现有方法。
ST-Veto:通过泰勒预测和视觉基础实现用于扩散多模态语言模型的时空令牌否决
专题命中 视觉定位与Grounding :grounding(title);VLM(abstract,abstract_cn);vision language model(abstract);multimodal large language model(abstract)
AI总结 研究针对扩散多模态大语言模型推理不足问题,提出无需训练的ST-Veto方法,利用二阶泰勒预测和图像注意力质量否决不稳定及弱基础令牌,与更安全候选交换,在多基准测试中优于其他方法,提升准确率且无额外成本。
Comments ICML 2026 - main
SplatReasoner:通过新颖视图合成增强具身推理与基础能力
机构 * POSTECH ; KAIST(韩国科学技术院) ; ETRI(韩国电子电信研究院) ; NVIDIA(英伟达)
专题命中 视觉定位与Grounding :grounding(title,abstract);VLM(abstract,abstract_cn);vision-language model(abstract);分类 cs.CV
AI总结 研究针对视觉语言模型应用于具身场景理解受固定视角限制的问题,提出SplatReasoner框架,利用3D高斯点云将新颖视图合成引入推理过程,经实验验证该方法能提升具身推理和3D基础能力。
Comments Accepted at ECCV 2026. Project page: https://splatreasoner.github.io/
BLPR: 通过置信度驱动的VLM回退实现视点和光照变化下的鲁棒车牌识别
机构 * Queen Mary University of London(伦敦玛丽女王大学)
专题命中 视觉定位与Grounding :VLM(title,title_cn);vision-language model(abstract);分类 cs.CV
AI总结 本文提出BLPR框架,通过合成数据预训练和真实数据微调提升Bolivian车牌识别鲁棒性,引入轻量视觉语言模型作为置信度回退机制,实现89.6%的字符识别准确率。
面向科学传播中图表-图像连贯性锚定的多模态推理类型学
机构 * New Jersey Institute of Technology(新泽西理工学院) ; Brooklyn College(布鲁克林学院)
专题命中 视觉定位与Grounding :grounding(title,abstract);VLM(abstract,abstract_cn);vision-language model(abstract);分类 cs.CV
AI总结 本研究提出R1-R5多模态推理类型学,刻画科学文献中图表、图像与文本的协同推理缺口,可系统识别连贯性锚定状态,缩小不同人群解读科学结论的认知差距。