Mitigating Object Hallucinations in Large Vision-Language Models with Assembly of Global and Local Attention
专题命中 视觉定位与Grounding :vision-language model(title,abstract);grounding(abstract);分类 cs.CV、cs.AI
Comments Accepted by CVPR 2025
AI 大模型
视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。
专题命中 视觉定位与Grounding :vision-language model(title,abstract);grounding(abstract);分类 cs.CV、cs.AI
Comments Accepted by CVPR 2025
专题命中 视觉定位与Grounding :grounding(title,abstract);vision-language model(abstract);分类 cs.CV、cs.LG
Comments Accepted by CVPR 2025
专题命中 视觉定位与Grounding :vision language model(title,abstract);VLM(abstract);分类 cs.CV、cs.AI
专题命中 视觉定位与Grounding :vision-language model(title,abstract);grounding(abstract);分类 cs.CV、cs.AI
Comments ACL 2023 Outstanding Paper
专题命中 视觉定位与Grounding :grounding(title);vision-language model(abstract);multimodal large language model(abstract);分类 cs.CV、cs.LG
专题命中 视觉定位与Grounding :visual language model(title,abstract);visual question answering(abstract);分类 cs.CV、cs.AI
Comments accepted by SafeGenAI workshop of NeurIPS 2024
专题命中 视觉定位与Grounding :vision-language model(title,abstract);VLM(abstract);分类 cs.CV、cs.AI
专题命中 视觉定位与Grounding :vision language model(title,abstract);VLM(abstract);分类 cs.CV、cs.AI
专题命中 视觉定位与Grounding :MLLM(title,abstract);multimodal large language model(abstract);分类 cs.CV、cs.LG
Comments Accepted for publication at WACV 2025
专题命中 视觉定位与Grounding :grounding(title,abstract);MLLM(abstract);分类 cs.CV、cs.LG
Comments 7 pages, 3 figures
专题命中 视觉定位与Grounding :MLLM(title,abstract);multimodal large language model(abstract);分类 cs.CV、cs.AI
专题命中 视觉定位与Grounding :vision-language model(title,abstract);grounding(abstract);分类 cs.CV、cs.AI
Comments Under peer review
专题命中 视觉定位与Grounding :multimodal large language model(title,abstract);grounding(abstract);分类 cs.CV、cs.AI
Comments ACL 2024 (Findings)
专题命中 视觉定位与Grounding :vision-language model(title,abstract);VLM(abstract);分类 cs.CV、cs.LG
Comments ICLR 2024 Spotlight
专题命中 视觉定位与Grounding :multimodal large language model(title,abstract);grounding(abstract);分类 cs.CV、cs.AI
Comments CVPR 2024 camera-ready version, code is available at https://github.com/RenShuhuai-Andy/TimeChat
专题命中 视觉定位与Grounding :grounding(title,abstract);vision-language model(abstract);分类 cs.CV、cs.AI
Comments Preprint. 48 pages, 22 figures, 10 tables
专题命中 视觉定位与Grounding :grounding(title,abstract);LLaVA(abstract);分类 cs.CV、cs.AI
专题命中 视觉定位与Grounding :grounding(title,abstract);vision-language model(abstract);分类 cs.CV、cs.AI
专题命中 视觉定位与Grounding :grounding(title,abstract);visual reasoning(abstract);分类 cs.CV、cs.AI
Comments In CVPR 2023
专题命中 视觉定位与Grounding :grounding(title,abstract);visual question answering(abstract);分类 cs.CV、cs.AI
Comments Camera ready for Findings of EMNLP 2021
专题命中 视觉定位与Grounding :grounding(title,abstract);visual reasoning(abstract);分类 cs.AI、cs.LG
Comments Code available at https://github.com/SeverTopan/SATNet
视觉定位:一项综述
专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV
AI总结 本综述梳理视觉定位的发展与背景,总结近年进展与新挑战,定义规范研究设置,介绍相关数据集与应用,提出未来方向,是该领域最全面的综述,适合不同阶段研究者。
Comments Accepted by TPAMI 2025. We keep tracing related works at https://github.com/linhuixiao/Awesome-Visual-Grounding, article publication page: https://ieeexplore.ieee.org/abstract/document/11235566
Journal ref IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 48, no. 3, pp. 2749-2771, March 2026
MM-Conv: 一种多模态数据集和基准,用于上下文感知的3D对话中指代解析
机构 * KTH Royal Institute of Technology(皇家理工学院)
专题命中 视觉定位与Grounding :grounding(title,abstract);vision-language model(abstract);分类 cs.CV;VLM(comments)
AI总结 本文提出了一种多模态数据集和基准,用于在动态3D环境中实现上下文感知的指代解析,通过引入包含6.7小时第一人称VR交互的同步语音、动作、注视和3D场景几何数据的基准,以及一个两阶段的指代解析流水线,改进了对话中的指代解析性能。
Comments Extended version of the paper published at LREC 2026 (Palma de Mallorca, Spain), with expanded VLM baselines and inter-annotator agreement analysis
Journal ref Proceedings of the 15th Language Resources and Evaluation Conference (LREC 2026), Palma de Mallorca, Spain
OVIP-SG:用于小型细粒度物体映射与检索的开放词汇实例保留场景图
专题命中 视觉定位与Grounding :VLM(summary_cn,abstract);vision-language model(abstract)
AI总结 本文提出OVIP-SG框架,通过VLM等技术实现小型细粒度物体的实例保留映射与语言引导检索,在Replica数据集上的多项指标优于现有方法,且经实际机器人实验验证有效。
Comments 15 pages, 6 figures, including appendix
面向物联网汽车应用的全自动、感知部署的测试流水线
专题命中 视觉定位与Grounding :VLM(summary_cn,abstract);vision-language model(abstract)
AI总结 针对物联网汽车应用的测试难题,提出结合LLM、VLM及人在回路机制的感知部署测试流水线,经CPDS案例验证可实现全需求覆盖与高准确率,适配OEM-供应商测试工作流。
Comments Submitted, accepted and presented at IoTBDS26
面向吉隆坡城市交通的隐私保护数据集构建:结合空间车辆上下文过滤的接地视觉-语言检测
机构 * Faculty of Artificial Intelligence, Universiti Teknologi Malaysia(马来西亚理工大学人工智能学院)
专题命中 视觉定位与Grounding :grounding(summary_cn,abstract);分类 cs.CV、cs.AI、cs.LG
AI总结 针对吉隆坡热带城市交通场景的隐私保护数据集构建难题,提出结合Grounding DINO与空间车辆ROI包含引擎的自动化匿名化框架,在1266帧图像上实现约95%的匿名化成功率。
FARM: 使用关系空间记忆找到任何物体
机构 * UC Berkeley(加州大学伯克利分校) ; Stanford University(斯坦福大学)
专题命中 视觉定位与Grounding :VLM(summary_cn,abstract);grounding(abstract)
AI总结 提出FARM系统,通过实时构建包含几何、视觉语言描述和视角证据的开放词汇物体级记忆,并利用VLM解析查询和显式空间约束,在44k语言查询中Recall@5和Recall@10分别提升164%和224%,Accuracy@1提升35%。
CaptionFormer:时空对象的统一分割、跟踪与描述
机构 * Inria, École Normale Supérieure, CNRS, PSL Research University(法国国家科学研究中心、巴黎高等师范学院、国家科学研究中心、巴黎综合理工研究所) ; Google DeepMind(谷歌DeepMind)
专题命中 视觉定位与Grounding :VLM(summary_cn,abstract);分类 cs.CV、cs.AI、cs.LG
AI总结 提出 CaptionFormer 模型,通过利用 VLM 生成合成描述并扩展数据集,实现视频中对象轨迹的联合检测、分割、跟踪与描述,在三个基准上达到最优。
Comments 17 pages, 10 figures
SceneAligner: 在真实场景中实现基于3D的平面定位
机构 * Cornell University(康奈尔大学) ; Kempner Institute, Harvard University(哈佛大学 Kempner 院)
专题命中 视觉定位与Grounding :grounding(summary_cn,abstract);分类 cs.CV、cs.AI、cs.LG
AI总结 本文提出了一种在真实场景中实现基于3D重建的平面定位方法,通过将任务 grounding 在场景的重建3D表示中,解决了现有方法在大规模建筑和栅格化平面图中应用受限的问题。
Comments Project Page: https://Cornell-VAILab.github.io/SceneAligner
RAR:用于视觉识别的检索与排序增强多模态大语言模型
机构 * Shanghai Jiao Tong University(上海交通大学) ; Shanghai AI Laboratory(上海人工智能实验室) ; The Chinese University of Hong Kong(香港中文大学) ; MThreads, Inc.(MThreads公司) ; Nanyang Technological University(南洋理工大学)
专题命中 视觉定位与Grounding :MLLM(summary_cn,abstract_cn);multimodal large language model(abstract);分类 cs.CV、cs.AI、cs.LG
AI总结 RAR结合CLIP的检索与MLLM的排序,提升细粒度识别精度,增强少样本和零样本任务表现。
Comments Project: https://github.com/Liuziyu77/RAR