Interactive Reinforcement Learning for Object Grounding via Self-Talking
专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI
Comments NIPS 2017 - Visually-Grounded Interaction and Language (ViGIL) Workshop
AI 大模型
视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。
专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI
Comments NIPS 2017 - Visually-Grounded Interaction and Language (ViGIL) Workshop
专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV
专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI
Comments accepted at NeSy 2017; The final version of this paper is available at http://ceur-ws.org/Vol-2003/
专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV
Comments Presented at NIPS 2017 Symposium on Interpretable Machine Learning
专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV
专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV
Comments Spotlight in ICCV 2017
专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI
Comments 8 pages, 4 figures, Accepted at RSS 2017 Workshop on Spatial-Semantic Representations in Robotics
专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI
专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI
Comments 21 pages, 7 figures, preliminary version appeared at GTTV'15
专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI
Comments 11 pages, 3 figures
专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI
专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI
Journal ref Journal Of Artificial Intelligence Research, Volume 38, pages 223-269, 2010
专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI
Comments Appears in the Proceedings of the 19th International Conference on Applications of Declarative Programming and Knowledge Management (INAP 2011)
引用溯源:通过法律引用图检测和减少LLM引用幻觉
机构 * LEX AI LLC
专题命中 视觉定位与Grounding :grounding(title,abstract)
AI总结 提出引用溯源(CG)指标,利用乌克兰法院判决的引用图(1.008亿判决,5.02亿边)检测LLM法律引用幻觉,并通过CG-DPO方法(基于真实判决构建偏好对)减少幻觉,在100个法律查询上CG为0.791-0.873,幻觉率13-21%。
Comments 21 pages, 4 figures, 5 tables. Substantially revised: title, framing and several v1 results changed. Adds a coverage sweep and a separability analysis; corrects the DPO configuration, the density-accuracy correlation and the qualitative examples. Code and data: https://huggingface.co/datasets/overthelex/citation-grounding-eval
专题命中 视觉定位与Grounding :grounding(title,abstract)
Comments This is the second part of the general framework for a constructive type theory presented in the paper Functionals and hardware arXiv:1501.03043. The version is final. The research on the grounding of Mathematics is continued in the paper {\em Asymptotic combinatorial constructions of Geometries} available at arXiv:1904.05173
ExtrinSplat:解耦几何与语义以实现3D高斯散射中的开放词汇理解
机构 * Guangdong Provincial Key Laboratory of Ultra High Definition Immersive Media Technology, Shenzhen Graduate School, Peking University(广东省超高清沉浸式媒体技术重点实验室,北京大学深圳研究生院) ; School of Computer Science and Technology, Tianjin University(天津大学计算机科学与技术学院) ; Guangdong Bohua UHD Innovation Center Co., Ltd.(广东博华超高清创新中心有限公司)
专题命中 视觉定位与Grounding :VLM(abstract,abstract_cn);vision-language model(abstract);分类 cs.CV、cs.AI
AI总结 ExtrinSplat通过解耦几何与语义,利用视觉语言模型生成轻量文本假设,提升3D高斯散射中开放词汇物体选择和语义分割的性能和效率。
Comments Accepted to CVPR 2026
Defake-o3:从推测性理由到可验证证据的可解释AIGI检测
机构 * Shanghai Jiao Tong University(上海交通大学) ; Ant Group(蚂蚁集团)
专题命中 视觉定位与Grounding :MLLM(abstract,abstract_cn);grounding(abstract);分类 cs.CV、cs.AI
AI总结 Defake-o3是结合交互式视觉搜索与证据验证器的可解释AIGI检测器,构建了GroundFake数据集和FakeFrontier基准,在多基准上同时提升了AIGI检测准确率与解释质量。
Comments Accepted by ACMMM 2026
MedPixel:用于医学推理与分割的统一像素-语言模型
专题命中 视觉定位与Grounding :vision-language model(abstract);visual reasoning(abstract);grounding(abstract);分类 cs.CV、cs.AI
AI总结 该研究提出统一医学像素-语言模型MedPixel,引入44万样本的MedPLG-440K数据集,通过联合多任务微调与像素级偏好优化训练,支持多类医学任务,性能优异且具备零样本迁移与鲁棒性。
面向胸部X射线报告生成的正例-无标签偏好优化
机构 * Columbia University(哥伦比亚大学) ; Stanford University(斯坦福大学) ; Emory University(埃默里大学)
专题命中 视觉定位与Grounding :VLM(summary_cn);vision-language model(abstract);分类 cs.CV、cs.LG
AI总结 该研究针对放射报告生成VLM的遗漏噪声问题,提出PU-DPO框架,将未提及项视为无标签,通过对比对优化,提升病理检测率与隐藏正例恢复能力,增强对遗漏噪声的鲁棒性。
面向社交媒体中统一的多模态虚假信息检测:基准数据集与基线模型
机构 * School of Computer Science and Information Engineering, Hefei University of Technology(计算机科学与信息工程学院,合肥工业大学) ; School of Computer Science and Technology, Northwestern Polytechnical University(计算机科学与技术学院,西北工业大学)
专题命中 视觉定位与Grounding :VLM(abstract,abstract_cn);vision-language model(abstract);分类 cs.CV、cs.AI
AI总结 该研究构建了含9.8万样本的OmniFake基准数据集,提出UMFDet框架,实现对人工与AI生成两类多模态虚假内容的统一检测,性能优于专用基线。
RefBench-PRO:面向感知与推理的指称表达理解基准
机构 * National Key Laboratory of Human-Machine Hybrid Augmented Intelligence(国家人机混合增强智能重点实验室) ; National Engineering Research Center for Visual Information and Applications(国家视觉信息与应用工程技术研究中心) ; Institute of Artificial Intelligence and Robotics(人工智能与机器人研究院) ; Xi’an Jiaotong University(西安交通大学) ; Institute of Artificial Intelligence (TeleAI)(人工智能研究院(TeleAI)) ; China Telecom(中国电信) ; Shanghai Jiao Tong University(上海交通大学) ; University of Science and Technology Beijing(北京科技大学)
专题命中 视觉定位与Grounding :MLLM(abstract,abstract_cn);grounding(abstract);分类 cs.CV、cs.AI
AI总结 RefBench-PRO提出一个面向感知与推理的指称表达理解基准,通过分解指称表达为感知和推理两个维度,设计六个逐步递增的挑战任务,并引入Ref-R1强化学习方案提升定位精度,实现对多模态大语言模型的可解释评估。
现代视觉语言模型中用于强像素级图像篡改检测的简单域泛化
机构 * Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学) ; University College London(伦敦大学学院) ; Weizmann Institute of Science(魏茨曼科学研究所)
专题命中 视觉定位与Grounding :VLM(abstract,abstract_cn);vision-language model(abstract);分类 cs.CV、cs.AI
AI总结 研究现代视觉语言模型中像素级图像篡改检测的域泛化,提出基于平衡小批量采样和后期注入策略的简单训练框架,大幅提升平均gIoU和cIoU,增强了篡改定位和分布外鲁棒性。
Comments Our code is available at https://github.com/VILA-Lab/PIXAR-DG
通过概念引导实现鲁棒的上下文分割
机构 * Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University(厦门大学多媒体可信感知与高效计算教育部重点实验室)
专题命中 视觉定位与Grounding :MLLM(abstract,abstract_cn);grounding(abstract);分类 cs.CV、cs.AI
AI总结 提出概念引导的上下文分割(CG-ICS),通过提取参考图像的高层语义概念而非仅依赖低层视觉匹配,结合文本概念与视觉示例,显著提升分割准确性和鲁棒性。
Comments ECCV 2026
超越视觉取证:审计多模态鲁棒性用于合成医学图像检测
机构 * University of Notre Dame(圣母大学) ; IBM Research(IBM研究院) ; Boston Children’s Hospital(波士顿儿童医院) ; Harvard Medical School(哈佛医学院) ; Optum AI, UnitedHealth Group(Optum AI, 联合健康集团)
专题命中 视觉定位与Grounding :VLM(abstract,abstract_cn);vision-language model(abstract);分类 cs.CV、cs.AI
AI总结 针对合成医学图像检测中多模态鲁棒性不足的问题,提出图像-记录配对基准,揭示视觉语言模型因过度依赖记录上下文而导致的预测偏差。
Comments Accepted at MICCAI 2026. Version 2 is a substantial journal extension of the MICCAI 2026 conference version, with additional provenance perturbations, paired statistical analysis, extended SAVC mitigation experiments, and broader deployment discussion. 19 pages, 3 figures, 2 tables
眼见为实:基于视觉锚点的提示重写对齐用于文本到图像生成
机构 * Peking University(北京大学) ; Tencent(腾讯) ; Dalian University of Technology(大连理工大学) ; Nanyang Technological University(南洋理工大学) ; University of Cambridge(剑桥大学) ; Zhejiang University(浙江大学)
专题命中 视觉定位与Grounding :MLLM(abstract,abstract_cn);grounding(abstract);分类 cs.CV、cs.AI
AI总结 提出FaithRewriter框架,利用多模态大模型生成中间视觉线索,结合大语言模型生成视觉锚定的增强提示,再蒸馏至小模型,以缩小用户意图与生成图像之间的差距。
EgoSim:用于具身交互生成的视角世界模拟器
机构 * Shanghai Jiao Tong University(上海交通大学) ; Shanghai AI Laboratory(上海人工智能实验室) ; The University of Hong Kong(香港大学)
专题命中 视觉定位与Grounding :grounding(summary_cn,abstract);分类 cs.CV、cs.AI
AI总结 EgoSim通过建模可更新的世界状态,解决现有视角模拟器在3D grounding和动态更新上的不足,生成空间一致的交互视频并支持跨具身迁移。
Comments Project Page: egosimulator.github.io
将神经符号程序蒸馏到3D多模态大语言模型中
机构 * University of California, Berkeley(加州大学伯克利分校)
专题命中 视觉定位与Grounding :MLLM(abstract,abstract_cn);grounding(abstract);分类 cs.CV、cs.AI
AI总结 提出APEIRIA,通过三阶段课程学习将符号推理模式蒸馏到3D多模态大语言模型中,实现透明推理与开放词汇空间推理的统一。
Comments To appear in ICML 2026
感知、判断与进化:基于事后洞察的自优化取证智能体用于AI生成图像检测
机构 * Zhejiang University(浙江大学) ; Alibaba Group(阿里巴巴集团)
专题命中 视觉定位与Grounding :MLLM(abstract,abstract_cn);multimodal large language model(abstract);分类 cs.CV、cs.AI
AI总结 提出ForeAgent框架,采用感知-判断架构融合多视图线索,并引入事后洞察驱动的自优化策略,通过采样-反思-进化范式持续提升检测能力,在多个基准上达到最优性能。
Comments 10 pages
LandslideAgent与多模态LandslideBench:一种面向自主滑坡识别与分析的领域规则增强型智能体
机构 * Central South University(中南大学)
专题命中 视觉定位与Grounding :VLM(abstract,abstract_cn);vision-language model(abstract);分类 cs.CV、cs.AI
AI总结 提出指令驱动智能体框架,包含多模态数据集LandslideBench、滑坡专用视觉语言模型LandslideVLM及领域规则增强智能体LandslideAgent,实现自主滑坡识别与分析。
MemoVAD: 边缘计算场景下基于动态语义记忆的资源高效视频异常检测
机构 * Institute of Artificial Intelligence and Future Networks, Beijing Normal University(北京师范大学人工智能与未来网络研究院) ; School of Computing and Artificial Intelligence, Southwest Jiaotong University(西南交通大学计算机与人工智能学院) ; Engineering Research Center of Cloud-Edge Intelligent Collaboration on Big Data, Ministry of Education, Beijing Normal University(北京师范大学大数据云边智能协同教育部工程研究中心)
专题命中 视觉定位与Grounding :VLM(abstract,abstract_cn);vision-language model(abstract);分类 cs.CV、cs.AI
AI总结 提出MemoVAD边缘-云协同框架,通过不确定性感知门控策略选择性调用云端视觉语言模型,并设计动态语义记忆缓存原型,在降低通信开销的同时提升视频异常检测性能。
Comments Accepted by IJCAI2026