Symbol Grounding via Chaining of Morphisms
专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI
AI 大模型
视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。
专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI
专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI
Comments 21 pages, 7 figures, preliminary version appeared at GTTV'15
专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI
Comments 11 pages, 3 figures
专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI
专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI
Journal ref Journal Of Artificial Intelligence Research, Volume 38, pages 223-269, 2010
专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI
Comments Appears in the Proceedings of the 19th International Conference on Applications of Declarative Programming and Knowledge Management (INAP 2011)
引用溯源:通过法律引用图检测和减少LLM引用幻觉
机构 * LEX AI LLC
专题命中 视觉定位与Grounding :grounding(title,abstract)
AI总结 提出引用溯源(CG)指标,利用乌克兰法院判决的引用图(1.008亿判决,5.02亿边)检测LLM法律引用幻觉,并通过CG-DPO方法(基于真实判决构建偏好对)减少幻觉,在100个法律查询上CG为0.791-0.873,幻觉率13-21%。
Comments 21 pages, 4 figures, 5 tables. Substantially revised: title, framing and several v1 results changed. Adds a coverage sweep and a separability analysis; corrects the DPO configuration, the density-accuracy correlation and the qualitative examples. Code and data: https://huggingface.co/datasets/overthelex/citation-grounding-eval
专题命中 视觉定位与Grounding :grounding(title,abstract)
Comments This is the second part of the general framework for a constructive type theory presented in the paper Functionals and hardware arXiv:1501.03043. The version is final. The research on the grounding of Mathematics is continued in the paper {\em Asymptotic combinatorial constructions of Geometries} available at arXiv:1904.05173
ExtrinSplat:解耦几何与语义以实现3D高斯散射中的开放词汇理解
机构 * Guangdong Provincial Key Laboratory of Ultra High Definition Immersive Media Technology, Shenzhen Graduate School, Peking University(广东省超高清沉浸式媒体技术重点实验室,北京大学深圳研究生院) ; School of Computer Science and Technology, Tianjin University(天津大学计算机科学与技术学院) ; Guangdong Bohua UHD Innovation Center Co., Ltd.(广东博华超高清创新中心有限公司)
专题命中 视觉定位与Grounding :VLM(abstract,abstract_cn);vision-language model(abstract);分类 cs.CV、cs.AI
AI总结 ExtrinSplat通过解耦几何与语义,利用视觉语言模型生成轻量文本假设,提升3D高斯散射中开放词汇物体选择和语义分割的性能和效率。
Comments Accepted to CVPR 2026
Defake-o3:从推测性理由到可验证证据的可解释AIGI检测
机构 * Shanghai Jiao Tong University(上海交通大学) ; Ant Group(蚂蚁集团)
专题命中 视觉定位与Grounding :MLLM(abstract,abstract_cn);grounding(abstract);分类 cs.CV、cs.AI
AI总结 Defake-o3是结合交互式视觉搜索与证据验证器的可解释AIGI检测器,构建了GroundFake数据集和FakeFrontier基准,在多基准上同时提升了AIGI检测准确率与解释质量。
Comments Accepted by ACMMM 2026
MedPixel:用于医学推理与分割的统一像素-语言模型
专题命中 视觉定位与Grounding :vision-language model(abstract);visual reasoning(abstract);grounding(abstract);分类 cs.CV、cs.AI
AI总结 该研究提出统一医学像素-语言模型MedPixel,引入44万样本的MedPLG-440K数据集,通过联合多任务微调与像素级偏好优化训练,支持多类医学任务,性能优异且具备零样本迁移与鲁棒性。
面向胸部X射线报告生成的正例-无标签偏好优化
机构 * Columbia University(哥伦比亚大学) ; Stanford University(斯坦福大学) ; Emory University(埃默里大学)
专题命中 视觉定位与Grounding :VLM(summary_cn);vision-language model(abstract);分类 cs.CV、cs.LG
AI总结 该研究针对放射报告生成VLM的遗漏噪声问题,提出PU-DPO框架,将未提及项视为无标签,通过对比对优化,提升病理检测率与隐藏正例恢复能力,增强对遗漏噪声的鲁棒性。
面向社交媒体中统一的多模态虚假信息检测:基准数据集与基线模型
机构 * School of Computer Science and Information Engineering, Hefei University of Technology(计算机科学与信息工程学院,合肥工业大学) ; School of Computer Science and Technology, Northwestern Polytechnical University(计算机科学与技术学院,西北工业大学)
专题命中 视觉定位与Grounding :VLM(abstract,abstract_cn);vision-language model(abstract);分类 cs.CV、cs.AI
AI总结 该研究构建了含9.8万样本的OmniFake基准数据集,提出UMFDet框架,实现对人工与AI生成两类多模态虚假内容的统一检测,性能优于专用基线。
RefBench-PRO:面向感知与推理的指称表达理解基准
机构 * National Key Laboratory of Human-Machine Hybrid Augmented Intelligence(国家人机混合增强智能重点实验室) ; National Engineering Research Center for Visual Information and Applications(国家视觉信息与应用工程技术研究中心) ; Institute of Artificial Intelligence and Robotics(人工智能与机器人研究院) ; Xi’an Jiaotong University(西安交通大学) ; Institute of Artificial Intelligence (TeleAI)(人工智能研究院(TeleAI)) ; China Telecom(中国电信) ; Shanghai Jiao Tong University(上海交通大学) ; University of Science and Technology Beijing(北京科技大学)
专题命中 视觉定位与Grounding :MLLM(abstract,abstract_cn);grounding(abstract);分类 cs.CV、cs.AI
AI总结 RefBench-PRO提出一个面向感知与推理的指称表达理解基准,通过分解指称表达为感知和推理两个维度,设计六个逐步递增的挑战任务,并引入Ref-R1强化学习方案提升定位精度,实现对多模态大语言模型的可解释评估。
现代视觉语言模型中用于强像素级图像篡改检测的简单域泛化
机构 * Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学) ; University College London(伦敦大学学院) ; Weizmann Institute of Science(魏茨曼科学研究所)
专题命中 视觉定位与Grounding :VLM(abstract,abstract_cn);vision-language model(abstract);分类 cs.CV、cs.AI
AI总结 研究现代视觉语言模型中像素级图像篡改检测的域泛化,提出基于平衡小批量采样和后期注入策略的简单训练框架,大幅提升平均gIoU和cIoU,增强了篡改定位和分布外鲁棒性。
Comments Our code is available at https://github.com/VILA-Lab/PIXAR-DG
通过概念引导实现鲁棒的上下文分割
机构 * Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University(厦门大学多媒体可信感知与高效计算教育部重点实验室)
专题命中 视觉定位与Grounding :MLLM(abstract,abstract_cn);grounding(abstract);分类 cs.CV、cs.AI
AI总结 提出概念引导的上下文分割(CG-ICS),通过提取参考图像的高层语义概念而非仅依赖低层视觉匹配,结合文本概念与视觉示例,显著提升分割准确性和鲁棒性。
Comments ECCV 2026
超越视觉取证:审计多模态鲁棒性用于合成医学图像检测
机构 * University of Notre Dame(圣母大学) ; IBM Research(IBM研究院) ; Boston Children’s Hospital(波士顿儿童医院) ; Harvard Medical School(哈佛医学院) ; Optum AI, UnitedHealth Group(Optum AI, 联合健康集团)
专题命中 视觉定位与Grounding :VLM(abstract,abstract_cn);vision-language model(abstract);分类 cs.CV、cs.AI
AI总结 针对合成医学图像检测中多模态鲁棒性不足的问题,提出图像-记录配对基准,揭示视觉语言模型因过度依赖记录上下文而导致的预测偏差。
Comments Accepted at MICCAI 2026. Version 2 is a substantial journal extension of the MICCAI 2026 conference version, with additional provenance perturbations, paired statistical analysis, extended SAVC mitigation experiments, and broader deployment discussion. 19 pages, 3 figures, 2 tables
眼见为实:基于视觉锚点的提示重写对齐用于文本到图像生成
机构 * Peking University(北京大学) ; Tencent(腾讯) ; Dalian University of Technology(大连理工大学) ; Nanyang Technological University(南洋理工大学) ; University of Cambridge(剑桥大学) ; Zhejiang University(浙江大学)
专题命中 视觉定位与Grounding :MLLM(abstract,abstract_cn);grounding(abstract);分类 cs.CV、cs.AI
AI总结 提出FaithRewriter框架,利用多模态大模型生成中间视觉线索,结合大语言模型生成视觉锚定的增强提示,再蒸馏至小模型,以缩小用户意图与生成图像之间的差距。
EgoSim:用于具身交互生成的视角世界模拟器
机构 * Shanghai Jiao Tong University(上海交通大学) ; Shanghai AI Laboratory(上海人工智能实验室) ; The University of Hong Kong(香港大学)
专题命中 视觉定位与Grounding :grounding(summary_cn,abstract);分类 cs.CV、cs.AI
AI总结 EgoSim通过建模可更新的世界状态,解决现有视角模拟器在3D grounding和动态更新上的不足,生成空间一致的交互视频并支持跨具身迁移。
Comments Project Page: egosimulator.github.io
将神经符号程序蒸馏到3D多模态大语言模型中
机构 * University of California, Berkeley(加州大学伯克利分校)
专题命中 视觉定位与Grounding :MLLM(abstract,abstract_cn);grounding(abstract);分类 cs.CV、cs.AI
AI总结 提出APEIRIA,通过三阶段课程学习将符号推理模式蒸馏到3D多模态大语言模型中,实现透明推理与开放词汇空间推理的统一。
Comments To appear in ICML 2026
感知、判断与进化:基于事后洞察的自优化取证智能体用于AI生成图像检测
机构 * Zhejiang University(浙江大学) ; Alibaba Group(阿里巴巴集团)
专题命中 视觉定位与Grounding :MLLM(abstract,abstract_cn);multimodal large language model(abstract);分类 cs.CV、cs.AI
AI总结 提出ForeAgent框架,采用感知-判断架构融合多视图线索,并引入事后洞察驱动的自优化策略,通过采样-反思-进化范式持续提升检测能力,在多个基准上达到最优性能。
Comments 10 pages
LandslideAgent与多模态LandslideBench:一种面向自主滑坡识别与分析的领域规则增强型智能体
机构 * Central South University(中南大学)
专题命中 视觉定位与Grounding :VLM(abstract,abstract_cn);vision-language model(abstract);分类 cs.CV、cs.AI
AI总结 提出指令驱动智能体框架,包含多模态数据集LandslideBench、滑坡专用视觉语言模型LandslideVLM及领域规则增强智能体LandslideAgent,实现自主滑坡识别与分析。
MemoVAD: 边缘计算场景下基于动态语义记忆的资源高效视频异常检测
机构 * Institute of Artificial Intelligence and Future Networks, Beijing Normal University(北京师范大学人工智能与未来网络研究院) ; School of Computing and Artificial Intelligence, Southwest Jiaotong University(西南交通大学计算机与人工智能学院) ; Engineering Research Center of Cloud-Edge Intelligent Collaboration on Big Data, Ministry of Education, Beijing Normal University(北京师范大学大数据云边智能协同教育部工程研究中心)
专题命中 视觉定位与Grounding :VLM(abstract,abstract_cn);vision-language model(abstract);分类 cs.CV、cs.AI
AI总结 提出MemoVAD边缘-云协同框架,通过不确定性感知门控策略选择性调用云端视觉语言模型,并设计动态语义记忆缓存原型,在降低通信开销的同时提升视频异常检测性能。
Comments Accepted by IJCAI2026
CoCoVideo: 基于商业模型的高质量对比基准用于AI生成视频检测
机构 * School of Informatics, Xiamen University(厦门大学信息学院) ; China Academy of Information and Communications Technology(中国信息通信技术研究院) ; AI Transcend Pte. Ltd.(AI Transcend有限公司)
专题命中 视觉定位与Grounding :MLLM(abstract,abstract_cn);multimodal large language model(abstract);分类 cs.CV、cs.AI
AI总结 针对现有数据集依赖低质量开源模型且商业样本带水印的问题,提出包含13个商业生成器的CoCoVideo-26K对比数据集,并设计结合对比学习与置信门控多模态大语言模型的CoCoDetect检测框架,实现高保真AI生成视频的鲁棒检测。
Comments Accepected by CVPR 2026
生成报告还是重复模板?测量和缓解三维CT报告生成中的模板崩溃
机构 * Technical University of Munich (TUM)(慕尼黑技术大学) ; TUM Hospital(TUM医院) ; Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心)
专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract_cn);grounding(abstract);分类 cs.CV、cs.AI
AI总结 针对三维CT报告生成中模型输出多样性低、病理检测能力差的模板崩溃问题,提出解耦框架CLarGen,通过分离临床检测与语言合成,显著提升临床准确性并保持报告流畅性。
PInVerify:面向主动实例验证的离线具身基准
机构 * University of Trento(特伦托大学)
专题命中 视觉定位与Grounding :MLLM(abstract,abstract_cn);multimodal large language model(abstract);分类 cs.CV、cs.AI
AI总结 提出主动实例验证任务,构建离线具身基准PInVerify,通过多视角导航和细粒度属性匹配评估具身智能体,并基于多模态大语言模型建立基线。
Comments Accepted as a poster at the Foundation Models Meet Embodied Agents (FMEA) Workshop, CVPR 2026. 44 pages including appendix. Code: https://github.com/Avalon-S/PInVerify
EchoPilot: 通过尺度空间语义提示和可靠性门控记忆实现无训练超声视频分割
机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) ; Third Affiliated Hospital of Sun Yat-Sen University(中山大学第三附属医院) ; Hong Kong Metropolitan University(香港 Metropolitan 大学)
专题命中 视觉定位与Grounding :VLM(abstract,abstract_cn);vision-language model(abstract);分类 cs.CV、cs.AI
AI总结 提出EchoPilot,一种无需训练、仅需单点点击和类别名称的超声视频分割框架,通过尺度空间语义提示解决初始化歧义,并引入可靠性门控记忆减少传播漂移,在多个数据集上达到最优性能。
Comments Early accepted to MICCAI 2026. Project page: https://keeplearning-again.github.io/EchoPilot/
不看而见:视觉-语言基准测试真的测试视觉吗?
机构 * University of Chicago(芝加哥大学) ; Stony Brook University(石溪大学) ; Toyota Technological Institute at Chicago(芝加哥丰田技术研究所)
专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract_cn);grounding(abstract);分类 cs.CV、cs.AI
AI总结 通过系统性实验发现,视觉-语言模型对细粒度视觉证据的依赖程度低于预期,表明当前基准测试不足以可靠评估其视觉基础能力。
Comments Accepted to GRAIL-V: Grounded Retrieval and Agentic Intelligence for Vision-Language, CVPR 2026 Workshop. accepted version
在情绪树中导航:用于多模态情绪识别的分层双曲RAG
机构 * Great Bay University(广东东莞大亚湾大学) ; Tencent Youtu Lab(腾讯优图实验室)
专题命中 视觉定位与Grounding :grounding(abstract,abstract_cn);multimodal large language model(abstract);分类 cs.CV、cs.LG
AI总结 本文提出HyperEmo-RAG,一种利用结构化情绪知识库的检索增强生成框架,通过双曲空间嵌入和证据图构建来提升多模态情绪识别的性能。
AnomalyClaw:通过工具引导的反驳实现通用视觉异常检测代理
机构 * Department of Computer Science and Engineering, Southern University of Science and Technology (SUSTech), Shenzhen, China(南方科技大学计算机科学与工程系,深圳,中国) ; School of EEE, Nanyang Technological University (NTU), Singapore(南洋理工大学电子工程学院,新加坡) ; CFAR, Agency for Science, Technology and Research (A*STAR), Singapore(科技研究局(A*STAR)的CFAR,新加坡)
专题命中 视觉定位与Grounding :VLM(abstract,abstract_cn);vision-language model(abstract);分类 cs.CV、cs.AI
AI总结 本文提出AnomalyClaw,一种无需训练的视觉异常检测代理,通过多轮反驳过程提升异常判断。在CrossDomainVAD-12基准上,AnomalyClaw在多个模型上均取得显著提升,且引入自演化扩展提升模型性能。
Comments We release the agent, the benchmark, and the analysis artifacts at https://github.com/jam-cc/AnomalyClaw