MouSi: Poly-Visual-Expert Vision-Language Models
专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV、cs.AI、cs.LG
AI 大模型
视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。
专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV、cs.AI、cs.LG
专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV、cs.AI、cs.LG
Comments In EMNLP 2023 main conference proceedings (to appear)
专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV、cs.AI、cs.LG
专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV、cs.AI、cs.LG
Comments Project website: https://chat-with-nerf.github.io/
专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV、cs.AI、cs.LG
Comments Extended version for ICLRW-2020 (BeTR-RL) paper
专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV、cs.AI、cs.LG
专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV、cs.AI、cs.LG
专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV、cs.AI、cs.LG
Comments Conference on Robot Learning (CoRL) 2021
专题命中 视觉定位与Grounding :grounding(title,abstract);visual question answering(abstract)
Comments Accepted at LANTERN2021
专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV、cs.AI、cs.LG
Comments Accepted at the International Symposium on Experimental Robotics (ISER) 2020. Videos at http://speechrobot.cs.uni-freiburg.de
专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV、cs.AI、cs.LG
Comments 4th Conference on Robot Learning (CoRL 2020), Cambridge MA, USA
专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV、cs.AI、cs.LG
Comments Accepted at WACV 2019. Also at NeurIPS 2017 workshop on Visually-Grounded Interaction and Language (ViGIL)
合成对象组合用于检测、分割和接地任务中的可扩展和准确学习
机构 * University of Washington(华盛顿大学) ; Allen Institute for Artificial Intelligence(人工智能研究院)
专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV、cs.AI
AI总结 合成对象组合(SOC)通过新的以对象为中心的组合策略,实现了准确且可扩展的数据合成流程,提升了检测、分割和接地任务的性能,尤其在低数据和封闭词汇场景中表现突出。
Comments Project website: https://github.com/weikaih04/Synthetic-Detection-Segmentation-Grounding-Data
专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV、cs.LG
Comments The second and the third authors contributed equally to this paper, and are listed in alphabetical order. Project Page: http://visual-physics-grounding.csail.mit.edu/
SEED: 用于可解释文本伪造检测的简单ViT与演化框架
机构 * State Key Laboratory of Internet of Things for Smart City, Department of Computer and Information Science, University of Macau(澳门大学计算机与信息科学系智慧城市物联网国家重点实验室) ; School of Information and Software Engineering, University of Electronic Science and Technology of China(电子科技大学信息与软件工程学院)
专题命中 视觉定位与Grounding :MLLM(summary_cn,abstract);分类 cs.CV
AI总结 提出SEED系统,结合相似性引导数据增强、单一ViT联合检测与定位、以及基于MLLM的演化报告生成,在ACM MM 2026文本伪造挑战赛中获得第三名。
EditFlow3D:保留轨迹的3D资产自动化局部编辑
专题命中 视觉定位与Grounding :VLM(summary_cn,abstract);分类 cs.CV
AI总结 EditFlow3D是一种无需训练的3D局部编辑框架,通过VLM驱动工作流生成掩码与视觉引导,结合掩码引导差分流和轨迹保留技术,在新基准EditFlow-Bench上验证了其编辑准确性与非目标区域保留能力优于现有方法。
LDU-Bench:面向不同布局电路背景下光刻缺陷理解的多模态大语言模型评估基准
专题命中 视觉定位与Grounding :multimodal large language model(abstract,abstract_cn);MLLM(summary_cn);分类 cs.CV
AI总结 本文提出LDU-Bench这一多模态基准,将光刻审查工作流程分解为四项任务,评估发现现有多模态大语言模型的缺陷分类能力无法稳定迁移至下游审查阶段,为工业MLLM评估提供了统一平台。
Comments 12 pages, 3 figures, and 5 tables, including appendices
OVEarth-Bench:面向开放词汇地球观测的类别广度与查询多样性评估
专题命中 视觉定位与Grounding :MLLM(summary_cn,abstract);分类 cs.CV
AI总结 本文提出OVEarth-Bench基准,从类别广度与查询多样性两方面扩展开放词汇地球观测评估,经实验发现MLLM类方法性能最优,为该领域未来研究提供指导。
TongueReenact:几何锚定的人脸重演舌部合成
机构 * University of Technology Sydney(悉尼科技大学) ; Shandong University of Science and Technology(山东科技大学)
专题命中 视觉定位与Grounding :VLM(summary_cn,abstract);分类 cs.CV
AI总结 本文提出首个跨身份舌部动态迁移框架,结合 foundation model 辅助的自举流水线与空间约束潜在掩码扩散模型,在舌部指标上较基线提升超两倍,且 VLM 评估证实其感知优势。
证据见证组合:针对闭集多模态答案的单预填充风险检测
专题命中 视觉定位与Grounding :MLLM(summary_cn,abstract_cn);multimodal large language model(abstract);分类 cs.CV
AI总结 该研究提出WEP方法,利用MLLM的白盒预填充路径,无需额外操作即可检测闭集视觉答案的推理风险,在3个MLLM和4个基准上提升了平均错误AP。
Comments 22 pages, 6 figures; includes supplementary material. Code: https://github.com/SouthWinter/WEP
RAD:面向机器人观测的真实异常检测数据集与基准
机构 * Massachusetts Institute of Technology(麻省理工学院) ; Peking University(北京大学) ; Carnegie Mellon University(卡内基梅隆大学) ; Great Bay University(大湾大学) ; Harvard University(哈佛大学) ; Tsinghua University(清华大学) ; Nanjing University(南京大学) ; Deakin University(德克萨斯大学)
专题命中 视觉定位与Grounding :VLM(summary_cn,abstract_cn);vision-language model(abstract);分类 cs.CV
AI总结 提出RAD数据集,包含13类日常物体和4种缺陷,从60多个机器人视角在非受控光照下采集,用于评估2D特征、3D重建和视觉语言模型在姿态无关的异常检测中的表现,发现2D方法优于3D和VLM方法。
循环代码取证:面向图像伪造检测的代理工具使用
机构 * University of Science and Technology of China(中国科学技术大学) ; Shanghai Innovation Institute(上海创新研究院) ; Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
专题命中 视觉定位与Grounding :MLLM(summary_cn,abstract_cn);multimodal large language model(abstract);分类 cs.AI
AI总结 本文提出ForenAgent框架,通过多轮交互使MLLM自主生成并迭代优化Python工具,提升图像伪造检测的灵活性和可解释性,构建FABench数据集进行系统训练和评估。
Comments 18 pages, 7 figures
VG3S:基于视觉几何的高斯散射用于语义占用预测
机构 * Department of Electronic and Computer Engineering, The Hong Kong University of Science and Technology(电子与计算机工程系,香港科学与技术大学)
专题命中 视觉定位与Grounding :grounding(summary_cn,abstract);分类 cs.CV
AI总结 VG3S 通过引入视觉基础模型的几何 grounding 能力,提升语义占用预测的准确性与泛化能力。
Comments Accepted by IROS 2026
Open-KNEAD:通过智能分解实现基于知识的营养估计
机构 * Purdue University(普渡大学)
专题命中 视觉定位与Grounding :MLLM(abstract,abstract_cn);grounding(abstract);multimodal large language model(abstract);分类 cs.CV
AI总结 研究探讨多模态大语言模型用于膳食评估时检索增强基础方法是否仍有价值。提出无需训练、可本地部署的Open-KNEAD框架,通过营养感知检索关联食物项与FNDDS代码,提高份量估计,还能恢复非美国烹饪风格估计偏差,优势显著并开源相关框架与知识库。
Comments 10 pages main paper, 5 pages supplementary
EvoGuard: 一种基于代理强化学习的可扩展框架,用于实际和演化的AI生成图像检测
机构 * The University of Tokyo(东京大学) ; National Institute of Informatics(国家信息研究所)
专题命中 视觉定位与Grounding :MLLM(summary_cn,abstract_cn);multimodal large language model(abstract);分类 cs.CV
AI总结 本文提出EvoGuard框架,利用多模态大语言模型和非MLLM检测器,通过能力感知的动态编排机制实现高效AIGI检测,优化后无需细粒度标注,实验表明其在准确性和可扩展性上均达最优。
Comments Template changed
停止猜测何时停止测试:用足够的数据进行高效模型评估
机构 * IBM Research(IBM研究院)
专题命中 视觉定位与Grounding :VLM(summary_cn,abstract);分类 cs.LG
AI总结 研究指出固定大小基准用于模型评估效率低,提出自适应评估框架,结合序贯测试统计范式与定制停止标准,在Open VLM Leaderboard上展示能自适应管理效率与可靠性权衡,降低计算成本并保持统计显著性。
花瓶博物馆:古希腊陶器数字智能博物馆
机构 * School of Computer Science and Technology, Beijing Jiaotong University(北京交通大学计算机科学与技术学院) ; School of Computer Science, Peking University(北京大学计算机科学学院) ; Huazhong University of Science and Technology(华中科技大学) ; Department of Computer Science and Information Technology, La Trobe University(拉筹伯大学计算机科学与信息技术系) ; UCAS-Terminus AI Lab, University of Chinese Academy of Sciences(中国科学院大学UCAS-Terminus人工智能实验室)
专题命中 视觉定位与Grounding :VLM(abstract,abstract_cn);vision-language model(abstract);grounding(abstract);分类 cs.CV
AI总结 针对古希腊陶器文化遗产领域中视觉语言模型面临的挑战,提出轻量级模块化多模态智能体框架花瓶博物馆,结合虚拟博物馆与花瓶智能体,经多模态感知等操作,提高引用有效性,减少幻觉并产生更中立答案。
Comments Code: https://github.com/AIGeeksGroup/VaseMuseum. Website: https://aigeeksgroup.github.io/VaseMuseum
边建图边思考:用于增量式3D场景图的异步视觉-语言智能体
机构 * University of Stuttgart(斯图加特大学) ; Graz University of Technology(格拉茨技术大学) ; IMPRS-IS(国际马克斯·普朗克智能系统研究学院)
专题命中 视觉定位与Grounding :VLM(abstract,abstract_cn);vision-language model(abstract);grounding(abstract);分类 cs.CV
AI总结 提出异步架构,轻量在线建图与重型语义精化并发运行,通过语义闭环、属性标注和空间关系推导,实现探索时可查询且语义渐增的3D场景图,在多个基准上超越现有方法。
Comments Accepted to ECCV 2026. Project page: https://denizbickici.github.io/thinkgraphs/
仅在需要时看:用于缓解LVLM中幻觉的上下文感知注意力干预
机构 * University of Chinese Academy of Sciences(中国科学院大学) ; University of Amsterdam(阿姆斯特丹大学) ; United Imaging Healthcare Co., Ltd.(联合影像医疗科技股份有限公司)
专题命中 视觉定位与Grounding :grounding(summary_cn,abstract);vision-language model(abstract);分类 cs.CV
AI总结 提出训练自由的上下文感知注意力干预(CAI),通过两轴选择性(看哪里和何时干预)在解码时针对性增强视觉 grounding,有效缓解物体幻觉并保持语言流畅性。
Journal ref ECCV 2026
MedP-CLIP:具有区域感知提示集成的医学CLIP
机构 * Sichuan University(四川大学) ; Xinjiang University(新疆大学) ; Alibaba Group(阿里巴巴集团) ; Fuzhou University(福州大学) ; Shanghai AI Laboratory(上海人工智能实验室) ; Southwest Jiaotong University(西南交通大学)
专题命中 视觉定位与Grounding :VLM(abstract,abstract_cn);vision-language model(abstract);multimodal large language model(abstract);分类 cs.CV
AI总结 提出MedP-CLIP,一种区域感知医学视觉语言模型,通过特征级区域提示集成机制,支持多种提示形式,在保持全局上下文的同时聚焦局部区域,在零样本识别、交互式分割等任务中显著优于基线方法。
Comments Accepted by Medical Image Analysis (MedIA)
Journal ref Medical Image Analysis, 113 (2026), 104193