arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2026-01-27 至 2026-01-27 共收录 81 信号源:cs.CV, cs.AI, cs.LG

1. 视觉问答 11 篇

2601.18240 2026-01-27 cs.CV 87%

V-Loop: Visual Logical Loop Verification for Hallucination Detection in Medical Visual Question Answering

V-Loop:用于医学视觉问答中幻觉检测的视觉逻辑循环验证

Mengyuan Jin, Zehui Liao, Yong Xia

机构 * Northwestern Polytechnical University(西北工业大学)

专题命中 视觉问答 :visual question answering(title,abstract);grounding(abstract);multimodal large language model(abstract);MLLM(abstract)

AI总结 V-Loop通过双向推理和视觉逻辑循环验证,提升医学视觉问答中幻觉检测的准确性和效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17383 2026-01-27 cs.CV cs.AI 84%

Physical Prompt Injection Attacks on Large Vision-Language Models

针对大视觉-语言模型的物理提示注入攻击

Chen Ling, Kai Hu, Hangcheng Liu, Xingshuo Han, Tianwei Zhang, Changhai Ou

机构 * School of Cyber Science and Engineering, Wuhan University(武汉大学计算机科学与工程学院) College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算与数据科学学院) College of Computer Science and Technology, Nanjing University of Aeronautics and Astronautics(南京航空航天大学计算机科学与技术学院)

专题命中 视觉问答 :vision-language model(title,abstract);visual question answering(abstract);分类 cs.CV、cs.AI

AI总结 本研究提出了一种无需访问模型或输入的物理提示注入攻击方法,通过物理物体注入恶意指令,成功攻击多种大视觉-语言模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18989 2026-01-27 cs.CV 83%

Ask Me Again Differently: GRAS for Measuring Bias in Vision Language Models on Gender, Race, Age, and Skin Tone

再次提问:GRAS用于测量视觉语言模型在性别、种族、年龄和肤色上的偏见

Shaivi Malik, Hasnat Md Abdullah, Sriparna Saha, Amit Sheth

机构 * Guru Gobind Singh Indraprastha University(古鲁·戈宾德·辛格印度教普拉斯塔大学) AI Institute, University of South Carolina(南卡罗来纳大学人工智能研究所) University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Indian Institute of Technology Patna, India(印度理工学院帕纳布分校)

专题命中 视觉问答 :vision language model(title,abstract);visual question answering(abstract);分类 cs.CV

AI总结 GRAS用于评估视觉语言模型在性别、种族、年龄和肤色上的偏见,通过基准测试揭示了模型的偏见水平,并提出了可解释的偏见评分方法。

Comments Accepted to the Findings of EACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.04192 2026-01-27 cs.CV 77%

PixFoundation: Are We Heading in the Right Direction with Pixel-level Vision Foundation Models?

PixFoundation: 我们在像素级视觉基础模型的方向正确吗?

Mennatullah Siam

专题命中 视觉问答 :visual question answering(abstract);grounding(abstract);MLLM(abstract);分类 cs.CV

AI总结 PixFoundation提出新的基准测试,揭示像素级基础模型在视觉问答中的不足,并提出可解释性工具分析接地机制。

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17089 2026-01-27 cs.CV 70%

GRASP: Guided Region-Aware Sparse Prompting for Adapting MLLMs to Remote Sensing

GRASP: 基于区域感知的稀疏提示引导方法用于适应遥感图像的多模态大语言模型

Qigan Sun, Chaoning Zhang, Jianwei Zhang, Xudong Wang, Jiehui Xie, Pengcheng Zheng, Haoyu Wang, Sungyoung Lee, Chi-lok Andy Tai, Yang Yang, Heng Tao Shen

机构 * School of Computing, Kyung Hee University(京畿大学计算机学院) School of Information and Software Engineering, University of Electronic Science and Technology of China(电子科技大学信息与软件工程学院) School of Computer Science and Engineering, University of Electronic Science and Technology of China(电子科技大学计算机科学与工程学院) College of Computer Science and Information Engineering, Harbin Normal University(哈尔滨师范大学计算机科学与信息工程学院) College of Professional and Continuing Education, The Hong Kong Polytechnic University(香港理工大学专业及继续教育学院) School of Computer Science and Technology, Tongji University(同济大学计算机科学与技术学院)

专题命中 视觉问答 :visual question answering(abstract);multimodal large language model(abstract);分类 cs.CV

AI总结 GRASP通过引导区域感知的稀疏提示方法,提升多模态大语言模型在遥感图像任务中的适应性与性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20298 2026-01-27 cs.CL cs.AI cs.CV 62%

MangaVQA and MangaLMM: A Benchmark and Specialized Model for Multimodal Manga Understanding

MangaVQA和MangaLMM:多模态漫画理解的基准和专用模型

Jeonghun Baek, Kazuki Egashira, Shota Onohara, Atsuyuki Miyai, Yuki Imajuku, Hikaru Ikuta, Kiyoharu Aizawa

机构 * The University of Tokyo(东京大学)

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.AI

AI总结 本文提出MangaVQA和MangaLMM,用于多模态漫画理解的基准和专用模型,通过视觉问答和文本识别任务提升漫画叙事理解能力。

Comments EACL 2026 Findings. Project page: https://manga109.github.io/MangaVQA_LMM/

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.18536 2026-01-27 cs.CV cs.CL cs.IR cs.LG 62%

FilterRAG: Zero-Shot Informed Retrieval-Augmented Generation to Mitigate Hallucinations in VQA

FilterRAG: 零样本引导式检索增强生成以缓解视觉问答中的幻觉

Nobin Sarwar

机构 * University of Maryland, Baltimore County(马里兰大学巴尔的摩分校)

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.LG

AI总结 FilterRAG通过结合BLIP-VQA与检索增强生成,利用外部知识源减少视觉问答中的幻觉问题,提升模型在知识驱动和分布外场景的鲁棒性。

Comments 12 pages, 6 figures and 2 tables; Accepted at ICCV 2025 Workshop on Building Foundation Models You Can Trust (T2FM)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12658 2026-01-27 cs.CL cs.AI 57%

Augmenting Question Answering with A Hybrid RAG Approach

通过混合RAG方法增强问答

Tianyi Yang, Nashrah Haque, Vaishnave Jonnalagadda, Yuya Jeremy Ong, Zhehui Chen, Yanzhao Wu, Lei Yu, Divyesh Jadav, Wenqi Wei

机构 * Plastic Lab(塑料实验室) Google(谷歌) Florida International University(佛罗里达国际大学) Rensselaer Polytechnic Institute(伦塞拉尔理工学院)

专题命中 视觉问答 :grounding(abstract);分类 cs.AI

AI总结 本文提出SSRAG方法,通过混合查询增强、代理路由和结构化检索技术,提升问答任务的响应质量。

Comments 10 pages, 5 tables, 2 figures; presented at IEEE CogMI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.09396 2026-01-27 cs.CV 57%

Too Many Frames, Not All Useful: Efficient Strategies for Long-Form Video QA

过多的帧,但并非都有用:长视频问答的高效策略

Jongwoo Park, Kanchana Ranasinghe, Kumara Kahatapitiya, Wonjeong Ryu, Donghyun Kim, Michael S. Ryoo

机构 * Stony Brook University(石溪大学) KAIST AI(韩国科学技术院人工智能研究所) Korea University(韩国大学)

专题命中 视觉问答 :vision language model(abstract);分类 cs.CV

AI总结 LVNet通过高效的关键帧选择器提升长视频问答效率,实现四个基准数据集上的最佳性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17191 2026-01-27 cs.CV cs.CL 57%

VaseVQA: Multimodal Agent and Benchmark for Ancient Greek Pottery

VaseVQA: 古代希腊陶器的多模态代理与基准

Jinchao Ge, Tengfei Cheng, Biao Wu, Zeyu Zhang, Shiya Huang, Judith Bishop, Gillian Shepherd, Meng Fang, Ling Chen, Yang Zhao

机构 * University of Adelaide(阿德莱德大学) University of Liverpool(利物浦大学) University of Technology Sydney(悉尼技术大学) The Australian National University(澳大利亚国立大学) La Trobe University(拉特罗布大学)

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

AI总结 VaseVQA通过引入VaseVL强化学习方法,提升对古代希腊陶器的多模态推理能力,尤其在复杂推理任务中表现更优。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17212 2026-01-27 cs.CL 50%

DF-RAG: Query-Aware Diversity for Retrieval-Augmented Generation

DF-RAG:基于查询的多样性检索增强生成

Saadat Hasan Khan, Spencer Hong, Jingyu Wu, Kevin Lybarger, Youbing Yin, Erin Babinsky, Daben Liu

机构 * George Mason University(乔治·马歇尔大学) Capital One

专题命中 视觉问答 :grounding(abstract)

AI总结 DF-RAG通过在检索阶段引入多样性,提升复杂推理问答任务的F1性能,相比传统RAG提升了4-10个百分点,并接近Oracle上限的91.3%

Comments Accepted to Findings of EACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 视觉推理 14 篇

2601.18356 2026-01-27 cs.LG 85%

Making medical vision-language models think causally across modalities with retrieval-augmented cross-modal reasoning

通过检索增强的跨模态推理使医疗视觉-语言模型在多模态中实现因果推理

Weiqin Yang, Haowen Xue, Qingyi Peng, Hexuan Hu, Qian Huang, Tingbo Zhang

机构 * University of Adelaide(阿德莱德大学) Hohai University(河海大学) Amap

专题命中 视觉推理 :vision-language model(title,abstract);visual question answering(abstract);grounding(abstract);分类 cs.LG

AI总结 本文提出多模态因果检索增强生成框架,通过整合因果推理原理与多模态检索,提升医疗VLMs在诊断预测和视觉问答中的事实准确性与鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13968 2026-01-27 cs.CV cs.AI cs.CL 84%

RotBench: Evaluating Multimodal Large Language Models on Identifying Image Rotation

RotBench: 对多模态大语言模型识别图像旋转能力的评估

Tianyi Niu, Jaemin Cho, Elias Stengel-Eskin, Mohit Bansal

机构 * UNC Chapel Hill(北卡罗来纳大学夏洛特分校) Allen Institute for Artificial Intelligence(人工智能研究院) The University of Texas at Austin(德克萨斯大学奥斯汀分校)

专题命中 视觉推理 :multimodal large language model(title,abstract);visual reasoning(abstract);分类 cs.CV、cs.AI

AI总结 RotBench评估了多模态大语言模型在识别图像旋转角度方面的性能,发现大多数模型难以区分90°和270°旋转,但能识别0°和180°图像,揭示了模型空间推理能力与人类的差距。

Comments EACL 2026 Camera-Ready. Code and data: https://github.com/tianyiniu/RotBench

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11020 2026-01-27 cs.CV cs.AI 81%

GeoVLMath: Enhancing Geometry Reasoning in Vision-Language Models via Cross-Modal Reward for Auxiliary Line Creation

GeoVLMath: 通过跨模态奖励增强视觉-语言模型中的几何推理以辅助线创建

Shasha Guo, Liang Pang, Xi Wang, Yanling Wang, Huawei Shen, Jing Zhang

机构 * Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所) Renmin University of China(中国人民大学) Zhipu AI(智谱AI)

专题命中 视觉推理 :vision-language model(title,abstract);分类 cs.CV、cs.AI

AI总结 GeoVLMath通过跨模态奖励模型提升视觉-语言模型在复杂立体几何问题中的几何推理能力。

Comments 19 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.18386 2026-01-27 cs.CV 70%

ARMOR: Agentic Reasoning for Methods Orchestration and Reparameterization for Robust Adversarial Attacks

ARMOR: 为对抗攻击的鲁棒性进行方法编排与重参数化中的代理推理

Gabriel Lee Jun Rong, Christos Korgialas, Dion Jia Xu Ho, Pai Chet Ng, Xiaoxiao Miao, Konstantinos N. Plataniotis

专题命中 视觉推理 :vision language model(abstract);VLM(abstract);分类 cs.CV

AI总结 ARMOR通过视觉语言模型引导的代理协作,提升对抗攻击的鲁棒性和跨架构迁移能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18619 2026-01-27 cs.AI 70%

Visual Attention Reasoning via Hierarchical Search and Self-Verification

通过分层搜索与自验证的视觉注意力推理

Wei Cai, Jian Zhao, Yuchen Yuan, Tianle Zhang, Ming Zhu, Haichuan Tang, Xuelong Li

专题命中 视觉推理 :grounding(abstract);multimodal large language model(abstract);分类 cs.AI

AI总结 本文提出通过分层搜索与自验证的视觉注意力推理框架,有效提升多模态大语言模型的视觉定位和推理能力,显著降低幻觉发生率。

Comments The paper is withdrawn by the authors after discovering a flaw in the theoretical derivation presented in the Method section. This incorrect step leads to conclusions that are not supported by the corrected derivation. The authors plan to reconstruct the argument and will release an updated version once the issue is fully resolved

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17197 2026-01-27 cs.CL cs.LG 70%

Reasoning Beyond Literal: Cross-style Multimodal Reasoning for Figurative Language Understanding

超越字面:跨风格多模态推理用于隐喻语言理解

Seyyed Saeid Cheshmi, Hahnemann Ortiz, James Mooney, Dongyeop Kang

机构 * University of Minnesota(明尼苏达大学)

专题命中 视觉推理 :vision-language model(abstract);VLM(abstract);分类 cs.LG

AI总结 本文提出一种三步框架,通过跨风格多模态推理提升隐喻语言理解能力,实验显示推理轨迹和跨风格训练能显著提升模型性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17123 2026-01-27 cs.HC cs.CV cs.RO 70%

Acoustic Field Video for Multimodal Scene Understanding

用于多模态场景理解的声学场视频

Daehwa Kim, Chris Harrison

专题命中 视觉推理 :vision-language model(abstract);VLM(abstract);分类 cs.CV

AI总结 本文提出声学场视频作为多模态场景理解的新输入方式,通过整合空间声学数据显著提升视觉-语言模型的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.24330 2026-01-27 cs.CV 70%

SenseNova-MARS: Empowering Multimodal Agentic Reasoning and Search via Reinforcement Learning

SenseNova-MARS: 通过强化学习赋能多模态代理推理与搜索

Yong Xien Chng, Tao Hu, Wenwen Tong, Xueheng Li, Jiandong Chen, Haojia Yu, Jiefan Lu, Hewei Guo, Hanming Deng, Chengjun Xie, Gao Huang, Dahua Lin, Lewei Lu

机构 * SenseTime Research(商汤科技研究院) Tsinghua University(清华大学) University of Science and Technology of China(中国科学技术大学)

专题命中 视觉推理 :vision-language model(abstract);visual reasoning(abstract);分类 cs.CV

AI总结 SenseNova-MARS通过强化学习赋能多模态代理推理与搜索,提升视觉-语言模型在复杂视觉任务中的表现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00555 2026-01-27 cs.LG cs.AI cs.CL cs.CV 67%

MMedAgent-RL: Optimizing Multi-Agent Collaboration for Multimodal Medical Reasoning

MMedAgent-RL: 优化多智能体协作以实现多模态医疗推理

Peng Xia, Jinglu Wang, Yibo Peng, Kaide Zeng, Zihan Dong, Xian Wu, Xiangru Tang, Hongtu Zhu, Yun Li, Linjun Zhang, Shujie Liu, Yan Lu, Huaxiu Yao

机构 * UNC-Chapel Hill(北卡罗来纳大学教堂山分校) Microsoft Research(微软研究院) CMU(卡内基梅隆大学) Rutgers University(罗格斯大学) Yale University(耶鲁大学)

专题命中 视觉推理 :vision-language model(abstract);分类 cs.CV、cs.AI、cs.LG

AI总结 MMedAgent-RL通过强化学习优化多智能体协作,提升多模态医疗推理性能,实现23.6%的性能提升。

Comments ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15304 2026-01-27 cs.IR 67%

MLLMRec: A Preference Reasoning Paradigm with Graph Refinement for Multimodal Recommendation

MLLMRec: 基于图细化的多模态推荐偏好推理范式

Yuzhuo Dang, Xin Zhang, Zhiqiang Pan, Yuxiao Duan, Wanyu Chen, Fei Cai, Honghui Chen

专题命中 视觉推理 :multimodal large language model(abstract);MLLM(abstract)

AI总结 MLLMRec通过图细化和多模态大语言模型提升多模态推荐的用户偏好推理与物品表示学习准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.06835 2026-01-27 cs.CV 57%

X-LeBench: A Benchmark for Extremely Long Egocentric Video Understanding

X-LeBench:一个用于极长第一人称视频理解的基准数据集

Wenqi Zhou, Kai Cao, Hao Zheng, Yunze Liu, Xinyi Zheng, Miao Liu, Per Ola Kristensson, Walterio Mayol-Cuevas, Fan Zhang, Weizhe Lin, Junxiao Shen

机构 * University of Bristol(布里斯托大学) University of Manchester(曼彻斯特大学) University of Cambridge(剑桥大学) College of AI, Tsinghua University(清华大学人工智能学院) Meta Memories.ai Research(Memories.ai研究)

专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.CV

AI总结 X-LeBench通过模拟真实日常生活场景,构建了首个极长第一人称视频理解基准数据集,揭示了长视频理解中的关键挑战。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17706 2026-01-27 cs.CL cs.CV 57%

A Computational Approach to Visual Metonymy

一种视觉隐喻的计算方法

Saptarshi Ghosh, Linfeng Liu, Tianyu Jiang

机构 * University of Cincinnati(辛辛那提大学)

专题命中 视觉推理 :vision-language model(abstract);分类 cs.CV

AI总结 本文提出了一种基于符号学理论的计算方法,通过构建ViMET数据集评估多模态语言模型在理解视觉隐喻方面的认知推理能力,并揭示了机器在处理间接视觉参考上的局限性。

Comments EACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17223 2026-01-27 cs.CL cs.AI 57%

Beyond Outcome Verification: Verifiable Process Reward Models for Structured Reasoning

超越结果验证:用于结构化推理的可验证过程奖励模型

Massimiliano Pronesti, Anya Belz, Yufang Hou

机构 * IBM Research Europe - Ireland(IBM欧洲研究院-爱尔兰) Dublin City University(都柏林城市大学) IT:U Interdisciplinary Transformation University Austria(IT:U跨学科转型大学奥地利)

专题命中 视觉推理 :grounding(abstract);分类 cs.AI

AI总结 本文提出可验证过程奖励模型,用于提升结构化推理的连贯性和准确性,实验显示其在医学证据评估中优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25975 2026-01-27 cs.CL cs.PL 50%

SymCode: A Neurosymbolic Approach to Mathematical Reasoning via Verifiable Code Generation

SymCode:通过可验证代码生成实现数学推理的神经符号方法

Sina Bagheri Nezhad, Yao Li, Ameeta Agrawal

机构 * Portland State University(波特兰州立大学) ElastixAI

专题命中 视觉推理 :grounding(abstract)

AI总结 SymCode通过可验证代码生成实现数学推理,显著提升准确性并增强模型的透明度和可靠性。

Comments camera-ready EACL 2026 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视觉定位与Grounding 21 篇

2510.11178 2026-01-27 cs.CV cs.CY 85%

BLEnD-Vis: Benchmarking Multimodal Cultural Understanding in Vision Language Models

BLEnD-Vis: 评估视觉语言模型在多模态文化理解中的基准测试

Bryan Chen Zhengyu Tan, Zheng Weihua, Zhengyuan Liu, Nancy F. Chen, Hwaran Lee, Kenny Tsu Wei Choo, Roy Ka-Wei Lee

专题命中 视觉定位与Grounding :vision language model(title);vision-language model(abstract);VLM(abstract);grounding(abstract)

AI总结 BLEnD-Vis通过多模态文化基准测试评估视觉语言模型在语言重述和视觉模态下的文化理解稳健性,揭示当前模型在文化知识上的脆弱性,并指导更文化胜任的模型发展。

Comments EACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17050 2026-01-27 cs.CV cs.AI 84%

Single-Pixel Vision-Language Model for Intrinsic Privacy-Preserving Behavioral Intelligence

单像素视觉-语言模型用于内在隐私保护的行为智能

Hongjun An, Yiliang Song, Jiawei Shao, Zhe Sun, Xuelong Li

机构 * Institute of Artificial Intelligence (TeleAI), China Telecom(人工智能研究院(TeleAI),中国电信)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);VLM(abstract);分类 cs.CV、cs.AI

AI总结 单像素视觉-语言模型通过内在隐私保护机制,在隐私敏感环境中实现安全监控与行为智能的平衡。

Comments Initial Version, Pending Updates. We welcome any feedback and suggestions for improvement. Please feel free to contact us at an.hongjun@foxmail.com

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17507 2026-01-27 cs.RO 82%

MetaWorld: Skill Transfer and Composition in a Hierarchical World Model for Grounding High-Level Instructions

MetaWorld: 一个用于地面指令基础的分层世界模型中的技能迁移与组合

Yutong Shen, Hangxu Liu, Kailin Pei, Ruizhe Xia, Tongtong Feng

机构 * Beijing University of Technology(北京理工大学) Fudan University(复旦大学) Tsinghua University(清华大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);VLM(abstract)

AI总结 MetaWorld通过整合语义规划与物理控制,利用专家策略迁移提升人形机器人在定位-操作任务中的性能。

Comments 8 pages, 4 figures, Submitted to ICLR 2026 World Model Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.16532 2026-01-27 cs.CV 79%

AnchoredDream: Zero-Shot 360° Indoor Scene Generation from a Single View via Geometric Grounding

AnchoredDream: 通过几何约束从单视角生成零样本360°室内场景

Runmao Yao, Junsheng Zhou, Zhen Dong, Yu-Shen Liu

机构 * School of Software, Tsinghua University(清华大学软件学院) Wuhan University(武汉大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

AI总结 AnchoredDream通过几何约束和外观-几何互促机制,实现从单视角零样本生成高质量360°室内场景。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15871 2026-01-27 cs.CV cs.MM 79%

Zero-Shot Visual Grounding in 3D Gaussians via View Retrieval

通过视图检索实现3D高斯体的零样本视觉定位

Liwei Liao, Xufeng Li, Xiaoyun Zheng, Boning Liu, Feng Gao, Ronggang Wang

机构 * Peking University(北京大学) Peng Cheng Laboratory(鹏城实验室) City University of Hongkong(香港城市大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

AI总结 本文提出GVR框架,通过视图检索将3DVG转化为2D检索任务,实现零样本3DGS的视觉定位,无需每场景训练和3D标注。

详情

展开后加载摘要…

URL PDF HTML 收藏