arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-10-29 至 2025-10-29 共收录 7 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7 篇

2510.23775 2025-10-29 cs.CV cs.AI eess.IV 84%

Explainable Detection of AI-Generated Images with Artifact Localization Using Faster-Than-Lies and Vision-Language Models for Edge Devices

Aryan Mathur, Asaduddin Ahmed, Pushti Amit Vasoya, Simeon Kandan Sonar, Yasir Z, Madesh Kuppusamy

机构 * Indian Institute of Technology Palakkad, India(印度帕拉卡德理工学院)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);VLM(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04897 2025-10-29 cs.CV 80%

From Objects to Anywhere: A Holistic Benchmark for Multi-level Visual Grounding in 3D Scenes

Tianxu Wang, Zhuofan Zhang, Ziyu Zhu, Yue Fan, Jing Xiong, Pengxiang Li, Xiaojian Ma, Qing Li

机构 * State Key Laboratory of General Artificial Intelligence, BIGAI(通用人工智能国家重点实验室,BIGAI) Tsinghua University(清华大学) Peking University(北京大学) Beijing Institute of Technology(北京理工大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV;multimodal large language model(comments)

Comments Update v3 of the NeurIPS 2025 Datasets and Benchmarks paper (v2), including additional evaluations of state-of-the-art multimodal large language models. Project page: https://anywhere-3d.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22672 2025-10-29 cs.CV cs.CL cs.RO 79%

Look and Tell: A Dataset for Multimodal Grounding Across Egocentric and Exocentric Views

Anna Deichler, Jonas Beskow

机构 * KTH Royal Institute of Technology(皇家理工学院)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments 10 pages, 6 figures, 2 tables. Accepted to the NeurIPS 2025 Workshop on SPACE in Vision, Language, and Embodied AI (SpaVLE). Dataset: https://huggingface.co/datasets/annadeichler/KTH-ARIA-referential

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15256 2025-10-29 cs.CV 79%

Normal and Abnormal Pathology Knowledge-Augmented Vision-Language Model for Anomaly Detection in Pathology Images

Jinsol Song, Jiamu Wang, Anh Tien Nguyen, Keunho Byeon, Sangjeong Ahn, Sung Hak Lee, Jin Tae Kwak

机构 * Korea University(韩国大学) The Catholic University of Korea(韩国天主大学)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV

Comments Accepted to ICCV 2025. Code is available at: https://github.com/QuIIL/ICCV2025_Ano-NAViLa

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24010 2025-10-29 cs.CV cs.AI cs.LG 67%

Mars-Bench: A Benchmark for Evaluating Foundation Models for Mars Science Tasks

Mirali Purohit, Bimal Gajera, Vatsal Malaviya, Irish Mehta, Kunal Kasodekar, Jacob Adler, Steven Lu, Umaa Rebbapragada, Hannah Kerner

机构 * School of Computing and Augmented Intelligence, Arizona State University(计算与增强智能学院,亚利桑那州立大学) School of Earth and Space Exploration, Arizona State University(地球与空间探索学院,亚利桑那州立大学) Jet Propulsion Laboratory, California Institute of Technology(喷气推进实验室,加州理工学院)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.AI、cs.LG

Comments Accepted at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24055 2025-10-29 cs.RO cs.LG 57%

Language-Conditioned Representations and Mixture-of-Experts Policy for Robust Multi-Task Robotic Manipulation

Xiucheng Zhang, Yang Jiang, Hongwei Qing, Jiashuo Bai

机构 * Xiucheng Zhang(未知) Yang Jiang(未知) Jiashuo Bai(未知)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG

Comments 8 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.15536 2025-10-29 cs.RO cs.AI 57%

GRS: Generating Robotic Simulation Tasks from Real-World Images

Alex Zook, Fan-Yun Sun, Josef Spjut, Valts Blukis, Stan Birchfield, Jonathan Tremblay

机构 * NVIDIA Stanford University(斯坦福大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏