arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-09-23 至 2025-09-23 共收录 17 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 17 篇

2506.21188 2025-09-23 cs.CV 79%

GroundFlow: A Plug-in Module for Temporal Reasoning on 3D Point Cloud Sequential Grounding

Zijun Lin, Shuting He, Cheston Tan, Bihan Wen

机构 * Nanyang Technological University(南洋理工大学) Centre for Frontier AI Research, A*STAR(前沿人工智能研究中心,A*STAR) MoE Key Laboratory of Interdisciplinary Research of Computation and Economics, Shanghai University of Finance and Economics(计算与经济学交叉研究关键实验室,上海财经大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16806 2025-09-23 cs.CV 79%

FOCUS: Unified Vision-Language Modeling for Interactive Editing Driven by Referential Segmentation

Fan Yang, Yousong Zhu, Xin Li, Yufei Zhan, Hongyin Zhao, Shurong Zheng, Yaowei Wang, Ming Tang, Jinqiao Wang

机构 * Foundation Model Research Center, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所基础模型研究中心) School of Artificial Intelligence, University of Chinese Academy of Science(中国科学院大学人工智能学院) Peng Cheng Laboratory, Shenzhen, China(鹏城实验室) Wuhan AI Research, Wuhan, China(武汉人工智能研究所)

专题命中 视觉定位与Grounding :vision-language model(title);vision language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16534 2025-09-23 cs.CL cs.AI 79%

InteGround: On the Evaluation of Verification and Retrieval Planning in Integrative Grounding

Cheng Jiayang, Qianqian Zhuang, Haoran Li, Chunkit Chan, Xin Liu, Lin Qiu, Yangqiu Song

机构 * The Hong Kong University of Science and Technology(香港科学与技术大学) Shanghai Jiaotong University(上海交通大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

Comments Accepted to EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21535 2025-09-23 eess.IV cs.CV cs.LG 62%

Exploring the Design Space of 3D MLLMs for CT Report Generation

Mohammed Baharoon, Jun Ma, Congyu Fang, Augustin Toma, Bo Wang

机构 * Vector Institute for Artificial Intelligence(向量人工智能研究所) Department of Biomedical Informatics, Harvard Medical School(生物医学信息学系,哈佛医学院) Peter Munk Cardiac Centre, University Health Network(皮特·蒙克心脏中心,大学健康网络) Medical Biophysics, University of Toronto(医学生物物理系,多伦多大学) Department of Computer Science, University of Toronto(计算机科学系,多伦多大学) Department of Laboratory Medicine and Pathobiology, University of Toronto(实验室医学与病理学系,多伦多大学) AI Hub, University Health Network(人工智能中心,大学健康网络)

专题命中 视觉定位与Grounding :multimodal large language model(abstract);分类 cs.CV、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18096 2025-09-23 cs.CV 57%

Seg4Diff: Unveiling Open-Vocabulary Segmentation in Text-to-Image Diffusion Transformers

Chaehyun Kim, Heeseong Shin, Eunbeen Hong, Heeji Yoon, Anurag Arnab, Paul Hongsuck Seo, Sunghwan Hong, Seungryong Kim

机构 * KAIST AI(韩国科学技术院人工智能研究所) Korea University(韩国大学) ETH Zürich(苏黎世联邦理工学院)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments NeurIPS 2025. Project page: https://cvlab-kaist.github.io/Seg4Diff/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17671 2025-09-23 cs.CL cs.AI 57%

Turk-LettuceDetect: A Hallucination Detection Models for Turkish RAG Applications

Selva Taş, Mahmut El Huseyni, Özay Ezerceli, Reyhan Bayraktar, Fatma Betül Terzioğlu

机构 * Hidden for Review(保密)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17615 2025-09-23 cs.CV 57%

From Benchmarks to Reality: Advancing Visual Anomaly Detection by the VAND 3.0 Challenge

Lars Heckler-Kram, Ashwin Vaidya, Jan-Hendrik Neudeck, Ulla Scheler, Dick Ameln, Samet Akcay, Paula Ramos

机构 * MVTec Software GmbH(MVTec软件公司) Technical University of Munich(慕尼黑技术大学) Intel(英特尔) Voxel51

专题命中 视觉定位与Grounding :vision language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17522 2025-09-23 cs.CV 57%

Chat-CBM: Towards Interactive Concept Bottleneck Models with Frozen Large Language Models

Hangzhou He, Lei Zhu, Kaiwen Li, Xinliang Zhang, Jiakui Hu, Ourui Fu, Zhengjian Yao, Yanye Lu

机构 * Department of Biomedical Engineering, College of Future Technology, Peking University(生物医学工程系,未来技术学院,北京大学) Institute of Medical Technology, Peking University Health Science Center, Peking University(医学技术研究所,北京大学医学部,北京大学) National Biomedical Imaging Center, College of Future Technology, Peking University(国家生物医学成像中心,未来技术学院,北京大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16810 2025-09-23 cs.AI 57%

Automated Procedural Analysis via Video-Language Models for AI-assisted Nursing Skills Assessment

Shen Chang, Dennis Liu, Renran Tian, Kristen L. Swartzell, Stacie L. Klingler, Amy M. Nagle, Nan Kong

机构 * Weldon School of Biomedical Engineering, Purdue University(普渡大学生物医学工程学院) Department of Industrial and Operations Engineering, University of Michigan(密歇根大学工业与运作工程系) Edward P. Fitts Department of Industrial and Systems Engineering, North Carolina State University(北卡罗来纳州立大学工业与系统工程系) School of Nursing, Purdue University(普渡大学护理学院)

专题命中 视觉定位与Grounding :VLM(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04484 2025-09-23 cs.CL cs.AI cs.CY 57%

The Good, the Bad and the Constructive: Automatically Measuring Peer Review's Utility for Authors

Abdelrahman Sadallah, Tim Baumgärtner, Iryna Gurevych, Ted Briscoe

机构 * NLP Department, Mohamed Bin Zayed University of Artificial Intelligence(马尔代夫比兹艾兹大学人工智能学院自然语言处理系) Ubiquitous Knowledge Processing Lab, Department of Computer Science(计算机科学系通用知识处理实验室) Hessian Center for AI (hessian.AI), TU Darmstadt(图尔努尔德马斯特大学海斯塞人工智能中心)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments EMNLP 2025 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15108 2025-09-23 cs.CL cs.AI cs.HC 57%

A Risk Ontology for Evaluating AI-Powered Psychotherapy Virtual Agents

Ian Steenstra, Timothy W. Bickmore

机构 * Northeastern University(东北大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments This is a preprint version of the paper accepted to IVA'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16502 2025-09-23 cs.LG 57%

GRIL: Knowledge Graph Retrieval-Integrated Learning with Large Language Models

Jialin Chen, Houyu Zhang, Seongjun Yun, Alejandro Mottini, Rex Ying, Xiang Song, Vassilis N. Ioannidis, Zheng Li, Qingjun Cui

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16297 2025-09-23 cs.CY cs.AI cs.CL 57%

How Large Language Models are Designed to Hallucinate

Richard Ackermann, Simeon Emanuilov

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments 23 pages, 2 tables, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17727 2025-09-23 cs.CY cs.IT math.IT 50%

Empirical AI Ethics: Reconfiguring Ethics towards a Situated, Plural, and Transformative Approach

Paula Helm, Selin Gerlek

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17523 2025-09-23 cs.CL eess.AS 50%

Leveraging Audio-Visual Data to Reduce the Multilingual Gap in Self-Supervised Speech Models

María Andrea Cruz Blandón, Zakaria Aldeneh, Jie Chi, Maureen de Seyssel

机构 * Tampere University Apple(塔尔库大学苹果)

专题命中 视觉定位与Grounding :grounding(abstract)

Comments 5 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16670 2025-09-23 cs.SD cs.MM eess.AS 50%

Speech-to-See: End-to-End Speech-Driven Open-Set Object Detection

Wenhuan Lu, Xinyue Song, Wenjun Ke, Zhizhi Yu, Wenhao Yang, Jianguo Wei

机构 * College of Intelligence and Computing, Tianjin University, Tianjin, China(智能与计算学院,天津大学,天津,中国) PipeChina Institute of Science and Technology, Tianjin, China(中石油科技研究院,天津,中国)

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.08851 2025-09-23 cs.RO 50%

OTAS: Open-vocabulary Token Alignment for Outdoor Segmentation

Simon Schwaiger, Stefan Thalhammer, Wilfried Wöber, Gerald Steinbauer-Wagner

机构 * Graz University of Technology, Faculty of Computer Science and Biomedical Engineering, Institute of Software Engineering and Artificial Intelligence(格拉茨技术大学,计算机科学与生物医学工程学院,软件工程与人工智能研究所) University of Applied Sciences Technikum Wien, Faculty of Industrial Engineering, Research Group Digital Manufacturing, Automation and Robotics(应用科学大学技术学院,工业工程学院,数字制造、自动化与机器人研究组) University of Natural Resources and Life Sciences, Department of Integrative Biology and Biodiversity Research, Institute for Integrative Nature Conservation Research(自然资源与生命科学大学,整合生物学与生物多样性研究部门,整合自然保护研究 institute)

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏