arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-10-24 至 2025-10-24 共收录 8 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 8 篇

2510.20244 2025-10-24 cs.CV cs.LG 84%

Empower Words: DualGround for Structured Phrase and Sentence-Level Temporal Grounding

Minseok Kang, Minhyeok Lee, Minjung Kim, Donghyeong Kim, Sangyoun Lee

机构 * Yonsei University(延世大学) LG Electronics(LG电子)

专题命中 视觉定位与Grounding :grounding(title,abstract);vision-language model(abstract);分类 cs.CV、cs.LG

Comments Comments: 28 pages, including appendix. 5 figures. Full version of the NeurIPS 2025 paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20612 2025-10-24 cs.CV cs.CL cs.LG 84%

Roboflow100-VL: A Multi-Domain Object Detection Benchmark for Vision-Language Models

Peter Robicheaux, Matvei Popov, Anish Madan, Isaac Robinson, Joseph Nelson, Deva Ramanan, Neehar Peri

机构 * Roboflow Carnegie Mellon University(卡内基梅隆大学)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);VLM(abstract);分类 cs.CV、cs.LG

Comments The first two authors contributed equally. This work has been accepted to the Neural Information Processing Systems (NeurIPS) 2025 Datasets & Benchmark Track. Project Page: https://rf100-vl.org/

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23010 2025-10-24 cs.CV 83%

SeG-SR: Integrating Semantic Knowledge into Remote Sensing Image Super-Resolution via Vision-Language Model

Bowen Chen, Keyan Chen, Mohan Yang, Zhengxia Zou, Zhenwei Shi

专题命中 视觉定位与Grounding :vision-language model(title,abstract);VLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19678 2025-10-24 cs.CL cs.CV 79%

Grounding Language with Vision: A Conditional Mutual Information Calibrated Decoding Strategy for Reducing Hallucinations in LVLMs

Hao Fang, Changle Zhou, Jiawei Kong, Kuofeng Gao, Bin Chen, Shu-Tao Xia

机构 * Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院,清华大学) Harbin Institute of Technology(哈尔滨工业大学)

专题命中 视觉定位与Grounding :grounding(title);vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20531 2025-10-24 cs.CV cs.AI 79%

Fake-in-Facext: Towards Fine-Grained Explainable DeepFake Analysis

Lixiong Qin, Yang Zhang, Mei Wang, Jiani Hu, Weihong Deng, Weiran Xu

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Beijing Normal University(北京师范大学)

专题命中 视觉定位与Grounding :grounding(abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments 25 pages, 9 figures, 17 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20229 2025-10-24 cs.CV cs.AI cs.CL 62%

Why LVLMs Are More Prone to Hallucinations in Longer Responses: The Role of Context

Ge Zheng, Jiaye Qian, Jiajin Tang, Sibei Yang

机构 * School of Computer Science and Engineering, Sun Yat-sen University(中山大学计算机科学与工程学院) ShanghaiTech University(上海理工大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.AI

Journal ref Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2025, pp. 4101-4113

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25033 2025-10-24 cs.CV cs.LG 62%

VT-FSL: Bridging Vision and Text with LLMs for Few-Shot Learning

Wenhao Li, Qiangchang Wang, Xianjing Meng, Zhibin Wu, Yilong Yin

机构 * School of Software, Shandong University(山东大学软件学院) Shenzhen Loop Area Institute(深圳河套学院) School of Computing and Artificial Intelligence, Shandong University of Finance and Economics(山东财经大学计算机与人工智能学院)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.LG

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20161 2025-10-24 cs.RO 50%

PathFormer: A Transformer with 3D Grid Constraints for Digital Twin Robot-Arm Trajectory Generation

Ahmed Alanazi, Duy Ho, Yugyung Lee

机构 * Department of Computer Science, University of Missouri–Kansas City (UMKC)(密苏里大学哥伦比亚分校计算机科学系) Department of Computer Science, California State University, Fullerton(加州州立大学富尔顿分校计算机科学系)

专题命中 视觉定位与Grounding :grounding(abstract)

Comments 8 pages, 7 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏