arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-10-28 至 2025-10-28 共收录 88 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 19 篇

2503.22194 2025-10-28 cs.CV cs.LG 81%

ORIGEN: Zero-Shot 3D Orientation Grounding in Text-to-Image Generation

Yunhong Min, Daehyeon Choi, Kyeongmin Yeo, Jihyun Lee, Minhyuk Sung

机构 * KAIST(韩国科学技术院)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV、cs.LG

Comments Project Page: https://origen2025.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13227 2025-10-28 cs.AI cs.CL cs.CV cs.HC 81%

Scaling Computer-Use Grounding via User Interface Decomposition and Synthesis

Tianbao Xie, Jiaqi Deng, Xiaochuan Li, Junlin Yang, Haoyuan Wu, Jixuan Chen, Wenjing Hu, Xinyuan Wang, Yuhui Xu, Zekun Wang, Yiheng Xu, Junli Wang, Doyen Sahoo, Tao Yu, Caiming Xiong

机构 * The University of Hong Kong(香港大学) Salesforce AI Research(Salesforce AI研究院)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV、cs.AI

Comments 49 pages, 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03201 2025-10-28 cs.CV 79%

AlignCAT: Visual-Linguistic Alignment of Category and Attribute for Weakly Supervised Visual Grounding

Yidan Wang, Chenyi Zhuang, Wutao Liu, Pan Gao, Nicu Sebe

机构 * Nanjing University of Aeronautics and Astronautics(南京航空航天大学) University of Trento(特伦托大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.08216 2025-10-28 cs.AI 79%

Grounding Methods for Neural-Symbolic AI

Rodrigo Castellano Ontiveros, Francesco Giannini, Marco Gori, Giuseppe Marra, Michelangelo Diligenti

机构 * University of Siena(锡耶纳大学) Scuola Normale Superiore(正规大学) KU Leuven(卢森堡大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

Journal ref Proceedings of the Thirty-Fourth International Joint Conference on Artificial Intelligence (IJCAI-25), pp. 4806-4814, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21076 2025-10-28 cs.CV 79%

DynamicVL: Benchmarking Multimodal Large Language Models for Dynamic City Understanding

Weihao Xuan, Junjue Wang, Heli Qi, Zihang Chen, Zhuo Zheng, Yanfei Zhong, Junshi Xia, Naoto Yokoya

机构 * The University of Tokyo(东京大学) RIKEN AIP(理化学研究所AIP) Waseda University(早稻田大学) Wuhan University(武汉大学) Stanford University(斯坦福大学)

专题命中 视觉定位与Grounding :multimodal large language model(title,abstract);分类 cs.CV

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23473 2025-10-28 cs.CV 70%

Video-Thinker: Sparking "Thinking with Videos" via Reinforcement Learning

Shijian Wang, Jiarui Jin, Xingjian Wang, Linxin Song, Runhao Fu, Hecheng Wang, Zongyuan Ge, Yuan Lu, Xuelian Cheng

机构 * Southeast University(东南大学) Monash University(墨尔本大学) Xiaohongshu Inc.(小红书公司) University of Southern California(南加州大学) Fudan University(复旦大学)

专题命中 视觉定位与Grounding :grounding(abstract);multimodal large language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23190 2025-10-28 cs.CV 70%

Evaluation of Vision-LLMs in Surveillance Video

Pascal Benschop, Cristian Meo, Justin Dauwels, Jelte P. Mense

机构 * Department of Computer Science(计算机科学系) Delft University of Technology(代尔夫特理工大学) LatentWorlds AI National Policelab AI & Model-Driven Decisions Lab(国家警务实验室AI与模型驱动决策实验室)

专题命中 视觉定位与Grounding :vision-language model(abstract);grounding(abstract);分类 cs.CV

Comments Accepted as poster in the NeurIPS 2025 Workshop on Space in Vision, Language, and Embodied AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15963 2025-10-28 cs.CV cs.AI cs.LG 67%

ESCA: Contextualizing Embodied Agents via Scene-Graph Generation

Jiani Huang, Amish Sethi, Matthew Kuo, Mayank Keoliya, Neelay Velingker, JungHo Jung, Ser-Nam Lim, Ziyang Li, Mayur Naik

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI、cs.LG

Comments Accepted as a Spotlight Paper at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22521 2025-10-28 cs.CV cs.AI cs.IR cs.LG 67%

Open Multimodal Retrieval-Augmented Factual Image Generation

Yang Tian, Fan Liu, Jingyuan Zhang, Wei Bi, Yupeng Hu, Liqiang Nie

机构 * Shandong University(山东大学) National University of Singapore(新加坡国立大学) Kuaishou Technology(快手科技) Harbin Institute of Technology(哈尔滨工业大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI、cs.LG

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17394 2025-10-28 cs.CV cs.AI 62%

HiProbe-VAD: Video Anomaly Detection via Hidden States Probing in Tuning-Free Multimodal LLMs

Zhaolin Cai, Fan Li, Ziwei Zheng, Yanjun Qin

机构 * Xinjiang University(新疆大学) Xi'an Jiaotong University(西安交通大学)

专题命中 视觉定位与Grounding :multimodal large language model(abstract);分类 cs.CV、cs.AI

Comments Accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20759 2025-10-28 cs.CV cs.AI 62%

PARTONOMY: Large Multimodal Models with Part-Level Visual Understanding

Ansel Blume, Jeonghwan Kim, Hyeonjeong Ha, Elen Chatikyan, Xiaomeng Jin, Khanh Duy Nguyen, Nanyun Peng, Kai-Wei Chang, Derek Hoiem, Heng Ji

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) University of California Los Angeles(加州大学洛杉矶分校)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI

Comments NeurIPS 2025 Spotlight; project page: https://wjdghks950.github.io/partonomy.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21763 2025-10-28 cs.CV cs.AI 62%

Proportion and Perspective Control for Flow-Based Image Generation

Julien Boudier, Hugo Caselles-Dupré

机构 * Obvious Research(Obvious研究)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.AI

Comments Technical report after open-source release

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23184 2025-10-28 cs.CV 57%

Finding 3D Scene Analogies with Multimodal Foundation Models

Junho Kim, Young Min Kim

机构 * Institute of New Media and Communications(新媒体与通讯研究所) Dept. of Electrical and Computer Engineering(电气与计算机工程系)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments Accepted to FM4RoboPlan workshop at RSS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19333 2025-10-28 cs.CV 57%

A Training-Free Framework for Open-Vocabulary Image Segmentation and Recognition with EfficientNet and CLIP

Ying Dai, Wei Yu Chen

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21900 2025-10-28 cs.CL cs.AI 57%

Deep Literature Survey Automation with an Iterative Workflow

Hongbo Zhang, Han Cui, Yidong Wang, Yijian Tian, Qi Guo, Cunxiang Wang, Jian Wu, Chiyu Song, Yue Zhang

机构 * Zhejiang University(浙江大学) School of Engineering, Westlake University(西湖大学工程学院) Peking University(北京大学) Institute of Advanced Technology, Westlake Institute for Advanced Study(西湖先进研究院技术研究所)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments Preprint version

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.03610 2025-10-28 cs.CV 57%

Learning Knowledge-based Prompts for Robust 3D Mask Presentation Attack Detection

Fangling Jiang, Qi Li, Bing Liu, Weining Wang, Caifeng Shan, Zhenan Sun, Ming-Hsuan Yang

机构 * School of Computer Science, University of South China(南方大学计算机科学学院) New Laboratory of Pattern Recognition, MAIS, CASIA(模式识别新实验室,MAIS,CASIA) School of Intelligence Science and Technology, Nanjing University(智能科学与技术学院,南京大学) Department of Computer Science and Engineering, University of California, Merced(加州大学默塞德分校计算机科学与工程系) Department of Computer Science and Engineering, Yonsei University(延世大学计算机科学与工程系)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments Accepted by TPAMI

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 文档图表理解 2 篇

2510.23066 2025-10-28 cs.IR 78%

Multi-Stage Field Extraction of Financial Documents with OCR and Compact Vision-Language Models

Yichao Jin, Yushuo Wang, Qishuai Zhong, Kent Chiu Jin-Chun, Kenneth Zhu Ke, Donald MacDonald

专题命中 文档图表理解 :vision-language model(title);vision language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21774 2025-10-28 cs.CV cs.AI 62%

OCR-Quality: A Human-Annotated Dataset for OCR Quality Assessment

Yulong Zhang

机构 * Beijing University of Posts and Telecommunications(北京邮电大学)

专题命中 文档图表理解 :vision-language model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

3. GUI与屏幕智能体 5 篇

2510.21722 2025-10-28 cs.HC cs.AI 83%

AquaVLM: Improving Underwater Situation Awareness with Mobile Vision Language Models

Beitong Tian, Lingzhi Zhao, Bo Chen, Haozhen Zheng, Jingcheng Yang, Mingyuan Wu, Deepak Vasisht, Klara Nahrstedt

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 GUI与屏幕智能体 :vision language model(title);vision-language model(abstract);VLM(abstract);分类 cs.AI

Comments 12 pages, 10 figures, under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05020 2025-10-28 cs.RO cs.AI 70%

Hierarchical Language Models for Semantic Navigation and Manipulation in an Aerial-Ground Robotic System

Haokun Liu, Zhaoqi Ma, Yunong Li, Junichiro Sugihara, Yicheng Chen, Jinjie Li, Moju Zhao

机构 * DRAGON Lab at Department of Mechanical Engineering, The University of Tokyo(东京大学机械工程系DRAGON实验室)

专题命中 GUI与屏幕智能体 :vision-language model(abstract);VLM(abstract);分类 cs.AI

Comments 18 pages, 10 figures

Journal ref Advanced Intelligent Systems, Oct. 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15566 2025-10-28 cs.CV cs.AI 62%

BTL-UI: Blink-Think-Link Reasoning Model for GUI Agent

Shaojie Zhang, Ruoceng Zhang, Pei Fu, Shaokang Wang, Jiahui Yang, Xin Du, Shiqi Cui, Bin Qin, Ying Huang, Zhenbo Luo, Jian Luan

机构 * Xiaomi Inc(小米公司)

专题命中 GUI与屏幕智能体 :multimodal large language model(abstract);分类 cs.CV、cs.AI

Comments Accepted at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21761 2025-10-28 cs.RO cs.AI cs.CV 62%

J-ORA: A Framework and Multimodal Dataset for Japanese Object Identification, Reference, Action Prediction in Robot Perception

Jesse Atuhurra, Hidetaka Kamigaito, Taro Watanabe, Koichiro Yoshino

机构 * Division of Information Science, NAIST(NAIST信息科学系) Guardian Robot Project, RIKEN(RIKEN守护机器人项目)

专题命中 GUI与屏幕智能体 :vision language model(abstract);分类 cs.CV、cs.AI

Comments Accepted to IROS2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21809 2025-10-28 cs.CV cs.RO 57%

Embodied Navigation with Auxiliary Task of Action Description Prediction

Haru Kondoh, Asako Kanezaki

机构 * Institute of Science Tokyo(东京科学研究所) RIKEN AIP(日本科学技术研究所AIP)

专题命中 GUI与屏幕智能体 :vision-language model(abstract);分类 cs.CV

Comments ICCV 2025 Poster

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 幻觉与鲁棒性 7 篇

2502.12520 2025-10-28 cs.CV 85%

SafeEraser: Enhancing Safety in Multimodal Large Language Models through Multimodal Machine Unlearning

Junkai Chen, Zhijie Deng, Kening Zheng, Yibo Yan, Shuliang Liu, PeiJun Wu, Peijie Jiang, Jia Liu, Xuming Hu

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) The Hong Kong University of Science and Technology(香港科技大学) Southeast University(东南大学) Ant Group, Alibaba(蚂蚁集团)

专题命中 幻觉与鲁棒性 :multimodal large language model(title,abstract);LLaVA(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22045 2025-10-28 cs.CV cs.AI 84%

VLM-SlideEval: Evaluating VLMs on Structured Comprehension and Perturbation Sensitivity in PPT

Hyeonsu Kang, Emily Bao, Anjan Goswami

专题命中 幻觉与鲁棒性 :VLM(title,abstract);vision-language model(abstract);分类 cs.CV、cs.AI

Comments 39th Conference on Neural Information Processing Systems (NeurIPS 2025) Workshop: Evaluating the Evolving LLM Lifecycle - Benchmarks, Emergent Abilities, and Scaling

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21906 2025-10-28 cs.NE cs.AI 83%

Structure-Aware Cooperative Ensemble Evolutionary Optimization on Combinatorial Problems with Multimodal Large Language Models

Jie Zhao, Kang Hao Cheong

机构 * School of Physical and Mathematical Sciences, Nanyang Technological University(南洋理工大学物理与数学科学学院) College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算与数据科学学院)

专题命中 幻觉与鲁棒性 :multimodal large language model(title,abstract);MLLM(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22785 2025-10-28 cs.CV 79%

Self-Calibrated Consistency can Fight Back for Adversarial Robustness in Vision-Language Models

Jiaxiang Liu, Jiawei Du, Xiao Liu, Prayag Tiwari, Mingkun Xu

机构 * Guangdong Institute of Intelligence Science and Technology(广东智能科学与技术研究院) Agency for Science, Technology and Research(科技研究局) School of Information Technology(信息技术学院)

专题命中 幻觉与鲁棒性 :vision-language model(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19911 2025-10-28 cs.CV 79%

Attention! Your Vision Language Model Could Be Maliciously Manipulated

Xiaosen Wang, Shaokang Wang, Zhijin Ge, Yuyang Luo, Shudong Zhang

专题命中 幻觉与鲁棒性 :vision language model(title);vision-language model(abstract);分类 cs.CV

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.11654 2025-10-28 cs.LG cs.AI eess.SP 62%

R-SFLLM: Jamming Resilient Framework for Split Federated Learning with Large Language Models

Aladin Djuhera, Vlad C. Andrei, Xinyang Li, Ullrich J. Mönich, Holger Boche, Walid Saad

专题命中 幻觉与鲁棒性 :vision language model(abstract);分类 cs.AI、cs.LG

Journal ref IEEE Transactions on Information Forensics and Security (Volume: 20), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23217 2025-10-28 cs.CL cs.AI 57%

Process Reward Models for Sentence-Level Verification of LVLM Radiology Reports

Alois Thomas, Maya Varma, Jean-Benoit Delbrouck, Curtis P. Langlotz

机构 * AIMI Center Stanford University(AIMI中心 斯坦福大学)

专题命中 幻觉与鲁棒性 :vision-language model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏