arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-07-29 至 2025-07-29 共收录 51 信号源:cs.CV, cs.AI, cs.LG

1. 视觉问答 7 篇

2502.04469 2025-07-29 cs.CV cs.AI 81%

Ask and Remember: A Questions-Only Replay Strategy for Continual Visual Question Answering

Imad Eddine Marouf, Enzo Tartaglione, Stephane Lathuiliere, Joost van de Weijer

机构 * LTCI, Télécom-Paris, Institut Polytechnique de Paris(LTCI, Télécom-Paris, Institut Polytechnique de Paris) Inria, LJK, Univ. Grenoble Alpes(Inria, LJK, 火车头 Grenoble Alpes) Universitat Autónoma de Barcelona(Autonomous University of Barcelona)

专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV、cs.AI

Comments ICCV 2025, 8 pages. Code: https://github.com/IemProg/QUAD

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24371 2025-07-29 cs.CV cs.AI 79%

Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering

Md Intisar Chowdhury, Kittinun Aukkapinyo, Hiroshi Fujimura, Joo Ann Woo, Wasu Wasusatein, Fadoua Ghourabi

机构 * AWL, Inc.(AWL公司)

专题命中 视觉问答 :vision-language model(abstract);VLM(abstract);LLaVA(abstract);分类 cs.CV、cs.AI

Comments Copyright 2025 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.03328 2025-07-29 cs.CV cs.AI cs.NE 73%

Visual Enumeration Remains Challenging for Multimodal Generative AI

Alberto Testolin, Kuinan Hou, Marco Zorzi

机构 * Department of General Psychology and Department of Mathematics University of Padova(帕多瓦大学心理学系和数学系) Department of General Psychology University of Padova(帕多瓦大学心理学系) Department of General Psychology and Padova Neuroscience Center University of Padova(帕多瓦大学心理学系和帕多瓦神经科学中心) IRCSS San Camillo Hospital, Venice-Lido(威尼斯利多医院IRCSS桑卡莫医院)

专题命中 视觉问答 :LLaVA(abstract);visual question answering(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19525 2025-07-29 cs.LG cs.AI 73%

MMCircuitEval: A Comprehensive Multimodal Circuit-Focused Benchmark for Evaluating LLMs

Chenchen Zhao, Zhengyuan Shi, Xiangyu Wen, Chengjie Liu, Yi Liu, Yunhao Zhou, Yuxiang Zhao, Hefei Feng, Yinan Zhu, Gwok-Waa Wan, Xin Cheng, Weiyu Chen, Yongqi Fu, Chujie Chen, Chenhao Xue, Guangyu Sun, Ying Wang, Yibo Lin, Jun Yang, Ning Xu, Xi Wang, Qiang Xu

机构 * Department of Computer Science and Engineering, The Chinese University of Hong Kong(中国香港中文大学计算机科学与工程系) School of Electronic Science and Engineering, Nanjing University(南京大学电子科学与工程学院) School of Integrated Circuits, Peking University(北京大学集成电路学院) School of Intergrated Circuits, Southeast University(东南大学集成电路学院) School of Computer Science and Engineering, Southeast University(东南大学计算机科学与工程学院) Department of Computer Science and Technology, University of Chinese Academy of Sciences(中国科学院大学计算机科学与技术系) National Center of Technology Innovation for EDA(EDA技术创新国家中心)

专题命中 视觉问答 :multimodal large language model(abstract);MLLM(abstract);分类 cs.AI、cs.LG

Comments 10 pages, 1 figure, 5 tables. To appear in ICCAD 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.10503 2025-07-29 cs.CV cs.CL cs.LG 62%

Everything is a Video: Unifying Modalities through Next-Frame Prediction

G. Thomas Hudson, Dean Slack, Thomas Winterbottom, Jamie Sterling, Chenghao Xiao, Junjie Shentu, Noura Al Moubayed

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.LG

Comments 10 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20939 2025-07-29 cs.CV 57%

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts

Yuying Ge, Yixiao Ge, Chen Li, Teng Wang, Junfu Pu, Yizhuo Li, Lu Qiu, Jin Ma, Lisheng Duan, Xinyu Zuo, Jinwen Luo, Weibo Gu, Zexuan Li, Xiaojing Zhang, Yangyu Tao, Han Hu, Di Wang, Ying Shan

机构 * ARC Lab, Tencent PCG(腾讯PCG ARC实验室) Search Application Department, Tencent CSIG(腾讯CSIG搜索应用部门) Tencent Hunyuan(腾讯文生视频) Big Data Platform Department, Tencent PCG(腾讯PCG大数据平台部门)

专题命中 视觉问答 :grounding(abstract);分类 cs.CV

Comments Project Page: https://tencentarc.github.io/posts/arc-video-announcement/

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19882 2025-07-29 cs.AI 57%

Causality-aligned Prompt Learning via Diffusion-based Counterfactual Generation

Xinshu Li, Ruoyu Wang, Erdun Gao, Mingming Gong, Lina Yao

机构 * The University of New South Wales(新南威尔士大学) The University of Adelaide(阿德莱德大学) The University of Melbourne(墨尔本大学) CSIRO’s Data 61(CSIRO数据61)

专题命中 视觉问答 :visual question answering(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 视觉推理 10 篇

2507.20342 2025-07-29 cs.AI cs.RO 83%

VLMPlanner: Integrating Visual Language Models with Motion Planning

Zhipeng Tang, Sha Zhang, Jiajun Deng, Chenjie Wang, Guoliang You, Yuting Huang, Xinrui Lin, Yanyong Zhang

机构 * University of Science and Technology of China(科学技术大学) University of Adelaide(阿德莱德大学) Institute of Artificial Intelligence, Hefei Comprehensive National Science Center(合肥综合性国家科学中心人工智能研究院)

专题命中 视觉推理 :visual language model(title);vision-language model(abstract);VLM(abstract);分类 cs.AI

Comments 8 pages, 3 figures, this paper has been accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.12789 2025-07-29 cs.CV 83%

Efficient Physics Simulation for 3D Scenes via MLLM-Guided Gaussian Splatting

Haoyu Zhao, Hao Wang, Xingyue Zhao, Hao Fei, Hongqiu Wang, Chengjiang Long, Hua Zou

机构 * School of Computer Science, Wuhan University(武汉大学计算机学院) Wuhan National Laboratory for Optoelectronics, Huazhong University of Science and Technology(华中科技大学光电研究院) Meta Reality Lab(Meta现实实验室) Xi’an Jiao Tong University(西安交通大学) National University of Singapore(新加坡国立大学) The Department of Systems Hub, Hong Kong University of Science and Technology (Guangzhou)(香港科技大学系统枢纽部门(广州))

专题命中 视觉推理 :MLLM(title,abstract);visual reasoning(abstract);分类 cs.CV

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20529 2025-07-29 cs.CV cs.AI 73%

Enhancing Spatial Reasoning through Visual and Textual Thinking

Xun Liang, Xin Guo, Zhongming Jin, Weihang Pan, Penghui Shang, Deng Cai, Binbin Lin, Jieping Ye

专题命中 视觉推理 :vision language model(abstract);visual question answering(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.03916 2025-07-29 cs.AI cs.CV 73%

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models

Yifan Jiang, Yibo Xue, Yukun Kang, Pin Zheng, Jian Peng, Feiran Wu, Changliang Xu

机构 * Hangzhou Institute for Advanced Study(杭州先进研究所) University of Chinese Academy of Sciences(中国科学院大学) Alibaba Group(阿里巴巴集团)

专题命中 视觉推理 :vision-language model(abstract);VLM(abstract);分类 cs.CV、cs.AI

Comments Appendix at: https://github.com/PAMPAS-Lab/ANA-PPT-Anamation/blob/main/Appendix.pdf

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.10341 2025-07-29 cs.RO cs.AI cs.LG 73%

Affordance-Guided Reinforcement Learning via Visual Prompting

Olivia Y. Lee, Annie Xie, Kuan Fang, Karl Pertsch, Chelsea Finn

机构 * Stanford University(斯坦福大学) Cornell University(康奈尔大学) University of California, Berkeley(加州大学伯克利分校)

专题命中 视觉推理 :vision-language model(abstract);visual reasoning(abstract);分类 cs.AI、cs.LG

Comments 8 pages, 6 figures. IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.13082 2025-07-29 cs.RO cs.AI cs.CV 62%

Free-form language-based robotic reasoning and grasping

Runyu Jiao, Alice Fasoli, Francesco Giuliari, Matteo Bortolon, Sergio Povoli, Guofeng Mei, Yiming Wang, Fabio Poiesi

机构 * Fondazione Bruno Kessler(布鲁诺·科塞拉基金会) University of Trento(特伦托大学) Istituto Italiano di Tecnologia(意大利技术研究院)

专题命中 视觉推理 :vision-language model(abstract);分类 cs.CV、cs.AI

Comments Accepted to IROS 2025. Project website: https://tev-fbk.github.io/FreeGrasp/

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20509 2025-07-29 cs.RO cs.AI cs.SY eess.SY 57%

LLMs-guided adaptive compensator: Bringing Adaptivity to Automatic Control Systems with Large Language Models

Zhongchao Zhou, Yuxi Lu, Yaonan Zhu, Yifan Zhao, Bin He, Liang He, Wenwen Yu, Yusuke Iwasawa

专题命中 视觉推理 :vision-language model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13629 2025-07-29 cs.CV 57%

FreeQ-Graph: Free-form Querying with Semantic Consistent Scene Graph for 3D Scene Understanding

Chenlu Zhan, Yufei Zhang, Gaoang Wang, Hongwei Wang

机构 * College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院) College of Biomedical Engineering and Instrument Science, Zhejiang University(浙江大学生物医学工程与仪器科学学院) Zhejiang University-University of Illinois Urbana-Champaign Institute, Zhejiang University(浙江大学-伊利诺伊大学厄巴纳-香槟分校联合学院)

专题命中 视觉推理 :grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19885 2025-07-29 cs.CL 50%

Zero-shot Performance of Generative AI in Brazilian Portuguese Medical Exam

Cesar Augusto Madid Truyts, Amanda Gomes Rabelo, Gabriel Mesquita de Souza, Daniel Scaldaferri Lages, Adriano Jose Pereira, Uri Adrian Prync Flato, Eduardo Pontes dos Reis, Joaquim Edson Vieira, Paulo Sergio Panse Silveira, Edson Amaro Junior

机构 * Einstein Global Advanced Technologies for Equity(埃因斯坦全球先进科技以公平为宗旨) Hospital Israelita Albert Einstein(埃因斯坦医院) Departamento de Pacientes Graves(重症患者部门) Stanford Center for Artificial Intelligence in Medicine and Imaging(斯坦福大学医学与成像人工智能中心) Departmento de Cirurgia(外科部门) Faculdade de Medicina, Universidade de São Paulo(圣保罗大学医学院) Faculdade Israelita de Ciências da Saúde Albert Einstein(埃因斯坦以色列健康科学学院)

专题命中 视觉推理 :multimodal large language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02984 2025-07-29 cs.CL 50%

From Answers to Rationales: Self-Aligning Multimodal Reasoning with Answer-Oriented Chain-of-Thought

Wentao Tan, Qiong Cao, Yibing Zhan, Chao Xue, Changxing Ding

专题命中 视觉推理 :multimodal large language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视觉定位与Grounding 14 篇

2507.15846 2025-07-29 cs.LG cs.AI cs.CL cs.CV cs.HC 82%

GUI-G$^2$: Gaussian Reward Modeling for GUI Grounding

Fei Tang, Zhangxuan Gu, Zhengxi Lu, Xuyang Liu, Shuheng Shen, Changhua Meng, Wen Wang, Wenqi Zhang, Yongliang Shen, Weiming Lu, Jun Xiao, Yueting Zhuang

机构 * Zhejiang University(浙江大学) Ant Group(蚂蚁集团)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20188 2025-07-29 cs.CV 79%

SAViL-Det: Semantic-Aware Vision-Language Model for Multi-Script Text Detection

Mohammed-En-Nadhir Zighem, Abdenour Hadid

机构 * Sorbonne Center for Artificial Intelligence, Sorbonne University Abu Dhabi, UAE(索邦人工智能中心,阿布扎比分校,阿联酋)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11261 2025-07-29 cs.CV 79%

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition

Ronggang Huang, Haoxin Yang, Yan Cai, Xuemiao Xu, Huaidong Zhang, Shengfeng He

机构 * South China University of Technology(华南理工大学) Guangdong Engineering Center for Large Model and GenAI Technology(广东省大模型与生成式人工智能技术工程中心) State Key Laboratory of Subtropical Building and Urban Science(亚热带建筑科学国家重点实验室) Ministry of Education Key Laboratory of Big Data and Intelligent Robot(教育部大数据与智能机器人重点实验室) Guangdong Provincial Key Lab of Computational Intelligence and Cyberspace Information(广东省计算智能与网络信息重点实验室) Singapore Management University(新加坡国立大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19599 2025-07-29 cs.CV 79%

Object-centric Video Question Answering with Visual Grounding and Referring

Haochen Wang, Qirui Chen, Cilin Yan, Jiayin Cai, Xiaolong Jiang, Yao Hu, Weidi Xie, Stratis Gavves

机构 * University of Amsterdam(阿姆斯特丹大学) SAI, Shanghai Jiao Tong University(上海交通大学SAI研究所) Xiaohongshu Inc(小红书公司)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20025 2025-07-29 cs.CV 70%

Region-based Cluster Discrimination for Visual Representation Learning

Yin Xie, Kaicheng Yang, Xiang An, Kun Wu, Yongle Zhao, Weimo Deng, Zimin Ran, Yumeng Wang, Ziyong Feng, Roy Miles, Ismail Elezi, Jiankang Deng

机构 * DeepGlint University of Technology Sydney(悉尼科技大学) Huawei London Research Center(华为伦敦研究中心) Imperial College London(伦敦帝国理工学院)

专题命中 视觉定位与Grounding :grounding(abstract);multimodal large language model(abstract);分类 cs.CV

Comments Accepted as a highlight paper at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20411 2025-07-29 cs.CL 67%

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning

George Ibrahim, Rita Ramos, Yova Kementchedjhieva

机构 * Department of Natural Language Processing, MBZUAI(自然语言处理部门,MBZUAI) INESC-ID, Instituto Superior Técnico, University of Lisbon(INESC-ID,理工学院,里斯本大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);grounding(abstract)

Comments Published as a conference paper at COLM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20913 2025-07-29 cs.CV cs.AI 62%

HAMLET-FFD: Hierarchical Adaptive Multi-modal Learning Embeddings Transformation for Face Forgery Detection

Jialei Cui, Jianwei Du, Yanzhe Li, Lei Gao, Hui Jiang, Chenfu Bao

机构 * Baidu Inc.(百度公司) Southeast University(东南大学) Tsinghua University(清华大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.21226 2025-07-29 cs.CV cs.AI 62%

MemeBLIP2: A novel lightweight multimodal system to detect harmful memes

Jiaqi Liu, Ran Tong, Aowei Shen, Shuzheng Li, Changlin Yang, Lisha Xu

机构 * Mathematics and Statistics Department, University of Texas at Dallas(德克萨斯大学达拉斯分校数学与统计学系)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.AI

Comments 11 pages, 3 figures. Accepted at the First Workshop on Multimodal Knowledge and Language Modeling (MKLM), IJCAI-25

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20397 2025-07-29 cs.CV 57%

VESPA: Towards un(Human)supervised Open-World Pointcloud Labeling for Autonomous Driving

Levente Tempfli, Esteban Rivera, Markus Lienkamp

机构 * Technical University of Munich(慕尼黑技术大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19939 2025-07-29 cs.CV 57%

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs

Jiaze Wang, Rui Chen, Haowang Cui

机构 * Tianjin University(天津大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13092 2025-07-29 cs.CV 57%

EventVAD: Training-Free Event-Aware Video Anomaly Detection

Yihua Shao, Haojin He, Sijie Li, Siyu Chen, Xinwei Long, Fanhu Zeng, Yuxuan Fan, Muyang Zhang, Ziyang Yan, Ao Ma, Xiaochen Wang, Hao Tang, Yan Wang, Shuyan Li

机构 * Peking University(北京大学) Guangdong University of Technology(广东工业大学) The University of Sheffield(谢菲尔德大学) University of Science and Technology Beijing(北京科技大学) Tsinghua University(清华大学) The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) Nanjing University(南京大学) University of Trento(特伦特大学) Queen's University Belfast(贝尔法斯特女王大学)

专题命中 视觉定位与Grounding :MLLM(abstract);分类 cs.CV

Comments Paper was accepted by ACM MM 2025; Code: https://github.com/YihuaJerry/EventVAD

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20306 2025-07-29 math.NA cs.NA 50%

A Hybrid Particle-Continuum Method for Simulating Fast Ice via Subgrid Iceberg Interaction

Carolin Mehlmann, Saskia Kahl

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12370 2025-07-29 cs.CL 50%

Understanding Common Ground Misalignment in Goal-Oriented Dialog: A Case-Study with Ubuntu Chat Logs

Rupak Sarkar, Neha Srikanth, Taylor Hudson, Rachel Rudinger, Claire Bonial, Philip Resnik

机构 * University of Maryland, College Park(马里兰大学 College Park 分校)

专题命中 视觉定位与Grounding :grounding(abstract)

Comments 8 pages

详情

展开后加载摘要…

URL PDF HTML 收藏