arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-07-29 至 2025-07-29 共收录 10 信号源:cs.CV, cs.AI, cs.LG

1. 视觉推理 10 篇

2507.20342 2025-07-29 cs.AI cs.RO 83%

VLMPlanner: Integrating Visual Language Models with Motion Planning

Zhipeng Tang, Sha Zhang, Jiajun Deng, Chenjie Wang, Guoliang You, Yuting Huang, Xinrui Lin, Yanyong Zhang

机构 * University of Science and Technology of China(科学技术大学) University of Adelaide(阿德莱德大学) Institute of Artificial Intelligence, Hefei Comprehensive National Science Center(合肥综合性国家科学中心人工智能研究院)

专题命中 视觉推理 :visual language model(title);vision-language model(abstract);VLM(abstract);分类 cs.AI

Comments 8 pages, 3 figures, this paper has been accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.12789 2025-07-29 cs.CV 83%

Efficient Physics Simulation for 3D Scenes via MLLM-Guided Gaussian Splatting

Haoyu Zhao, Hao Wang, Xingyue Zhao, Hao Fei, Hongqiu Wang, Chengjiang Long, Hua Zou

机构 * School of Computer Science, Wuhan University(武汉大学计算机学院) Wuhan National Laboratory for Optoelectronics, Huazhong University of Science and Technology(华中科技大学光电研究院) Meta Reality Lab(Meta现实实验室) Xi’an Jiao Tong University(西安交通大学) National University of Singapore(新加坡国立大学) The Department of Systems Hub, Hong Kong University of Science and Technology (Guangzhou)(香港科技大学系统枢纽部门(广州))

专题命中 视觉推理 :MLLM(title,abstract);visual reasoning(abstract);分类 cs.CV

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20529 2025-07-29 cs.CV cs.AI 73%

Enhancing Spatial Reasoning through Visual and Textual Thinking

Xun Liang, Xin Guo, Zhongming Jin, Weihang Pan, Penghui Shang, Deng Cai, Binbin Lin, Jieping Ye

专题命中 视觉推理 :vision language model(abstract);visual question answering(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.03916 2025-07-29 cs.AI cs.CV 73%

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models

Yifan Jiang, Yibo Xue, Yukun Kang, Pin Zheng, Jian Peng, Feiran Wu, Changliang Xu

机构 * Hangzhou Institute for Advanced Study(杭州先进研究所) University of Chinese Academy of Sciences(中国科学院大学) Alibaba Group(阿里巴巴集团)

专题命中 视觉推理 :vision-language model(abstract);VLM(abstract);分类 cs.CV、cs.AI

Comments Appendix at: https://github.com/PAMPAS-Lab/ANA-PPT-Anamation/blob/main/Appendix.pdf

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.10341 2025-07-29 cs.RO cs.AI cs.LG 73%

Affordance-Guided Reinforcement Learning via Visual Prompting

Olivia Y. Lee, Annie Xie, Kuan Fang, Karl Pertsch, Chelsea Finn

机构 * Stanford University(斯坦福大学) Cornell University(康奈尔大学) University of California, Berkeley(加州大学伯克利分校)

专题命中 视觉推理 :vision-language model(abstract);visual reasoning(abstract);分类 cs.AI、cs.LG

Comments 8 pages, 6 figures. IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.13082 2025-07-29 cs.RO cs.AI cs.CV 62%

Free-form language-based robotic reasoning and grasping

Runyu Jiao, Alice Fasoli, Francesco Giuliari, Matteo Bortolon, Sergio Povoli, Guofeng Mei, Yiming Wang, Fabio Poiesi

机构 * Fondazione Bruno Kessler(布鲁诺·科塞拉基金会) University of Trento(特伦托大学) Istituto Italiano di Tecnologia(意大利技术研究院)

专题命中 视觉推理 :vision-language model(abstract);分类 cs.CV、cs.AI

Comments Accepted to IROS 2025. Project website: https://tev-fbk.github.io/FreeGrasp/

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20509 2025-07-29 cs.RO cs.AI cs.SY eess.SY 57%

LLMs-guided adaptive compensator: Bringing Adaptivity to Automatic Control Systems with Large Language Models

Zhongchao Zhou, Yuxi Lu, Yaonan Zhu, Yifan Zhao, Bin He, Liang He, Wenwen Yu, Yusuke Iwasawa

专题命中 视觉推理 :vision-language model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13629 2025-07-29 cs.CV 57%

FreeQ-Graph: Free-form Querying with Semantic Consistent Scene Graph for 3D Scene Understanding

Chenlu Zhan, Yufei Zhang, Gaoang Wang, Hongwei Wang

机构 * College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院) College of Biomedical Engineering and Instrument Science, Zhejiang University(浙江大学生物医学工程与仪器科学学院) Zhejiang University-University of Illinois Urbana-Champaign Institute, Zhejiang University(浙江大学-伊利诺伊大学厄巴纳-香槟分校联合学院)

专题命中 视觉推理 :grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19885 2025-07-29 cs.CL 50%

Zero-shot Performance of Generative AI in Brazilian Portuguese Medical Exam

Cesar Augusto Madid Truyts, Amanda Gomes Rabelo, Gabriel Mesquita de Souza, Daniel Scaldaferri Lages, Adriano Jose Pereira, Uri Adrian Prync Flato, Eduardo Pontes dos Reis, Joaquim Edson Vieira, Paulo Sergio Panse Silveira, Edson Amaro Junior

机构 * Einstein Global Advanced Technologies for Equity(埃因斯坦全球先进科技以公平为宗旨) Hospital Israelita Albert Einstein(埃因斯坦医院) Departamento de Pacientes Graves(重症患者部门) Stanford Center for Artificial Intelligence in Medicine and Imaging(斯坦福大学医学与成像人工智能中心) Departmento de Cirurgia(外科部门) Faculdade de Medicina, Universidade de São Paulo(圣保罗大学医学院) Faculdade Israelita de Ciências da Saúde Albert Einstein(埃因斯坦以色列健康科学学院)

专题命中 视觉推理 :multimodal large language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02984 2025-07-29 cs.CL 50%

From Answers to Rationales: Self-Aligning Multimodal Reasoning with Answer-Oriented Chain-of-Thought

Wentao Tan, Qiong Cao, Yibing Zhan, Chao Xue, Changxing Ding

专题命中 视觉推理 :multimodal large language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏