arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 46122 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 4657 篇

2507.01925 2025-07-03 cs.RO 50%

A Survey on Vision-Language-Action Models: An Action Tokenization Perspective

Yifan Zhong, Fengshuo Bai, Shaofei Cai, Xuchuan Huang, Zhang Chen, Xiaowei Zhang, Yuanfei Wang, Shaoyang Guo, Tianrui Guan, Ka Nam Lui, Zhiquan Qi, Yitao Liang, Yuanpei Chen, Yaodong Yang

机构 * Institute for AI, Peking University(人工智能研究院,北京大学) PKU-PsiBot Joint Lab(北京大学PsiBot联合实验室) School of Computer Science, Peking University(北京大学计算机学院)

专题命中 图文多模态 :multimodal(abstract)

Comments 70 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13725 2025-06-17 cs.RO 50%

CEED-VLA: Consistency Vision-Language-Action Model with Early-Exit Decoding

Wenxuan Song, Jiayi Chen, Pengxiang Ding, Yuxin Huang, Han Zhao, Donglin Wang, Haoang Li

机构 * HKUST(GZ)(香港科技大学(广州)) Westlake University(西湖大学) Zhejiang University(浙江大学)

专题命中 图文多模态 :multimodal(abstract)

Comments 16 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01844 2025-06-03 cs.LG cs.RO 50%

SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics

Mustafa Shukor, Dana Aubakirova, Francesco Capuano, Pepijn Kooijmans, Steven Palma, Adil Zouitine, Michel Aractingi, Caroline Pascal, Martino Russi, Andres Marafioti, Simon Alibert, Matthieu Cord, Thomas Wolf, Remi Cadene

机构 * Hugging Face Sorbonne University(索邦大学) École Normale Supérieure Paris-Saclay(巴黎萨克雷高等师范学院)

专题命中 图文多模态 :multimodal(abstract)

Comments 24 pages. Code and assets: https://github.com/huggingface/lerobot

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15304 2025-06-02 cs.RO 50%

Saliency-Aware Quantized Imitation Learning for Efficient Robotic Control

Seongmin Park, Hyungmin Kim, Sangwoo Kim, Wonseok Jeon, Juyoung Yang, Byeongwook Jeon, Yoonseon Oh, Jungwook Choi

机构 * Hanyang University(翰阳大学) Hyundai Motor Company(现代汽车公司)

专题命中 图文多模态 :multi-modal(abstract)

Comments arXiv admin note: text overlap with arXiv:2412.01034

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23828 2025-06-02 cs.CR 50%

Spa-VLM: Stealthy Poisoning Attacks on RAG-based VLM

Lei Yu, Yechao Zhang, Ziqi Zhou, Yang Wu, Wei Wan, Minghui Li, Shengshan Hu, Pei Xiaobing, Jing Wang

专题命中 图文多模态 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.04999 2025-05-30 cs.RO cs.LG 50%

DynaMem: Online Dynamic Spatio-Semantic Memory for Open World Mobile Manipulation

Peiqi Liu, Zhanqiu Guo, Mohit Warke, Soumith Chintala, Chris Paxton, Nur Muhammad Mahi Shafiullah, Lerrel Pinto

机构 * New York University(纽约大学) Meta Inc.(Meta公司) Hello Robot Inc.(Hello Robot公司)

专题命中 图文多模态 :multimodal(abstract)

Comments Website: https://dynamem.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.02465 2025-05-14 cs.RO cs.SY eess.SY 50%

UAV-VLRR: Vision-Language Informed NMPC for Rapid Response in UAV Search and Rescue

Yasheerah Yaqoot, Muhammad Ahsan Mustafa, Oleg Sautenkov, Artem Lykov, Valerii Serpiva, Dzmitry Tsetserukou

机构 * Intelligent Space Robotics Laboratory, Center for Digital Engineering, Skolkovo Institute of Science and Technology(智能空间机器人实验室、数字工程中心、斯克尔科沃科学与技术研究所)

专题命中 图文多模态 :multimodal(abstract)

Comments UAV-VLRR

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.12894 2025-05-13 cs.SE cs.RO 50%

VLATest: Testing and Evaluating Vision-Language-Action Models for Robotic Manipulation

Zhijie Wang, Zhehua Zhou, Jiayang Song, Yuheng Huang, Zhan Shu, Lei Ma

机构 * University of Alberta(阿尔伯塔大学) The University of Tokyo(东京大学)

专题命中 图文多模态 :multi-modal(abstract)

Comments To appear in FSE '25 (Proceedings of ACM Software Engineering, Vol. 2, Issue FSE, Article FSE073), 24 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.05937 2025-05-12 cs.HC 50%

MER-CLIP: AU-Guided Vision-Language Alignment for Micro-Expression Recognition

Shifeng Liu, Xinglong Mao, Sirui Zhao, Peiming Li, Tong Xu, Enhong Chen

专题命中 图文多模态 :cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.03460 2025-05-07 cs.RO 50%

LogisticsVLN: Vision-Language Navigation For Low-Altitude Terminal Delivery Based on Agentic UAVs

Xinyuan Zhang, Yonglin Tian, Fei Lin, Yue Liu, Jing Ma, Kornélia Sára Szatmáry, Fei-Yue Wang

机构 * School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所多模态人工智能系统国家重点实验室) Department of Engineering Science, Faculty of Innovation Engineering, Macau University of Science and Technology(澳门科技大学创新工程学院工程科学系) China Ship Research and Development Academy(中国船舶科研 Academy) Obuda University(奥布达大学) State Key Laboratory for Management and Control of Complex Systems, Chinese Academy of Sciences(中国科学院复杂系统管理与控制国家重点实验室)

专题命中 图文多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.03181 2025-05-07 cs.LG 50%

VLM Q-Learning: Aligning Vision-Language Models for Interactive Decision-Making

Jake Grigsby, Yuke Zhu, Michael Ryoo, Juan Carlos Niebles

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校) Salesforce AI Research(Salesforce人工智能研究)

专题命中 图文多模态 :multi-modal(abstract)

Comments SSI-FM Workshop ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.02569 2025-05-06 cs.RO cs.HC 50%

HapticVLM: VLM-Driven Texture Recognition Aimed at Intelligent Haptic Interaction

Muhammad Haris Khan, Miguel Altamirano Cabrera, Dmitrii Iarchuk, Yara Mahmoud, Daria Trinitatova, Issatay Tokmurziyev, Dzmitry Tsetserukou

机构 * Intelligent Space Robotics Laboratory, Center for Digital Engineering, Skolkovo Institute of Science and Technology(智能空间机器人实验室、数字工程中心、斯克尔科沃科学与技术研究所)

专题命中 图文多模态 :multimodal(abstract)

Comments Submitted to IEEE conf

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.09715 2025-05-05 cs.IT cs.GT math.IT 50%

Generative Semantic Communication via Textual Prompts: Latency Performance Tradeoffs

Mengmeng Ren, Li Qiao, Long Yang, Zhen Gao, Jian Chen, Mahdi Boloursaz Mashhadi, Pei Xiao, Rahim Tafazolli, Mehdi Bennis

专题命中 图文多模态 :multi-modal(abstract)

Comments Accepted by IEEE Transactions on Vehicular Technology

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.17171 2025-04-29 cs.HC 50%

Augmenting Captions with Emotional Cues: An AR Interface for Real-Time Accessible Communication

Sunday David Ubur

专题命中 图文多模态 :multimodal(abstract)

Comments Minor correction to references for better citation matching

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.16054 2025-04-23 cs.LG cs.RO 50%

$π_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Physical Intelligence, Kevin Black, Noah Brown, James Darpinian, Karan Dhabalia, Danny Driess, Adnan Esmail, Michael Equi, Chelsea Finn, Niccolo Fusai, Manuel Y. Galliker, Dibya Ghosh, Lachy Groom, Karol Hausman, Brian Ichter, Szymon Jakubczak, Tim Jones, Liyiming Ke, Devin LeBlanc, Sergey Levine, Adrian Li-Bell, Mohith Mothukuri, Suraj Nair, Karl Pertsch, Allen Z. Ren, Lucy Xiaoyang Shi, Laura Smith, Jost Tobias Springenberg, Kyle Stachowicz, James Tanner, Quan Vuong, Homer Walke, Anna Walling, Haohuan Wang, Lili Yu, Ury Zhilinsky

机构 * Physical Intelligence

专题命中 图文多模态 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.14517 2025-03-26 cs.LG cs.CR 50%

TUNI: A Textual Unimodal Detector for Identity Inference in CLIP Models

Songze Li, Ruoxi Cheng, Xiaojun Jia

专题命中 图文多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.07343 2025-02-12 cs.DB 50%

DEG: Efficient Hybrid Vector Search Using the Dynamic Edge Navigation Graph

Ziqi Yin, Jianyang Gao, Pasquale Balsebre, Gao Cong, Cheng Long

专题命中 图文多模态 :image-text(abstract)

Comments Accepted by sigmod 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.15065 2025-02-12 stat.ML cs.LG math.ST stat.TH 50%

The Benefits of Balance: From Information Projections to Variance Reduction

Lang Liu, Ronak Mehta, Soumik Pal, Zaid Harchaoui

专题命中 图文多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.19457 2025-02-10 cs.RO 50%

A Parameter-Efficient Tuning Framework for Language-guided Object Grounding and Robot Grasping

Houjian Yu, Mingen Li, Alireza Rezazadeh, Yang Yang, Changhyun Choi

专题命中 图文多模态 :multimodal(abstract)

Comments Accepted for ICRA 2025. Project page: https://sites.google.com/umn.edu/etog-etrg/home

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.03721 2025-02-07 cs.CR cs.LG 50%

Detecting Backdoor Attacks via Similarity in Semantic Communication Systems

Ziyang Wei, Yili Jiang, Jiaqi Huang, Fangtian Zhong, Sohan Gyawali

专题命中 图文多模态 :image-text(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.13851 2025-01-24 cs.LG 50%

Large Vision-Language Models for Knowledge-Grounded Data Annotation of Memes

Shiling Deng, Serge Belongie, Peter Ebert Christensen

专题命中 图文多模态 :cross-modal(abstract)

Comments 18 pages, 5 figures, 13 tables, GitHub repository: https://github.com/Seefreem/meme_text_retrieval_p1

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.04897 2025-01-10 cs.LG 50%

Online Continual Learning: A Systematic Literature Review of Approaches, Challenges, and Benchmarks

Seyed Amir Bidaki, Amir Mohammadkhah, Kiyan Rezaee, Faeze Hassani, Sadegh Eskandari, Maziar Salahi, Mohammad M. Ghassemi

专题命中 图文多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.12610 2024-11-01 cs.RO 50%

A Joint Modeling of Vision-Language-Action for Target-oriented Grasping in Clutter

Kechun Xu, Shuqi Zhao, Zhongxiang Zhou, Zizhang Li, Huaijin Pi, Yue Wang, Rong Xiong

专题命中 图文多模态 :multi-modal(abstract)

Comments Accepted by ICRA 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.19438 2024-10-15 cs.NE 50%

Neuro-Vision to Language: Enhancing Brain Recording-based Visual Reconstruction and Language Interaction

Guobin Shen, Dongcheng Zhao, Xiang He, Linghao Feng, Yiting Dong, Jihang Wang, Qian Zhang, Yi Zeng

专题命中 图文多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.11838 2024-10-15 cs.RO cs.HC 50%

The Conversation is the Command: Interacting with Real-World Autonomous Robot Through Natural Language

Linus Nwankwo, Elmar Rueckert

专题命中 图文多模态 :multimodal(abstract)

Comments 5 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.01190 2024-10-03 cs.IR cs.DL 50%

Integrating Visual and Textual Inputs for Searching Large-Scale Map Collections with CLIP

Jamie Mahowald, Benjamin Charles Germain Lee

专题命中 图文多模态 :multimodal(abstract)

Comments 18 pages, 7 figures, accepted at the Computational Humanities Research Conference (CHR 2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.13383 2024-07-18 cs.LG 50%

Gradient Projection For Continual Parameter-Efficient Tuning

Jingyang Qiao, Zhizhong Zhang, Xin Tan, Yanyun Qu, Wensheng Zhang, Zhi Han, Yuan Xie

专题命中 图文多模态 :cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.05874 2024-06-11 cs.CR 50%

Stealthy Targeted Backdoor Attacks against Image Captioning

Wenshu Fan, Hongwei Li, Wenbo Jiang, Meng Hao, Shui Yu, Xiao Zhang

专题命中 图文多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.06904 2024-04-11 cs.RO 50%

Vision-Language Model-based Physical Reasoning for Robot Liquid Perception

Wenqiang Lai, Yuan Gao, Tin Lun Lam

专题命中 图文多模态 :multimodal(abstract)

Comments 8 pages, 6 figures, submitted to IROS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.14883 2023-11-28 cs.SI 50%

Predicting Potential School Shooters from Social Media Posts

Alana Cedeno, Rachel Liang, Sheikh Rabiul Islam

专题命中 图文多模态 :multimodal(abstract)

Journal ref IEEE Big Data 2023

详情

展开后加载摘要…

URL PDF HTML 收藏