arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4672 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 4672 篇

2510.14254 2025-10-17 cs.LG 50%

Generalist vs Specialist Time Series Foundation Models: Investigating Potential Emergent Behaviors in Assessing Human Health Using PPG Signals

Saurabh Kataria, Yi Wu, Zhaoliang Chen, Hyunjung Gloria Kwak, Yuhao Xu, Lovely Yeswanth Panchumarthi, Ran Xiao, Jiaying Lu, Ayca Ermis, Anni Zhao, Runze Yan, Alex Federov, Zewen Liu, Xu Wu, Wei Jin, Carl Yang, Jocelyn Grunwell, Stephanie R. Brown, Amit Shah, Craig Jabaley, Tim Buchman, Sivasubramanium V Bhavani, Randall J. Lee, Xiao Hu

机构 * Nell Hodgson Woodruff School of Nursing(Nell Hodgson Woodruff护理学院) School of Computer Science(计算机科学学院) Department of Pediatrics(儿科系) Department of Computer Science(计算机科学系) Department of Epidemiology(流行病学系) Department of Anesthesiology(麻醉学系) Department of Surgery(外科系) Department of Medicine(医学系) School of Medicine(医学院)

专题命中 图文多模态 :cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04710 2025-10-07 cs.LG 50%

ViTs: Teaching Machines to See Time Series Anomalies Like Human Experts

Zexin Wang, Changhua Pei, Yang Liu, Hengyue Jiang, Quan Zhou, Haotian Si, Hang Cui, Jianhui Li, Gaogang Xie, Jingjing Li, Dan Pei

机构 * Computer Network Information Center, Chinese Academy of Sciences(中国科学院计算机网络信息中心) Hangzhou Institute for Advanced Study, University of Chinese Academy of Sciences(中国科学院大学杭州先进研究所) Tsinghua University(清华大学)

专题命中 图文多模态 :image-text(abstract)

Comments 13 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15130 2025-09-29 cs.LG 50%

Few-Shot Adversarial Low-Rank Fine-Tuning of Vision-Language Models

Sajjad Ghiasvand, Haniyeh Ehsani Oskouie, Mahnoosh Alizadeh, Ramtin Pedarsani

专题命中 图文多模态 :cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14967 2025-09-22 cs.RO cs.HC 50%

Affordance-Based Disambiguation of Surgical Instructions for Collaborative Robot-Assisted Surgery

Ana Davila, Jacinto Colan, Yasuhisa Hasegawa

机构 * Nagoya University, Japan(名古屋大学)

专题命中 图文多模态 :multimodal(abstract)

Comments To be presented at the 1st Workshop on Intelligent Cobodied Assistance and Robotic Empowerment (iCARE). 2025 Conference on Robot Learning (CoRL)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06937 2025-09-19 cs.RO 50%

Handle Object Navigation as Weighted Traveling Repairman Problem

Ruimeng Liu, Xinhang Xu, Shenghai Yuan, Lihua Xie

机构 * Centre for Advanced Robotics Technology Innovation (CARTIN), School of Electrical and Electronic Engineering, Nanyang Technological University(先进机器人技术创新中心(CARTIN)、电子与电气工程学院、南洋理工大学)

专题命中 图文多模态 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11065 2025-09-16 cs.SE cs.PL 50%

ViScratch: Using Large Language Models and Gameplay Videos for Automated Feedback in Scratch

Yuan Si, Daming Li, Hanyuan Shi, Jialu Zhang

专题命中 图文多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06768 2025-09-09 cs.RO 50%

Embodied Hazard Mitigation using Vision-Language Models for Autonomous Mobile Robots

Oluwadamilola Sotomi, Devika Kodi, Kiruthiga Chandra Shekar, Aliasghar Arab

机构 * Department of Mechanical and Aerospace Engineering, Tandon School of Engineering, New York University(机械与航空航天工程系,坦顿工程学院,纽约大学) GenAuto.ai by General Autonomy Inc.(General Autonomy Inc. 的 GenAuto.ai)

专题命中 图文多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04162 2025-09-05 cs.AR 50%

Real Time FPGA Based Transformers & VLMs for Vision Tasks: SOTA Designs and Optimizations

Safa Mohammed Sali, Mahmoud Meribout, Ashiyana Abdul Majeed

专题命中 图文多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02805 2025-09-04 cs.LG 50%

Challenges in Understanding Modality Conflict in Vision-Language Models

Trang Nguyen, Jackson Michaels, Madalina Fiterau, David Jensen

机构 * Manning College of Information \& Computer Sciences, University of Massachusetts Amherst, Amherst, U.S.

专题命中 图文多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01361 2025-08-05 cs.RO 50%

VLH: Vision-Language-Haptics Foundation Model

Luis Francisco Moreno Fuentes, Muhammad Haris Khan, Miguel Altamirano Cabrera, Valerii Serpiva, Dmitri Iarchuk, Yara Mahmoud, Issatay Tokmurziyev, Dzmitry Tsetserukou

机构 * Intelligent Space Robotics Laboratory(智能空间机器人实验室) Skolkovo Institute of Science and Technology(斯克尔科沃科学与技术研究所)

专题命中 图文多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10011 2025-08-05 cs.RO 50%

KeyMPs: One-Shot Vision-Language Guided Motion Generation by Sequencing DMPs for Occlusion-Rich Tasks

Edgar Anarossi, Yuhwan Kwon, Hirotaka Tahara, Shohei Tanaka, Keisuke Shirai, Masashi Hamaya, Cristian C. Beltran-Hernandez, Atsushi Hashimoto, Takamitsu Matsubara

机构 * Division of Information Science, Graduate School of Science and Technology, Nara Institute of Science and Technology(信息科学系,科学技术研究生学校,科学技术研究所) Department of Electrical and Electronic Engineering, Faculty of Engineering Science, Kansai University(电气电子工程系,工学科学大学) Department of Electronics, Kobe City College of Technology(电子系,神户市立技术学院) OMRON SINIC X Corporation(OMRON SINIC X公司)

专题命中 图文多模态 :multimodal(abstract)

Comments Published in IEEE Access, Jul 14 2025

Journal ref IEEE Access, vol. 13, pp. 125420-125441, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21053 2025-08-04 cs.LG cs.RO 50%

Flow Matching Policy Gradients

David McAllister, Songwei Ge, Brent Yi, Chung Min Kim, Ethan Weber, Hongsuk Choi, Haiwen Feng, Angjoo Kanazawa

机构 * UC Berkeley(伯克利大学) Max Planck Institute for Intelligent Systems(智能系统马克斯·普朗克研究所)

专题命中 图文多模态 :multimodal(abstract)

Comments See our blog post at https://flowreinforce.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.23859 2025-08-04 astro-ph.IM 50%

radio-llava: Advancing Vision-Language Models for Radio Astronomical Source Analysis

S. Riggi, T. Cecconello, A. Pilzer, S. Palazzo, N. Gupta, A. M. Hopkins, C. Trigilio, G. Umana

专题命中 图文多模态 :multimodal(abstract)

Comments 19 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22304 2025-07-31 cs.CR 50%

Invisible Injections: Exploiting Vision-Language Models Through Steganographic Prompt Embedding

Chetan Pathade

专题命中 图文多模态 :multimodal(abstract)

Comments 14 Pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.08505 2025-07-15 cs.LG 50%

Efficient Deployment of Vision-Language Models on Mobile Devices: A Case Study on OnePlus 13R

Pablo Robin Guerrero, Yueyang Pan, Sanidhya Kashyap

机构 * École Polytechnique Fédérale de Lausanne(联邦理工学院洛桑分校)

专题命中 图文多模态 :MLLM(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.01925 2025-07-03 cs.RO 50%

A Survey on Vision-Language-Action Models: An Action Tokenization Perspective

Yifan Zhong, Fengshuo Bai, Shaofei Cai, Xuchuan Huang, Zhang Chen, Xiaowei Zhang, Yuanfei Wang, Shaoyang Guo, Tianrui Guan, Ka Nam Lui, Zhiquan Qi, Yitao Liang, Yuanpei Chen, Yaodong Yang

机构 * Institute for AI, Peking University(人工智能研究院,北京大学) PKU-PsiBot Joint Lab(北京大学PsiBot联合实验室) School of Computer Science, Peking University(北京大学计算机学院)

专题命中 图文多模态 :multimodal(abstract)

Comments 70 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13725 2025-06-17 cs.RO 50%

CEED-VLA: Consistency Vision-Language-Action Model with Early-Exit Decoding

Wenxuan Song, Jiayi Chen, Pengxiang Ding, Yuxin Huang, Han Zhao, Donglin Wang, Haoang Li

机构 * HKUST(GZ)(香港科技大学(广州)) Westlake University(西湖大学) Zhejiang University(浙江大学)

专题命中 图文多模态 :multimodal(abstract)

Comments 16 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01844 2025-06-03 cs.LG cs.RO 50%

SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics

Mustafa Shukor, Dana Aubakirova, Francesco Capuano, Pepijn Kooijmans, Steven Palma, Adil Zouitine, Michel Aractingi, Caroline Pascal, Martino Russi, Andres Marafioti, Simon Alibert, Matthieu Cord, Thomas Wolf, Remi Cadene

机构 * Hugging Face Sorbonne University(索邦大学) École Normale Supérieure Paris-Saclay(巴黎萨克雷高等师范学院)

专题命中 图文多模态 :multimodal(abstract)

Comments 24 pages. Code and assets: https://github.com/huggingface/lerobot

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15304 2025-06-02 cs.RO 50%

Saliency-Aware Quantized Imitation Learning for Efficient Robotic Control

Seongmin Park, Hyungmin Kim, Sangwoo Kim, Wonseok Jeon, Juyoung Yang, Byeongwook Jeon, Yoonseon Oh, Jungwook Choi

机构 * Hanyang University(翰阳大学) Hyundai Motor Company(现代汽车公司)

专题命中 图文多模态 :multi-modal(abstract)

Comments arXiv admin note: text overlap with arXiv:2412.01034

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23828 2025-06-02 cs.CR 50%

Spa-VLM: Stealthy Poisoning Attacks on RAG-based VLM

Lei Yu, Yechao Zhang, Ziqi Zhou, Yang Wu, Wei Wan, Minghui Li, Shengshan Hu, Pei Xiaobing, Jing Wang

专题命中 图文多模态 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.04999 2025-05-30 cs.RO cs.LG 50%

DynaMem: Online Dynamic Spatio-Semantic Memory for Open World Mobile Manipulation

Peiqi Liu, Zhanqiu Guo, Mohit Warke, Soumith Chintala, Chris Paxton, Nur Muhammad Mahi Shafiullah, Lerrel Pinto

机构 * New York University(纽约大学) Meta Inc.(Meta公司) Hello Robot Inc.(Hello Robot公司)

专题命中 图文多模态 :multimodal(abstract)

Comments Website: https://dynamem.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.02465 2025-05-14 cs.RO cs.SY eess.SY 50%

UAV-VLRR: Vision-Language Informed NMPC for Rapid Response in UAV Search and Rescue

Yasheerah Yaqoot, Muhammad Ahsan Mustafa, Oleg Sautenkov, Artem Lykov, Valerii Serpiva, Dzmitry Tsetserukou

机构 * Intelligent Space Robotics Laboratory, Center for Digital Engineering, Skolkovo Institute of Science and Technology(智能空间机器人实验室、数字工程中心、斯克尔科沃科学与技术研究所)

专题命中 图文多模态 :multimodal(abstract)

Comments UAV-VLRR

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.12894 2025-05-13 cs.SE cs.RO 50%

VLATest: Testing and Evaluating Vision-Language-Action Models for Robotic Manipulation

Zhijie Wang, Zhehua Zhou, Jiayang Song, Yuheng Huang, Zhan Shu, Lei Ma

机构 * University of Alberta(阿尔伯塔大学) The University of Tokyo(东京大学)

专题命中 图文多模态 :multi-modal(abstract)

Comments To appear in FSE '25 (Proceedings of ACM Software Engineering, Vol. 2, Issue FSE, Article FSE073), 24 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.05937 2025-05-12 cs.HC 50%

MER-CLIP: AU-Guided Vision-Language Alignment for Micro-Expression Recognition

Shifeng Liu, Xinglong Mao, Sirui Zhao, Peiming Li, Tong Xu, Enhong Chen

专题命中 图文多模态 :cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.03460 2025-05-07 cs.RO 50%

LogisticsVLN: Vision-Language Navigation For Low-Altitude Terminal Delivery Based on Agentic UAVs

Xinyuan Zhang, Yonglin Tian, Fei Lin, Yue Liu, Jing Ma, Kornélia Sára Szatmáry, Fei-Yue Wang

机构 * School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所多模态人工智能系统国家重点实验室) Department of Engineering Science, Faculty of Innovation Engineering, Macau University of Science and Technology(澳门科技大学创新工程学院工程科学系) China Ship Research and Development Academy(中国船舶科研 Academy) Obuda University(奥布达大学) State Key Laboratory for Management and Control of Complex Systems, Chinese Academy of Sciences(中国科学院复杂系统管理与控制国家重点实验室)

专题命中 图文多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.03181 2025-05-07 cs.LG 50%

VLM Q-Learning: Aligning Vision-Language Models for Interactive Decision-Making

Jake Grigsby, Yuke Zhu, Michael Ryoo, Juan Carlos Niebles

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校) Salesforce AI Research(Salesforce人工智能研究)

专题命中 图文多模态 :multi-modal(abstract)

Comments SSI-FM Workshop ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.02569 2025-05-06 cs.RO cs.HC 50%

HapticVLM: VLM-Driven Texture Recognition Aimed at Intelligent Haptic Interaction

Muhammad Haris Khan, Miguel Altamirano Cabrera, Dmitrii Iarchuk, Yara Mahmoud, Daria Trinitatova, Issatay Tokmurziyev, Dzmitry Tsetserukou

机构 * Intelligent Space Robotics Laboratory, Center for Digital Engineering, Skolkovo Institute of Science and Technology(智能空间机器人实验室、数字工程中心、斯克尔科沃科学与技术研究所)

专题命中 图文多模态 :multimodal(abstract)

Comments Submitted to IEEE conf

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.09715 2025-05-05 cs.IT cs.GT math.IT 50%

Generative Semantic Communication via Textual Prompts: Latency Performance Tradeoffs

Mengmeng Ren, Li Qiao, Long Yang, Zhen Gao, Jian Chen, Mahdi Boloursaz Mashhadi, Pei Xiao, Rahim Tafazolli, Mehdi Bennis

专题命中 图文多模态 :multi-modal(abstract)

Comments Accepted by IEEE Transactions on Vehicular Technology

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.17171 2025-04-29 cs.HC 50%

Augmenting Captions with Emotional Cues: An AR Interface for Real-Time Accessible Communication

Sunday David Ubur

专题命中 图文多模态 :multimodal(abstract)

Comments Minor correction to references for better citation matching

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.16054 2025-04-23 cs.LG cs.RO 50%

$π_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Physical Intelligence, Kevin Black, Noah Brown, James Darpinian, Karan Dhabalia, Danny Driess, Adnan Esmail, Michael Equi, Chelsea Finn, Niccolo Fusai, Manuel Y. Galliker, Dibya Ghosh, Lachy Groom, Karol Hausman, Brian Ichter, Szymon Jakubczak, Tim Jones, Liyiming Ke, Devin LeBlanc, Sergey Levine, Adrian Li-Bell, Mohith Mothukuri, Suraj Nair, Karl Pertsch, Allen Z. Ren, Lucy Xiaoyang Shi, Laura Smith, Jost Tobias Springenberg, Kyle Stachowicz, James Tanner, Quan Vuong, Homer Walke, Anna Walling, Haohuan Wang, Lili Yu, Ury Zhilinsky

机构 * Physical Intelligence

专题命中 图文多模态 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏