arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 46122 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 4657 篇

2602.02786 2026-02-04 cs.LG 50%

LEMON: Local Explanations via Modality-aware OptimizatioN

LEMON:通过模态感知优化实现局部解释

Yu Qin, Phillip Sloan, Raul Santos-Rodriguez, Majid Mirmehdi, Telmo de Menezes e Silva Filho

机构 * University of Bristol(布里斯托大学)

专题命中 图文多模态 :multimodal(abstract)

AI总结 LEMON是一种高效的多模态局部解释框架,通过模态感知优化生成统一解释,减少计算成本并提升解释忠实性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21794 2026-01-30 cs.LG 50%

Knowledge Vector Weakening: Efficient Training-free Unlearning for Large Vision-Language Models

知识向量削弱:高效无训练卸载方法用于大视觉-语言模型

Yejin Kim, Dongjun Hwang, Sungmin Cha, Junsuk Choe

机构 * Sogang University(ソガン大学) New York University(纽约大学)

专题命中 图文多模态 :multimodal(abstract)

AI总结 KVW提出了一种无需训练的高效卸载方法,通过削弱模型中被激活的知识向量,有效防止模型利用有害知识,提升计算效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11908 2026-01-27 cs.RO 50%

Safe Learning for Contact-Rich Robot Tasks: A Survey from Classical Learning-Based Methods to Safe Foundation Models

接触丰富机器人任务的安全学习:从经典学习方法到安全基础模型的综述

Heng Zhang, Rui Dai, Gokhan Solak, Pokuang Zhou, Yu She, Arash Ajoudani

机构 * Human-Robot Interfaces and Interaction Lab, Istituto Italiano di Tecnologia, Genova, Italy(人类-机器人接口与交互实验室,意大利技术研究院,热那亚,意大利) Ph.D. program of national interest in Robotics and Intelligent Machines (DRIM) and Università di Genova, Genoa, Italy(机器人与智能机器国家利益博士项目(DRIM)和热那亚大学,热那亚,意大利) Edwardson School of Industrial Engineering, Purdue University, West Lafayette, IN, USA(工业工程埃德华森学校,普渡大学,西拉法伊斯,美国)

专题命中 图文多模态 :multimodal(abstract)

AI总结 本文综述了接触丰富机器人任务的安全学习方法,探讨了从经典学习方法到安全基础模型的发展,分析了安全探索与执行的关键技术及未来方向。

Comments version 2

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.27240 2026-01-08 cs.LG 50%

FedSM: Robust Semantics-Guided Feature Mixup for Bias Reduction in Federated Learning with Long-Tail Data

FedSM: 用于在长尾数据联邦学习中减少偏见的语义引导特征混合

Jingrui Zhang, Yimeng Xu, Shujie Li, Feng Liang, Haihan Duan, Yanjie Dong, Victor C. M. Leung, Xiping Hu

专题命中 图文多模态 :image-text(abstract)

AI总结 FedSM通过语义引导的特征混合和轻量级分类器重训练,有效减少联邦学习中长尾数据的偏见问题。

Journal ref IEEE Internet of Things Journal, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21450 2025-12-29 cs.LG 50%

RLLaVA: An RL-central Framework for Language and Vision Assistants

RLLaVA: 一种面向语言和视觉助手的强化学习中心框架

Lei Zhao, Zihao Ma, Boyu Lin, Yuhe Liu, Wenjun Wu, Lei Huang

机构 * SKLCCSE, Institute of Artificial Intelligence, Beihang University, Beijing, China(信息与电子技术学院,人工智能研究院,北京航空航天大学,北京,中国) Beijing Advanced Innovation Center for Future Blockchain and Privacy Computing, Beihang University(未来区块链与隐私计算先进创新中心,北京航空航天大学) Hangzhou International Innovation Institute, Beihang University, Hangzhou, China(杭州国际创新研究院,北京航空航天大学,杭州,中国)

专题命中 图文多模态 :multi-modal(abstract)

AI总结 RLLaVA 提出了一种强化学习中心框架,通过解耦算法逻辑与模型架构,实现高效训练和多任务扩展,提升视觉-语言模型性能。

Comments The code is available at https://github.com/TinyLoopX/RLLaVA

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.14661 2025-12-17 cs.AR 50%

Focus: A Streaming Concentration Architecture for Efficient Vision-Language Models

聚焦:一种高效的视觉-语言模型流式集中架构

Chiyue Wei, Cong Guo, Junyao Zhang, Haoxuan Shan, Yifan Xu, Ziyue Zhang, Yudong Liu, Qinsi Wang, Changchun Zhou, Hai "Helen" Li, Yiran Chen

专题命中 图文多模态 :cross-modal(abstract)

AI总结 Focus提出了一种高效的视觉-语言模型流式集中架构,通过分层压缩和细粒度冗余消除,实现2.4倍速度提升和3.3倍能效提升。

Comments HPCA 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11109 2025-12-15 cs.LG 50%

Limits and Gains of Test-Time Scaling in Vision-Language Reasoning

视觉语言推理中测试时扩展的极限与收益

Mohammadjavad Ahmadpour, Amirmahdi Meighani, Payam Taebi, Omid Ghahroodi, Amirmohammad Izadi, Mahdieh Soleymani Baghshah

专题命中 图文多模态 :multimodal(abstract)

AI总结 研究探讨了测试时扩展在视觉语言推理中的效果,发现其在不同任务和模型上表现不一,需根据具体需求定制策略。

Comments Mohammadjavad Ahmadpour and Amirmadhi Meighani contributed equally to this work

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.09927 2025-12-11 cs.RO 50%

Token Expand-Merge: Training-Free Token Compression for Vision-Language-Action Models

令牌扩展-合并:面向视觉-语言-动作模型的免训练 令牌压缩

Yifan Ye, Jiaqi Ma, Jun Cen, Zhihe Lu

机构 * College of Science and Engineering, Hamad Bin Khalifa University(1 科学与工程学院,哈马德·本·卡西姆大学) Mohamed bin Zayed University of Artificial Intelligence(2 摩萨·本·扎耶德人工智能大学) College of Computer Science and Technology, Zhejiang University(3 计算机科学与技术学院,浙江大学)

专题命中 图文多模态 :multimodal(abstract)

AI总结 TEAM-VLA通过动态令牌扩展与合并机制,实现无需训练的视觉-语言-动作模型高效推理,提升速度并保持任务性能。

Comments 8 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.09619 2025-12-11 cs.RO 50%

GLaD: Geometric Latent Distillation for Vision-Language-Action Models

GLaD:面向视觉-语言-动作模型的几何潜在蒸馏

Minghao Guo, Meng Cao, Jiachen Tao, Rongtao Xu, Yan Yan, Xiaodan Liang, Ivan Laptev, Xiaojun Chang

机构 * MBZUAI University of Illinois Chicago(伊利诺伊大学芝加哥分校)

专题命中 图文多模态 :multimodal(abstract)

AI总结 GLaD通过引入几何意识的预训练机制,提升了视觉-语言-动作模型的空间推理和策略泛化能力,无需依赖深度传感器或3D标注。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01331 2025-12-02 cs.RO cs.LG 50%

RobustVLA: Robustness-Aware Reinforcement Post-Training for Vision-Language-Action Models

RobustVLA: 为视觉-语言-动作模型引入鲁棒性感知的强化学习后训练

Hongyin Zhang, Shuo Zhang, Junxi Jin, Qixin Zeng, Runze Li, Donglin Wang

机构 * Westlake University(西湖大学)

专题命中 图文多模态 :multi-modal(abstract)

AI总结 RobustVLA通过引入鲁棒性感知的强化学习后训练方法,提升视觉-语言-动作模型在环境不确定性下的鲁棒性和可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.27256 2025-11-03 cs.LG cs.HC 50%

ECVL-ROUTER: Scenario-Aware Routing for Vision-Language Models

Xin Tang, Youfang Han, Fangfei Gou, Wei Zhao, Xin Meng, Yang Yu, Jinguo Zhang, Yuanchun Shi, Yuntao Wang, Tengxiang Zhang

机构 * Tsinghua University(清华大学) Goertek Inc(歌尔股份有限公司)

专题命中 图文多模态 :multimodal(abstract)

Comments 23 pages, 13 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13317 2025-10-28 cs.LG 50%

Unlabeled Data vs. Pre-trained Knowledge: Rethinking SSL in the Era of Large Models

Song-Lin Lv, Rui Zhu, Tong Wei, Yu-Feng Li, Lan-Zhe Guo

机构 * School of Intelligence Science and Technology, Nanjing University, China(智能科学与技术学院,南京大学) School of Artificial Intelligence, Nanjing University, China(人工智能学院,南京大学) National Key Laboratory for Novel Software Technology, Nanjing University, China(新型软件技术国家重点实验室,南京大学) School of Computer Science and Engineering, Southeast University, Nanjing, China(计算机科学与工程学院,东南大学)

专题命中 图文多模态 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20278 2025-10-24 cs.LG 50%

KCM: KAN-Based Collaboration Models Enhance Pretrained Large Models

Guangyu Dai, Siliang Tang, Yueting Zhuang

机构 * Zhejiang University(浙江大学)

专题命中 图文多模态 :cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16394 2025-10-21 eess.IV 50%

FSAR-Cap: A Fine-Grained Two-Stage Annotated Dataset for SAR Image Captioning

Jinqi Zhang, Lamei Zhang, Bin Zou

专题命中 图文多模态 :image-text(abstract)

Comments 5pages,4figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15508 2025-10-20 cs.LG cs.NA math.NA stat.ML 50%

Theoretical Refinement of CLIP by Utilizing Linear Structure of Optimal Similarity

Naoki Yoshida, Satoshi Hayakawa, Yuhta Takida, Toshimitsu Uesaka, Hiromi Wakaki, Yuki Mitsufuji

机构 * The University of Tokyo(东京大学) Sony Group Corporation(索尼集团公司) Sony AI(索尼人工智能)

专题命中 图文多模态 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14254 2025-10-17 cs.LG 50%

Generalist vs Specialist Time Series Foundation Models: Investigating Potential Emergent Behaviors in Assessing Human Health Using PPG Signals

Saurabh Kataria, Yi Wu, Zhaoliang Chen, Hyunjung Gloria Kwak, Yuhao Xu, Lovely Yeswanth Panchumarthi, Ran Xiao, Jiaying Lu, Ayca Ermis, Anni Zhao, Runze Yan, Alex Federov, Zewen Liu, Xu Wu, Wei Jin, Carl Yang, Jocelyn Grunwell, Stephanie R. Brown, Amit Shah, Craig Jabaley, Tim Buchman, Sivasubramanium V Bhavani, Randall J. Lee, Xiao Hu

机构 * Nell Hodgson Woodruff School of Nursing(Nell Hodgson Woodruff护理学院) School of Computer Science(计算机科学学院) Department of Pediatrics(儿科系) Department of Computer Science(计算机科学系) Department of Epidemiology(流行病学系) Department of Anesthesiology(麻醉学系) Department of Surgery(外科系) Department of Medicine(医学系) School of Medicine(医学院)

专题命中 图文多模态 :cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04710 2025-10-07 cs.LG 50%

ViTs: Teaching Machines to See Time Series Anomalies Like Human Experts

Zexin Wang, Changhua Pei, Yang Liu, Hengyue Jiang, Quan Zhou, Haotian Si, Hang Cui, Jianhui Li, Gaogang Xie, Jingjing Li, Dan Pei

机构 * Computer Network Information Center, Chinese Academy of Sciences(中国科学院计算机网络信息中心) Hangzhou Institute for Advanced Study, University of Chinese Academy of Sciences(中国科学院大学杭州先进研究所) Tsinghua University(清华大学)

专题命中 图文多模态 :image-text(abstract)

Comments 13 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15130 2025-09-29 cs.LG 50%

Few-Shot Adversarial Low-Rank Fine-Tuning of Vision-Language Models

Sajjad Ghiasvand, Haniyeh Ehsani Oskouie, Mahnoosh Alizadeh, Ramtin Pedarsani

专题命中 图文多模态 :cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14967 2025-09-22 cs.RO cs.HC 50%

Affordance-Based Disambiguation of Surgical Instructions for Collaborative Robot-Assisted Surgery

Ana Davila, Jacinto Colan, Yasuhisa Hasegawa

机构 * Nagoya University, Japan(名古屋大学)

专题命中 图文多模态 :multimodal(abstract)

Comments To be presented at the 1st Workshop on Intelligent Cobodied Assistance and Robotic Empowerment (iCARE). 2025 Conference on Robot Learning (CoRL)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06937 2025-09-19 cs.RO 50%

Handle Object Navigation as Weighted Traveling Repairman Problem

Ruimeng Liu, Xinhang Xu, Shenghai Yuan, Lihua Xie

机构 * Centre for Advanced Robotics Technology Innovation (CARTIN), School of Electrical and Electronic Engineering, Nanyang Technological University(先进机器人技术创新中心(CARTIN)、电子与电气工程学院、南洋理工大学)

专题命中 图文多模态 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11065 2025-09-16 cs.SE cs.PL 50%

ViScratch: Using Large Language Models and Gameplay Videos for Automated Feedback in Scratch

Yuan Si, Daming Li, Hanyuan Shi, Jialu Zhang

专题命中 图文多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06768 2025-09-09 cs.RO 50%

Embodied Hazard Mitigation using Vision-Language Models for Autonomous Mobile Robots

Oluwadamilola Sotomi, Devika Kodi, Kiruthiga Chandra Shekar, Aliasghar Arab

机构 * Department of Mechanical and Aerospace Engineering, Tandon School of Engineering, New York University(机械与航空航天工程系,坦顿工程学院,纽约大学) GenAuto.ai by General Autonomy Inc.(General Autonomy Inc. 的 GenAuto.ai)

专题命中 图文多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04162 2025-09-05 cs.AR 50%

Real Time FPGA Based Transformers & VLMs for Vision Tasks: SOTA Designs and Optimizations

Safa Mohammed Sali, Mahmoud Meribout, Ashiyana Abdul Majeed

专题命中 图文多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02805 2025-09-04 cs.LG 50%

Challenges in Understanding Modality Conflict in Vision-Language Models

Trang Nguyen, Jackson Michaels, Madalina Fiterau, David Jensen

机构 * Manning College of Information \& Computer Sciences, University of Massachusetts Amherst, Amherst, U.S.

专题命中 图文多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01361 2025-08-05 cs.RO 50%

VLH: Vision-Language-Haptics Foundation Model

Luis Francisco Moreno Fuentes, Muhammad Haris Khan, Miguel Altamirano Cabrera, Valerii Serpiva, Dmitri Iarchuk, Yara Mahmoud, Issatay Tokmurziyev, Dzmitry Tsetserukou

机构 * Intelligent Space Robotics Laboratory(智能空间机器人实验室) Skolkovo Institute of Science and Technology(斯克尔科沃科学与技术研究所)

专题命中 图文多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10011 2025-08-05 cs.RO 50%

KeyMPs: One-Shot Vision-Language Guided Motion Generation by Sequencing DMPs for Occlusion-Rich Tasks

Edgar Anarossi, Yuhwan Kwon, Hirotaka Tahara, Shohei Tanaka, Keisuke Shirai, Masashi Hamaya, Cristian C. Beltran-Hernandez, Atsushi Hashimoto, Takamitsu Matsubara

机构 * Division of Information Science, Graduate School of Science and Technology, Nara Institute of Science and Technology(信息科学系,科学技术研究生学校,科学技术研究所) Department of Electrical and Electronic Engineering, Faculty of Engineering Science, Kansai University(电气电子工程系,工学科学大学) Department of Electronics, Kobe City College of Technology(电子系,神户市立技术学院) OMRON SINIC X Corporation(OMRON SINIC X公司)

专题命中 图文多模态 :multimodal(abstract)

Comments Published in IEEE Access, Jul 14 2025

Journal ref IEEE Access, vol. 13, pp. 125420-125441, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21053 2025-08-04 cs.LG cs.RO 50%

Flow Matching Policy Gradients

David McAllister, Songwei Ge, Brent Yi, Chung Min Kim, Ethan Weber, Hongsuk Choi, Haiwen Feng, Angjoo Kanazawa

机构 * UC Berkeley(伯克利大学) Max Planck Institute for Intelligent Systems(智能系统马克斯·普朗克研究所)

专题命中 图文多模态 :multimodal(abstract)

Comments See our blog post at https://flowreinforce.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.23859 2025-08-04 astro-ph.IM 50%

radio-llava: Advancing Vision-Language Models for Radio Astronomical Source Analysis

S. Riggi, T. Cecconello, A. Pilzer, S. Palazzo, N. Gupta, A. M. Hopkins, C. Trigilio, G. Umana

专题命中 图文多模态 :multimodal(abstract)

Comments 19 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22304 2025-07-31 cs.CR 50%

Invisible Injections: Exploiting Vision-Language Models Through Steganographic Prompt Embedding

Chetan Pathade

专题命中 图文多模态 :multimodal(abstract)

Comments 14 Pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.08505 2025-07-15 cs.LG 50%

Efficient Deployment of Vision-Language Models on Mobile Devices: A Case Study on OnePlus 13R

Pablo Robin Guerrero, Yueyang Pan, Sanidhya Kashyap

机构 * École Polytechnique Fédérale de Lausanne(联邦理工学院洛桑分校)

专题命中 图文多模态 :MLLM(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏