arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-08-04 至 2025-08-04 共收录 45 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 9 篇

2506.15170 2025-08-04 cs.CR 50%

From LLMs to MLLMs to Agents: A Survey of Emerging Paradigms in Jailbreak Attacks and Defenses within LLM Ecosystem

Yanxu Mao, Tiehan Cui, Peipei Liu, Datao You, Hongsong Zhu

专题命中 多模态评测 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 多模态Agent 3 篇

2508.00669 2025-08-04 cs.CL cs.AI cs.CV cs.LG 67%

Medical Reasoning in the Era of LLMs: A Systematic Review of Enhancement Techniques and Applications

Wenxuan Wang, Zizhan Ma, Meidan Ding, Shiyi Zheng, Shengyuan Liu, Jie Liu, Jiaming Ji, Wenting Chen, Xiang Li, Linlin Shen, Yixuan Yuan

机构 * Renmin University of China(中国人民大学) The Chinese University of Hong Kong(香港中文大学) Shenzhen University(深圳大学) City University of Hong Kong(香港城市大学) Peking University(北京大学) Massachusetts General Hospital and Harvard Medical School(麻省总医院和哈佛医学院)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.08451 2025-08-04 cs.CV 57%

GUIOdyssey: A Comprehensive Dataset for Cross-App GUI Navigation on Mobile Devices

Quanfeng Lu, Wenqi Shao, Zitao Liu, Lingxiao Du, Fanqing Meng, Boxuan Li, Botong Chen, Siyuan Huang, Kaipeng Zhang, Ping Luo

机构 * Shanghai AI Laboratory(上海人工智能实验室) The University of Hong Kong(香港大学) Nanjing University(南京大学) Shanghai Jiao Tong University(上海交通大学) Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳))

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV

Comments 22 pages, 14 figures, ICCV 2025, a cross-app GUI navigation dataset

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.23454 2025-08-04 cs.HC cs.CY cs.ET cs.GR q-bio.NC 50%

Breaking the mould of Social Mixed Reality - State-of-the-Art and Glossary

Marta Bieńkiewicz, Julia Ayache, Panayiotis Charalambous, Cristina Becchio, Marco Corragio, Bertram Taetz, Francesco De Lellis, Antonio Grotta, Anna Server, Daniel Rammer, Richard Kulpa, Franck Multon, Azucena Garcia-Palacios, Jessica Sutherland, Kathleen Bryson, Stéphane Donikian, Didier Stricker, Benoît Bardy

专题命中 多模态Agent :multi-modal(abstract)

Comments pre-print

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 多模态训练与对齐 8 篇

2508.00447 2025-08-04 cs.CV cs.LG 79%

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text

Anju Rani, Daniel Ortiz-Arroyo, Petar Durdevic

机构 * Department of Energy Technology(能源技术系) Aalborg University(奥尔堡大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments 11 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22268 2025-08-04 cs.IR cs.AI 79%

Multi-modal Relational Item Representation Learning for Inferring Substitutable and Complementary Items

Junting Wang, Chenghuan Guo, Jiao Yang, Yanhui Guo, Yan Gao, Hari Sundaram

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Amazon(亚马逊)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.19651 2025-08-04 cs.LG cs.CL 79%

Unlocking Multi-Modal Potentials for Link Prediction on Dynamic Text-Attributed Graphs

Yuanyuan Xu, Wenjie Zhang, Ying Zhang, Xuemin Lin, Xiwei Xu

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00230 2025-08-04 cs.LG cs.CL cs.CV 62%

Towards Higher Effective Rank in Parameter-efficient Fine-tuning using Khatri--Rao Product

Paul Albert, Frederic Z. Zhang, Hemanth Saratchandran, Anton van den Hengel, Ehsan Abbasnejad

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.CL

Comments To appear in ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00248 2025-08-04 cs.CV 57%

Guided Depth Map Super-Resolution via Multi-Scale Fusion U-shaped Mamba Network

Chenggang Guo, Hao Xu, XianMing Wan

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00053 2025-08-04 cs.CV 57%

A Quality-Guided Mixture of Score-Fusion Experts Framework for Human Recognition

Jie Zhu, Yiyang Su, Minchul Kim, Anil Jain, Xiaoming Liu

机构 * Department of Computer Science and Engineering, Michigan State University(计算机科学与工程系,密歇根州立大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments Accepted to ICCV 2025. 11 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00258 2025-08-04 cs.RO 50%

Topology-Inspired Morphological Descriptor for Soft Continuum Robots

Zhiwei Wu, Siyi Wei, Jiahao Luo, Jinhui Zhang

专题命中 多模态训练与对齐 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.02198 2025-08-04 cs.RO 50%

FalconGym: A Photorealistic Simulation Framework for Zero-Shot Sim-to-Real Vision-Based Quadrotor Navigation

Yan Miao, Will Shen, Sayan Mitra

机构 * Department of Electrical and Computer Engineering, University of Illinois at Urbana-Champaign(电气与计算机工程系,伊利诺伊大学厄巴纳-香槟分校)

专题命中 多模态训练与对齐 :multi-modal(abstract)

Comments Accepted in IROS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 其他多模态 3 篇

2508.00665 2025-08-04 cs.AI cs.HC cs.LG 79%

Transparent Adaptive Learning via Data-Centric Multimodal Explainable AI

Maryam Mosleh, Marie Devlin, Ellis Solaiman

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00778 2025-08-04 cs.CE 71%

τ-Ring: A Smart Ring Platform for Multimodal Physiological and Behavioral Sensing

Jiankai Tang, Zhe He, Mingyu Zhang, Wei Geng, Chengchi Zhou, Weinan Shi, Yuanchun Shi, Yuntao Wang

专题命中 其他多模态 :multimodal(title)

Comments UbiComp 25

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.17667 2025-08-04 cs.RO 50%

Learning Goal-Directed Object Pushing in Cluttered Scenes With Location-Based Attention

Nils Dengler, Juan Del Aguila Ferrandis, João Moura, Sethu Vijayakumar, Maren Bennewitz

机构 * Humanoid Robots Lab, University of Bonn, Germany(波恩大学人形机器人实验室) School of Informatics, The University of Edinburgh, Edinburgh, UK(爱丁堡大学信息学院) The Alan Turing Institute, London, UK(艾伦·图灵研究所) The Lamarr Institute, Bonn, Germany(波恩兰马尔研究所) The Center for Robotics, University of Bonn, Germany(波恩大学机器人中心)

专题命中 其他多模态 :multimodal(abstract)

Comments Accepted to IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)2025

详情

展开后加载摘要…

URL PDF HTML 收藏