arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4895 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 其他多模态 4895 篇

2408.07666 2026-01-01 cs.LG cs.AI cs.CL cs.CV 67%

Model Merging in LLMs, MLLMs, and Beyond: Methods, Theories, Applications and Opportunities

在大语言模型、多模态大语言模型及更广泛的领域中进行模型融合:方法、理论、应用与机遇

Enneng Yang, Li Shen, Guibing Guo, Xingwei Wang, Xiaochun Cao, Jie Zhang, Dacheng Tao

机构 * Shenzhen Campus of Sun Yat-sen University, China(中山大学深圳校区) Northeastern University China(东北大学) Shenzhen Campus of Sun Yat-sen University China(中山大学深圳校区) Nanyang Technological University Singapore(南洋理工大学) Northeastern University(东北大学) Shenzhen Campus of Sun Yat-sen University(中山大学深圳校区) Nanyang Technological University(南洋理工大学) Institute for Clarity in Documentation Dublin Ohio USA(文档清晰研究所) Inria Paris-Rocquencourt Rocquencourt France(巴黎-罗quentourt研究所) Rajiv Gandhi University Doimukh Arunachal Pradesh India(拉贾·甘地大学) Tsinghua University Haidian Qu Beijing Shi China(清华大学) Palmer Research Laboratories San Antonio Texas USA(帕勒研究中心) Institute for Clarity in Documentation(文档清晰研究所) Inria Paris-Rocquencourt(巴黎-罗quentourt研究所) Rajiv Gandhi University(拉贾·甘地大学) Tsinghua University(清华大学) Palmer Research Laboratories(帕勒研究中心)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 本文综述了模型融合的方法、理论、应用及未来方向,提出新的分类方法并探讨其在多个机器学习领域的应用及挑战。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.06850 2025-09-16 cs.NE 67%

Visual Evolutionary Optimization on Graph-Structured Combinatorial Problems with MLLMs: A Case Study of Influence Maximization

Jie Zhao, Kang Hao Cheong

专题命中 其他多模态 :multimodal(abstract);MLLM(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05182 2025-08-22 cs.IR 67%

On the Comprehensibility of Multi-structured Financial Documents using LLMs and Pre-processing Tools

Shivani Upadhyay, Messiah Ataey, Syed Shariyar Murtaza, Yifan Nie, Jimmy Lin

专题命中 其他多模态 :multi-modal(abstract);MLLM(abstract)

Comments 15 pages, 5 figures, 9 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.20448 2025-07-22 cs.LG 67%

Knockout: A simple way to handle missing inputs

Minh Nguyen, Batuhan K. Karaman, Heejong Kim, Alan Q. Wang, Fengbei Liu, Mert R. Sabuncu

机构 * Cornell University(康奈尔大学) Weill Cornell Medicine(韦尔医学院)

专题命中 其他多模态 :multimodal(abstract);multi-modal(abstract)

Comments Accepted at TMLR

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11737 2025-06-16 cs.CV cs.CL cs.MM 67%

Quizzard@INOVA Challenge 2025 -- Track A: Plug-and-Play Technique in Interleaved Multi-Image Model

Dinh Viet Cuong, Hoang-Bao Le, An Pham Ngoc Nguyen, Liting Zhou, Cathal Gurrin

机构 * School of Computing, Dublin City University(都柏林城市大学计算机学院) ADAPT Centre(ADAPT中心)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.13852 2025-05-22 cs.CL cs.AI cs.CV cs.LG 67%

Retrospective Learning from Interactions

Zizhao Chen, Mustafa Omer Gul, Yiwei Chen, Gloria Geng, Anne Wu, Yoav Artzi

机构 * Department of Computer Science and Cornell Tech, Cornell University(计算机科学系和康奈尔科技学院,康奈尔大学)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.10154 2025-05-13 cs.RO 67%

A Clinical Tuning Framework for Continuous Kinematic and Impedance Control of a Powered Knee-Ankle Prosthesis

Emma Reznick, T. Kevin Best, Robert Gregg

机构 * IEEE

专题命中 其他多模态 :multimodal(abstract);multi-modal(abstract)

Comments Published in IEEE JTEHM. IEEE Journal of Translational Engineering in Health and Medicine (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.09829 2025-04-25 cs.RO cs.LG cs.SY eess.SY 67%

SE(3)-Equivariant Robot Learning and Control: A Tutorial Survey

Joohwan Seo, Soochul Yoo, Junwoo Chang, Hyunseok An, Hyunwoo Ryu, Soomi Lee, Arvind Kruthiventy, Jongeun Choi, Roberto Horowitz

专题命中 其他多模态 :multimodal(abstract);multi-modal(abstract)

Comments Accepted to International Journcal of Control, Automation and Systems (IJCAS)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.12316 2025-04-18 cs.CL cs.AI cs.CV 67%

Data Metabolism: An Efficient Data Design Schema For Vision Language Model

Jingyuan Zhang, Hongzhi Zhang, Zhou Haonan, Chenxi Sun, Xingguang ji, Jiakang Wang, Fanheng Kong, Yahui Liu, Qi Wang, Fuzheng Zhang

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments To be presented at ICLR 2025, First Workshop on Open Science for Foundation Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.01759 2025-03-12 cs.SI cs.CL cs.CV cs.MM 67%

VGA: Vision and Graph Fused Attention Network for Rumor Detection

Lin Bai, Caiyan Jia, Ziying Song, Chaoqun Cui

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.04353 2025-02-10 cs.CL cs.AI cs.CV 67%

CognArtive: Large Language Models for Automating Art Analysis and Decoding Aesthetic Elements

Afshin Khadangi, Amir Sartipi, Igor Tchappi, Gilbert Fridgen

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.12735 2024-12-18 cs.CV cs.AI cs.CL 67%

GIRAFFE: Design Choices for Extending the Context Length of Visual Language Models

Mukai Li, Lei Li, Shansan Gong, Qi Liu

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Working in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.18908 2024-12-04 cs.HC 67%

DuetML: Human-LLM Collaborative Machine Learning Framework for Non-Expert Users

Wataru Kawabe, Yusuke Sugano

专题命中 其他多模态 :multimodal(abstract);MLLM(abstract)

Comments 22 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.23883 2024-11-01 cs.CL cs.AI cs.LG cs.MM 67%

'No' Matters: Out-of-Distribution Detection in Multimodality Long Dialogue

Rena Gao, Xuetong Wu, Siwen Luo, Caren Han, Feng Liu

专题命中 其他多模态 :multimodal(abstract);分类 cs.CL、cs.AI、cs.MM

Comments 16 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.08424 2024-10-28 cs.RO cs.LG 67%

Conditional Neural Expert Processes for Learning Movement Primitives from Demonstration

Yigit Yildirim, Emre Ugur

专题命中 其他多模态 :multimodal(abstract);multi-modal(abstract)

Comments This work has been submitted to the IEEE RA-L for possible publication. Submitted to Robotics and Automation Letters on July 5, 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.18931 2024-09-30 cs.SI cs.CY 67%

Social Media Bot Policies: Evaluating Passive and Active Enforcement

Kristina Radivojevic, Christopher McAleer, Catrell Conley, Cormac Kennedy, Paul Brenner

专题命中 其他多模态 :multimodal(abstract);multimodal foundation model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.16567 2024-08-30 cs.RO cs.LG 67%

Identifying Terrain Physical Parameters from Vision -- Towards Physical-Parameter-Aware Locomotion and Navigation

Jiaqi Chen, Jonas Frey, Ruyi Zhou, Takahiro Miki, Georg Martius, Marco Hutter

专题命中 其他多模态 :multi-modal(abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.10993 2024-05-21 q-bio.QM 67%

No winners: Performance of lung cancer prediction models depends on screening-detected, incidental, and biopsied pulmonary nodule use cases

Thomas Z. Li, Kaiwen Xu, Aravind Krishnan, Riqiang Gao, Michael N. Kammer, Sanja Antic, David Xiao, Michael Knight, Yency Martinez, Rafael Paez, Robert J. Lentz, Stephen Deppen, Eric L. Grogan, Thomas A. Lasko, Kim L. Sandler, Fabien Maldonado, Bennett A. Landman

专题命中 其他多模态 :multimodal(abstract);multi-modal(abstract)

Comments Submitted to Radiology: AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.17971 2024-04-03 cs.CV cs.AI cs.CL 67%

All in an Aggregated Image for In-Image Learning

Lei Wang, Wanyu Xu, Zhiqiang Hu, Yihuai Lan, Shan Dong, Hao Wang, Roy Ka-Wei Lee, Ee-Peng Lim

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.12383 2023-12-20 cs.AI cs.CL cs.CV 67%

Visual AI and Linguistic Intelligence Through Steerability and Composability

David Noever, Samantha Elizabeth Miller Noever

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.11441 2023-11-07 cs.CV cs.AI cs.CL cs.HC 67%

Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V

Jianwei Yang, Hao Zhang, Feng Li, Xueyan Zou, Chunyuan Li, Jianfeng Gao

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.18842 2023-05-31 cs.CL cs.AI cs.CV 67%

Generate then Select: Open-ended Visual Question Answering Guided by World Knowledge

Xingyu Fu, Sheng Zhang, Gukyeong Kwon, Pramuditha Perera, Henghui Zhu, Yuhao Zhang, Alexander Hanbo Li, William Yang Wang, Zhiguo Wang, Vittorio Castelli, Patrick Ng, Dan Roth, Bing Xiang

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted to ACL 2023 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.03701 2022-10-10 cs.RO 67%

VIRDO++: Real-World, Visuo-tactile Dynamics and Perception of Deformable Objects

Youngsun Wi, Andy Zeng, Pete Florence, Nima Fazeli

专题命中 其他多模态 :multimodal(abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.09559 2022-05-31 stat.ME stat.CO stat.ML 67%

Continuously-Tempered PDMP Samplers

Matthew Sutton, Robert Salomone, Augustin Chevallier, Paul Fearnhead

专题命中 其他多模态 :multimodal(abstract);multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.01295 2022-02-16 cs.CV cs.AI cs.CL 67%

Information Symmetry Matters: A Modal-Alternating Propagation Network for Few-Shot Learning

Zhong Ji, Zhishen Hou, Xiyao Liu, Yanwei Pang, Jungong Han

专题命中 其他多模态 :cross-modal(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.11403 2021-10-25 cs.CV cs.AI cs.CL cs.LG 67%

SCENIC: A JAX Library for Computer Vision Research and Beyond

Mostafa Dehghani, Alexey Gritsenko, Anurag Arnab, Matthias Minderer, Yi Tay

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.08013 2021-09-17 cs.CV cs.CL cs.LG cs.MM 67%

Detecting Propaganda Techniques in Memes

Dimitar Dimitrov, Bishr Bin Ali, Shaden Shaar, Firoj Alam, Fabrizio Silvestri, Hamed Firooz, Preslav Nakov, Giovanni Da San Martino

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.MM

Comments propaganda, disinformation, fake news, memes, multimodality. arXiv admin note: text overlap with arXiv:2105.09284

Journal ref ACL-2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.08614 2021-04-20 cs.SD cs.AI cs.CL cs.LG cs.RO eess.AS 67%

Cetacean Translation Initiative: a roadmap to deciphering the communication of sperm whales

Jacob Andreas, Gašper Beguš, Michael M. Bronstein, Roee Diamant, Denley Delaney, Shane Gero, Shafi Goldwasser, David F. Gruber, Sarah de Haas, Peter Malkin, Roger Payne, Giovanni Petri, Daniela Rus, Pratyusha Sharma, Dan Tchernov, Pernille Tønnesen, Antonio Torralba, Daniel Vogt, Robert J. Wood

专题命中 其他多模态 :multimodal(abstract);分类 cs.CL、cs.AI、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2009.10370 2020-09-23 cs.CV cs.AI cs.HC cs.MM 67%

Visual Methods for Sign Language Recognition: A Modality-Based Review

Bassem Seddik, Najoua Essoukri Ben Amara

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV、cs.AI、cs.MM

Comments This survey paper is accepted as Springer book chapter, currently under edition

详情

展开后加载摘要…

URL PDF HTML 收藏
1909.09070 2019-09-20 cs.AI cs.CL cs.CV 67%

Look, Read and Enrich. Learning from Scientific Figures and their Captions

Jose Manuel Gomez-Perez, Raul Ortega

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted in the 10th International Conference on Knowledge capture (K-CAP 2019)

详情

展开后加载摘要…

URL PDF HTML 收藏