arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-09-09 至 2025-09-09 共收录 75 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 9 篇

2503.16376 2025-09-09 cs.CV 74%

LaPIG: Cross-Modal Generation of Paired Thermal and Visible Facial Images

Leyang Wang, Joice Lin

机构 * University College London(伦敦大学学院) Xiamen University(厦门大学)

专题命中 多模态生成 :cross-modal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05333 2025-09-09 cs.CV cs.AI 62%

RT-VLM: Re-Thinking Vision Language Model with 4-Clues for Real-World Object Recognition Robustness

Junghyun Park, Tuan Anh Nguyen, Dugki Min

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05321 2025-09-09 cs.CV cs.AI 62%

A Dataset Generation Scheme Based on Video2EEG-SPGN-Diffusion for SEED-VD

Yunfei Guo, Tao Zhang, Wu Huang, Yao Song

机构 * Chengdu Techman Software Co., Ltd.(成都技漫软件有限公司) Sichuan University(四川大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.03726 2025-09-09 cs.LG cs.AI q-bio.BM 57%

Diffusion on language model encodings for protein sequence generation

Viacheslav Meshchaninov, Pavel Strashnov, Andrey Shevtsov, Fedor Nikolaev, Nikita Ivanisenko, Olga Kardymon, Dmitry Vetrov

机构 * Constructor University, Bremen, Germany(Constructor大学,不来梅,德国)

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.01312 2025-09-09 cs.LG 50%

Sampling from Energy-based Policies using Diffusion

Vineet Jain, Tara Akhound-Sadegh, Siamak Ravanbakhsh

专题命中 多模态生成 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 多模态评测 17 篇

2503.13111 2025-09-09 cs.CV cs.CL cs.LG 84%

MM-Spatial: Exploring 3D Spatial Understanding in Multimodal LLMs

Erik Daxberger, Nina Wenzel, David Griffiths, Haiming Gang, Justin Lazarow, Gefen Kohavi, Kai Kang, Marcin Eichner, Yinfei Yang, Afshin Dehghan, Peter Grasch

机构 * Apple(苹果公司)

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.CL

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.04180 2025-09-09 cs.CV 83%

Slice-100K: A Multimodal Dataset for Extrusion-based 3D Printing

Anushrut Jignasu, Kelly O. Marshall, Ankush Kumar Mishra, Lucas Nerone Rillo, Baskar Ganapathysubramanian, Aditya Balu, Chinmay Hegde, Adarsh Krishnamurthy

机构 * Iowa State University(爱荷华州立大学) New York University(纽约大学)

专题命中 多模态评测 :multimodal(title,abstract);multimodal foundation model(abstract);分类 cs.CV

Comments Accepted to NeurIPS 2024. For codebase, see https://github.com/idealab-isu/Slice-100K

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06079 2025-09-09 cs.CL cs.CV 81%

Multimodal Reasoning for Science: Technical Report and 1st Place Solution to the ICML 2025 SeePhys Challenge

Hao Liang, Ruitao Wu, Bohan Zeng, Junbo Niu, Wentao Zhang, Bin Dong

机构 * Peking University(北京大学) Beihang University(北航) Zhongguancun Academy(中关村学院)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.00102 2025-09-09 cs.CV cs.CL cs.LG 81%

ElectroVizQA: How well do Multi-modal LLMs perform in Electronics Visual Question Answering?

Pragati Shuddhodhan Meshram, Swetha Karthikeyan, Bhavya Bhavya, Suma Bhat

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05513 2025-09-09 cs.CV cs.AI cs.RO 81%

OpenEgo: A Large-Scale Multimodal Egocentric Dataset for Dexterous Manipulation

Ahad Jawaid, Yu Xiang

机构 * Department of Computer Science, The University of Texas at Dallas(德克萨斯大学达拉斯分校计算机科学系) Physical Automation, Inc.(Physical Automation 公司)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 4 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02100 2025-09-09 cs.HC cs.CL 79%

E-THER: A Multimodal Dataset for Empathic AI -- Towards Emotional Mismatch Awareness

Sharjeel Tahir, Judith Johnson, Jumana Abu-Khalaf, Syed Afaq Ali Shah

机构 * Centre for AI and ML, Edith Cowan University(人工智能与机器学习中心,埃德温·科温大学) University of Manchester(曼彻斯特大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

Comments 15 pages, 4 figures. Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.13470 2025-09-09 eess.SP cs.CV cs.LG 79%

Multimodal Latent Fusion of ECG Leads for Early Assessment of Pulmonary Hypertension

Mohammod N. I. Suvon, Shuo Zhou, Prasun C. Tripathi, Wenrui Fan, Samer Alabed, Bishesh Khanal, Venet Osmani, Andrew J. Swift, Chen, Chen, Haiping Lu

机构 * School of Computer Science, University of Sheffield(谢菲尔德大学计算机科学学院) Centre for Machine Intelligence, University of Sheffield(谢菲尔德大学智能中心) Department of Electrical & Computer Science Engineering, IITRAM Ahmedabad(印度阿赫迈德亚布德IITRAM电子与计算机科学工程系) Digital Environment Research Institute, Queen Mary University of London(伦敦玛丽女王大学数字环境研究院) Nepal Applied Mathematics and Informatics Institute for research (NAAMII), Nepal(尼泊尔应用数学与信息技术研究所) Department of Computing, Imperial College London(伦敦帝国学院计算机系) School of Medicine and Population Health, University of Sheffield(谢菲尔德大学医学与人口健康学院) Department of Clinical Radiology, Sheffield Teaching Hospitals(谢菲尔德教学医院放射科) National Institute for Health and Care Research (NIHR), Sheffield Biomedical Research Centre(英国国家健康与护理研究所(NIHR)谢菲尔德生物医学研究中心)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05773 2025-09-09 cs.CV 79%

PictOBI-20k: Unveiling Large Multimodal Models in Visual Decipherment for Pictographic Oracle Bone Characters

Zijian Chen, Wenjie Hua, Jinhao Li, Lirong Deng, Fan Du, Tingzhu Chen, Guangtao Zhai

机构 * Shanghai Jiao Tong University(上海交通大学) Shanghai AI Lab(上海人工智能实验室) Wuhan University(武汉大学) East China Normal Unversity(华东师范大学) Macao Polytechnic University(澳门 polytechnic university) Southern University of Science and Technology(南方科技大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments 6 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05330 2025-09-09 cs.AI 79%

MVRS: The Multimodal Virtual Reality Stimuli-based Emotion Recognition Dataset

Seyed Muhammad Hossein Mousavi, Atiye Ilanloo

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17675 2025-09-09 cs.LG 78%

Towards Synthesizing Normative Data for Cognitive Assessments Using Generative Multimodal Large Language Models

Victoria Yan, Honor Chotkowski, Fengran Wang, Xinhui Li, Carl Yang, Jiaying Lu, Runze Yan, Xiao Hu, Alex Fedorov

机构 * The Westminster Schools(韦斯敏斯特学校) Center for Data Science, Nell Hodgson Woodruff School of Nursing, Emory University(数据科学中心、恩莫森护理学院、埃默里大学) Department of Computer Science, Emory University(计算机科学系、埃默里大学) School of Electrical and Computer Engineering, Georgia Institute of Technology(电气与计算机工程学院、佐治亚理工学院)

专题命中 多模态评测 :multimodal(title,abstract)

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16193 2025-09-09 cs.CV cs.MM 76%

LMM4Edit: Benchmarking and Evaluating Multimodal Image Editing with LMMs

Zitong Xu, Huiyu Duan, Bingnan Liu, Guangji Ma, Jiarui Wang, Liu Yang, Shiqi Gao, Xiaoyu Wang, Jia Wang, Xiongkuo Min, Guangtao Zhai, Weisi Lin

机构 * Institute of Image Communication and Network Engineering, Shanghai JiaoTong University(上海交通大学图像通信与网络工程研究所) University of Electronic and Science Technology of China(电子科技大学) Nanyang Technological University(南洋理工大学)

专题命中 多模态评测 :multimodal(title);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06830 2025-09-09 cs.CV cs.LG 74%

Curia: A Multi-Modal Foundation Model for Radiology

Corentin Dancette, Julien Khlaut, Antoine Saporta, Helene Philippe, Elodie Ferreres, Baptiste Callard, Théo Danielou, Léo Alberge, Léo Machado, Daniel Tordjman, Julie Dupuis, Korentin Le Floch, Jean Du Terrail, Mariam Moshiri, Laurent Dercle, Tom Boeken, Jules Gregory, Maxime Ronot, François Legou, Pascal Roux, Marc Sapoval, Pierre Manceron, Paul Hérent

机构 * Raidium Hôpital Européen Georges Pompidou, AP-HP(欧洲乔治·蓬皮杜医院, AP-HP) Université Paris-Cité(巴黎城市大学) INRIA(法国国家信息与自动化技术研究所) INSERM(法国国家卫生与医学研究中心) Beaujon Hospital(贝琼医院) Medical University of South Carolina(南卡罗来纳医科大学) Columbia University Irving Medical Center(哥伦比亚大学伊万斯医学中心) Centre Cardiologique du Nord(北心脑血管中心)

专题命中 多模态评测 :multi-modal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06351 2025-09-09 cs.CV cs.LG 74%

A Multi-Modal Deep Learning Framework for Colorectal Pathology Diagnosis: Integrating Histological and Colonoscopy Data in a Pilot Study

Krithik Ramesh, Ritvik Koneru

专题命中 多模态评测 :multi-modal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03494 2025-09-09 cs.CV 70%

Parameter-Efficient Adaptation of mPLUG-Owl2 via Pixel-Level Visual Prompts for NR-IQA

Yahya Benmahane, Mohammed El Hassouni

机构 * Computer Science Department Faculty of Sciences, Rabat(科学学院计算机科学系,拉巴特) Computer Science Department FLSH(计算机科学系FLSH)

专题命中 多模态评测 :multimodal(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.01766 2025-09-09 cs.CL cs.CV 62%

Support or Refute: Analyzing the Stance of Evidence to Detect Out-of-Context Mis- and Disinformation

Xin Yuan, Jie Guo, Weidong Qiu, Zheng Huang, Shujun Li

专题命中 多模态评测 :multi-modal(abstract);分类 cs.CV、cs.CL

Comments Accepted and published by EMNLP 2023. Details can be found in https://aclanthology.org/2023.emnlp-main.259

Journal ref In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 4268-4280, Singapore. Association for Computational Linguistics

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05747 2025-09-09 cs.CV cs.AI cs.LG cs.MA cs.RO 62%

InterAct: A Large-Scale Dataset of Dynamic, Expressive and Interactive Activities between Two People in Daily Scenarios

Leo Ho, Yinghao Huang, Dafei Qin, Mingyi Shi, Wangpok Tse, Wei Liu, Junichi Yamagishi, Taku Komura

机构 * The University of Hong Kong(香港大学) Centre for Transformative Garment Production(转型服装生产中心) Great Bay University(大湾大学) Dongguan Key Laboratory for Intelligence and Information Technology(东莞智能与信息科技重点实验室) Shandong University(山东大学) National Institute of Informatics(国家信息研究所)

专题命中 多模态评测 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments The first two authors contributed equally to this work

Journal ref Proceedings of the ACM on Computer Graphics and Interactive Techniques 8.4 (2025) 53:1-27

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06456 2025-09-09 cs.CV 57%

Cross3DReg: Towards a Large-scale Real-world Cross-source Point Cloud Registration Benchmark

Zongyi Xu, Zhongpeng Lang, Yilong Chen, Shanshan Zhao, Xiaoshui Huang, Yifan Zuo, Yan Zhang, Qianni Zhang, Xinbo Gao

专题命中 多模态评测 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 多模态Agent 8 篇

2508.05557 2025-09-09 cs.AI 83%

MV-Debate: Multi-view Agent Debate with Dynamic Reflection Gating for Multimodal Harmful Content Detection in Social Media

Rui Lu, Jinhe Bi, Yunpu Ma, Feng Xiao, Yuntao Du, Yijun Tian

专题命中 多模态Agent :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20521 2025-09-09 cs.AI cs.CL 81%

Project Riley: Multimodal Multi-Agent LLM Collaboration with Emotional Reasoning and Voting

Ana Rita Ortigoso, Gabriel Vieira, Daniel Fuentes, Luis Frazão, Nuno Costa, António Pereira

机构 * Polytechnic University of Leiria(莱里亚理工大学)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments 28 pages, 5 figures. Submitted for review to Information Fusion

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06269 2025-09-09 cs.AI 57%

REMI: A Novel Causal Schema Memory Architecture for Personalized Lifestyle Recommendation Agents

Vishal Raman, Vijai Aravindh R, Abhijith Ragav

机构 * Radian Group Inc.(Radian集团) Sri Sivasubramaniya Nadar College Of Engineering(Sri Sivasubramaniya纳达尔工程学院) Amazon(亚马逊)

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

Comments 8 pages, 2 figures, Accepted at the OARS Workshop, KDD 2025, Paper link: https://oars-workshop.github.io/papers/Raman2025.pdf

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03990 2025-09-09 cs.AI 57%

Meta-Policy Reflexion: Reusable Reflective Memory and Rule Admissibility for Resource-Efficient LLM Agent

Chunlong Wu, Ye Luo, Zhibo Qu, Min Wang

机构 * Tongji University(同济大学)

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00054 2025-09-09 cs.RO cs.AI 57%

Robotic Fire Risk Detection based on Dynamic Knowledge Graph Reasoning: An LLM-Driven Approach with Graph Chain-of-Thought

Haimei Pan, Jiyun Zhang, Qinxi Wei, Xiongnan Jin, Chen Xinkai, Jie Cheng

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

Comments We have decided to withdraw this paper as the work is still undergoing further refinement. To ensure the clarity of the results, we prefer to make additional improvements before resubmission. We appreciate the readers' understanding

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.03460 2025-09-09 cs.AI 57%

Multi-Agent Reasoning for Cardiovascular Imaging Phenotype Analysis

Weitong Zhang, Mengyun Qiao, Chengqi Zang, Steven Niederer, Paul M Matthews, Wenjia Bai, Bernhard Kainz

机构 * Department of Computing, Imperial College London, London, UK(帝国理工学院计算机系) Department of Mechanical Engineering, University College London, London, UK(伦敦大学学院机械工程系) Department of Brain Sciences, Imperial College London, London, UK(帝国理工学院脑科学系) Data Science Institute, Imperial College London, London, UK(帝国理工学院数据科学研究所) University of Tokyo, Tokyo, JP(东京大学) National Heart and Lung Institute, Imperial College London, London, UK(帝国理工学院国家心脏和肺研究所) FAU Erlangen-Nürnberg, Erlangen, DE(埃朗根-纽伦堡大学) UK Dementia Research Institute, Imperial College London, London, UK(英国痴呆研究所在伦敦帝国理工学院) Rosalind Franklin Institute, Harwell Science and Innovation Campus, Didcot, UK(罗莎琳德·弗兰克林研究所)

专题命中 多模态Agent :multi-modal(abstract);分类 cs.AI

Comments accepted by MICCAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05338 2025-09-09 cs.RO cs.AI 57%

Plantbot: Integrating Plant and Robot through LLM Modular Agent Networks

Atsushi Masumori, Norihiro Maruyama, Itsuki Doi, johnsmith, Hiroki Sato, Takashi Ikegami

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03864 2025-09-09 cs.AI 57%

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety

Zhenyu Pan, Yiting Zhang, Yutong Zhang, Jianshu Zhang, Haozheng Luo, Yuwei Han, Dennis Wu, Hong-Yu Chen, Philip S. Yu, Manling Li, Han Liu

机构 * Northwestern University(西北大学) University of Illinois at Chicago(伊利诺伊大学香槟分校)

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

Comments accepted by the Trustworthy FMs workshop in ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏