arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-08-21 至 2025-08-21 共收录 39 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 3 篇

2408.06569 2025-08-21 cs.CL cs.AI 84%

Social Debiasing for Fair Multi-modal LLMs

Harry Cheng, Yangyang Guo, Qingpei Guo, Ming Yang, Tian Gan, Weili Guan, Liqiang Nie

专题命中 图文多模态 :multi-modal(title,abstract);MLLM(abstract);分类 cs.CL、cs.AI

Comments Project page: https://github.com/xaCheng1996/Social_Debiasing_For_Fair_MLLMs

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14045 2025-08-21 cs.CL cs.CV 62%

From Image Captioning to Visual Storytelling

Admitos Passadakis, Yingjin Song, Albert Gatt

机构 * TUDelft(代尔夫特理工大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL

Comments 16 pages (including references), 5 figures and 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.03984 2025-08-21 cs.CV 57%

CoT-Segmenter: Enhancing OOD Detection in Dense Road Scenes via Chain-of-Thought Reasoning

Jeonghyo Song, Kimin Yun, DaeUng Jo, Jinyoung Kim, Youngjoon Yoo

机构 * Department of Artificial Intelligence, Chung-Ang University(Chung-Ang大学人工智能系) Visual Intelligence Lab., ETRI(ETRI视觉智能实验室) University of Science and Technology (UST)(科技大学(UST)) School of Electronics Engineering, Kyungpook National University(Kyungpook国立大学电子工程学院)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments 6 pages, 3 figures. Accepted at IEEE International Conference on Advanced Visual and Signal-Based Systems 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 音频语音多模态 4 篇

2507.10016 2025-08-21 cs.CR cs.SD eess.AS 83%

The Man Behind the Sound: Demystifying Audio Private Attribute Profiling via Multimodal Large Language Model Agents

Lixu Wang, Kaixiang Yao, Xinfeng Li, Dong Yang, Haoyang Li, Xiaofeng Wang, Wei Dong

机构 * Nanyang Technological University(南洋理工大学) The University of Tokyo(东京大学) Hong Kong Polytechnic University(香港理工大学)

专题命中 音频语音多模态 :multimodal(title,abstract);MLLM(abstract);分类 eess.AS

Comments 22 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14631 2025-08-21 cs.SE 78%

Towards a DSL to Formalize Multimodal Requirements

Marcos Gomez-Vazquez, Jordi Cabot

专题命中 音频语音多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.21154 2025-08-21 cs.HC 78%

Hypergraph Multi-Modal Learning for EEG-based Emotion Recognition in Conversation

Zijian Kang, Yueyang Li, Shengyu Gong, Weiming Zeng, Hongjie Yan, Lingbin Bian, Zhiguo Zhang, Wai Ting Siok, Nizhuan Wang

专题命中 音频语音多模态 :multi-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14130 2025-08-21 eess.AS cs.LG 70%

EmoSLLM: Parameter-Efficient Adaptation of LLMs for Speech Emotion Recognition

Hugo Thimonier, Antony Perzo, Renaud Seguier

机构 * Emobot CentraleSupélec IETR (UMR CNRS 6164)(IETR)

专题命中 音频语音多模态 :multimodal(abstract);multi-modal(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频多模态 4 篇

2508.14395 2025-08-21 cs.HC cs.AI 79%

NoteIt: A System Converting Instructional Videos to Interactable Notes Through Multimodal Video Understanding

Running Zhao, Zhihan Jiang, Xinchen Zhang, Chirui Chang, Handi Chen, Weipeng Deng, Luyao Jin, Xiaojuan Qi, Xun Qian, Edith C. H. Ngai

机构 * The University of Hong Kong(香港大学) The Chinese University of Hong Kong(香港中文大学) Google(谷歌)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

Comments Accepted to UIST 2025. Project website: https://zhaorunning.github.io/NoteIt/

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14609 2025-08-21 cs.CV 57%

AnchorSync: Global Consistency Optimization for Long Video Editing

Zichi Liu, Yinggui Wang, Tao Wei, Chao Ma

机构 * MoE Key Lab of Artificial Intelligence, AI Institute Shanghai Jiao Tong University Shanghai China(人工智能联合实验室,人工智能研究院,上海交通大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments ACM MM 2025; Code is released at https://github.com/VISION-SJTU/AnchorSync

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14442 2025-08-21 cs.HC cs.AI 57%

Detecting Reading-Induced Confusion Using EEG and Eye Tracking

Haojun Zhuang, Dünya Baradari, Nataliya Kosmyna, Arnav Balyan, Constanze Albrecht, Stephanie Chen, Pattie Maes

机构 * University of California, Berkeley(加州大学伯克利分校) MIT Media Lab(MIT媒体实验室) Princeton University(普林斯顿大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14432 2025-08-21 cs.LG 50%

Personalized Counterfactual Framework: Generating Potential Outcomes from Wearable Data

Ajan Subramanian, Amir M. Rahmani

机构 * Dept. of Computer Science, University of California, Irvine(加州大学伊市分校计算机科学系) School of Nursing, University of California, Irvine(加州大学伊市分校护理学院)

专题命中 视频多模态 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 跨模态检索 1 篇

2508.14515 2025-08-21 cs.IR cs.AI 79%

MISS: Multi-Modal Tree Indexing and Searching with Lifelong Sequential Behavior for Retrieval Recommendation

Chengcheng Guo, Junda She, Kuo Cai, Shiyao Wang, Qigen Hu, Qiang Luo, Kun Gai, Guorui Zhou

机构 * Kuaishou Inc.(快手公司)

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.AI

Comments CIKM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 多模态生成 7 篇

2504.02906 2025-08-21 cs.CL cs.AI 84%

Boosting Chart-to-Code Generation in MLLM via Dual Preference-Guided Refinement

Zhihan Zhang, Yixin Cao, Lizi Liao

机构 * Singapore Management University(新加坡国立管理学院) Fudan University(复旦大学)

专题命中 多模态生成 :MLLM(title);multimodal(abstract);cross-modal(abstract);分类 cs.CL、cs.AI

Comments Accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.13034 2025-08-21 cs.CL cs.AI cs.CV cs.LG 67%

Natural Language Generation from Visual Events: State-of-the-Art and Key Open Questions

Aditya K Surikuchi, Raquel Fernández, Sandro Pezzelle

机构 * Institute for Logic, Language and Computation (ILLC), University of Amsterdam(逻辑、语言与计算研究所(ILLC),阿姆斯特丹大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14718 2025-08-21 cs.CL 57%

The Digital Sous Chef -- A Comparative Study on Fine-Tuning Language Models for Recipe Generation

Shubham Pundhir, Ganesh Bagler

机构 * Indraprastha Institute of Information Technology(印度理工学院信息技术研究所)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CL

Comments 8 pages, 4 figures. Code is available at: https://github.com/shubh-iiit/RecipeGPT2-Your-Own-AI-Chef

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14405 2025-08-21 cs.CV 57%

CTA-Flux: Integrating Chinese Cultural Semantics into High-Quality English Text-to-Image Communities

Yue Gong, Shanyuan Liu, Liuzhuozheng Li, Jian Zhu, Bo Cheng, Liebucha Wu, Xiaoyu Wu, Yuhang Ma, Dawei Leng, Yuhui Yin

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14393 2025-08-21 cs.CV 57%

Img2ST-Net: Efficient High-Resolution Spatial Omics Prediction from Whole Slide Histology Images via Fully Convolutional Image-to-Image Learning

Junchao Zhu, Ruining Deng, Junlin Guo, Tianyuan Yao, Juming Xiong, Chongyu Qu, Mengmeng Yin, Yu Wang, Shilin Zhao, Haichun Yang, Daguang Xu, Yucheng Tang, Yuankai Huo

机构 * Department of Computer Science, Vanderbilt University(计算机科学系,范德比尔特大学) Weill Cornell Medicine(韦尔·科恩医学中心) Department of Electrical and Computer Engineering, Vanderbilt University(电气与计算机工程系,范德比尔特大学) Department of Biostatistics, Vanderbilt University Medical Center(生物统计学系,范德比尔特大学医学中心) NVIDIA

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14359 2025-08-21 cs.CV 57%

Taming Transformer for Emotion-Controllable Talking Face Generation

Ziqi Zhang, Cheng Deng

机构 * Xidian University(西安电子科技大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.13799 2025-08-21 math.ST math.PR stat.ML stat.TH 50%

Non-asymptotic bounds for forward processes in denoising diffusions: Ornstein-Uhlenbeck is hard to beat

Miha Brešar, Aleksandar Mijatović

专题命中 多模态生成 :multi-modal(abstract)

Comments new Subsection 4.1 on Kinetic Langevin diffusion as forward process included; to appear in Annals of Applied Probability; 25 pages, 4 figures; see short YouTube videos https://youtu.be/hQvfpwI0UPk?si=tfL-DrH2EzqCuGSN and https://youtu.be/xjzVPOEkl44?si=fq9l3kZFg8eELYG3 explaining the main results and ideas of proofs

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 多模态评测 11 篇

2508.14706 2025-08-21 cs.CL cs.AI cs.CV cs.LG cs.MM 83%

ShizhenGPT: Towards Multimodal LLMs for Traditional Chinese Medicine

Junying Chen, Zhenyang Cai, Zhiheng Liu, Yunjin Yang, Rongsheng Wang, Qingying Xiao, Xiangyi Feng, Zhan Su, Jing Guo, Xiang Wan, Guangjun Yu, Haizhou Li, Benyou Wang

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.13368 2025-08-21 cs.CV cs.LG 79%

MetaWild: A Multimodal Dataset for Animal Re-Identification with Environmental Metadata

Yuzhuo Li, Di Zhao, Tingrui Qiao, Yihao Wu, Bo Pang, Yun Sing Koh

机构 * University of Auckland(奥克兰大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments 7 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.07865 2025-08-21 cs.CV cs.RO 79%

AnoVox: A Benchmark for Multimodal Anomaly Detection in Autonomous Driving

Daniel Bogdoll, Iramm Hamdard, Lukas Namgyu Rößler, Felix Geisler, Muhammed Bayram, Felix Wang, Jan Imhof, Miguel de Campos, Anushervon Tabarov, Yitian Yang, Hanno Gottschalk, J. Marius Zöllner

机构 * FZI Research Center for Information Technology(弗劳恩霍夫研究所信息技术研究中心) Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院) Technical University of Berlin(柏林技术大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments Daniel Bogdoll, Iramm Hamdard, and Lukas Namgyu Rößler contributed equally. Accepted for publication at ECCV 2024 W-CODA workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.11370 2025-08-21 cs.CL 79%

G-LLaVA: Solving Geometric Problem with Multi-Modal Large Language Model

Jiahui Gao, Renjie Pi, Jipeng Zhang, Jiacheng Ye, Wanjun Zhong, Yufei Wang, Lanqing Hong, Jianhua Han, Hang Xu, Zhenguo Li, Lingpeng Kong

机构 * Noah’s Ark Lab(诺亚 Ark 实验室) The University of Hong Kong(香港大学) The Hong Kong University of Science and Technology(香港科技大学)

专题命中 多模态评测 :multi-modal(title);multimodal(abstract);分类 cs.CL

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14058 2025-08-21 cs.IR cs.AI 79%

Dual-Phase Playtime-guided Recommendation: Interest Intensity Exploration and Multimodal Random Walks

Jingmao Zhang, Zhiting Zhao, Yunqi Lin, Jianghong Ma, Tianjun Wei, Haijun Zhang, Xiaofeng Zhang

机构 * Harbin Institute of Technology(哈尔滨工业大学) Nanyang Technological University(南洋理工大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

Comments Accepted for publication at ACM Multimedia (ACM MM) 2025. 10 pages, 5 figures. Code and dataset: https://github.com/zqxwcevrtyui/DP2Rec

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14080 2025-08-21 cs.LG 67%

KnowDR-REC: A Benchmark for Referring Expression Comprehension with Real-World Knowledge

Guanghao Jin, Jingpei Wu, Tianpei Guo, Yiyi Niu, Weidong Zhou, Guoyang Liu

专题命中 多模态评测 :multimodal(abstract);multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.16126 2025-08-21 cs.CV 57%

VisioPhysioENet: Visual Physiological Engagement Detection Network

Alakhsimar Singh, Kanav Goyal, Nischay Verma, Puneet Kumar, Xiaobai Li, Amritpal Singh

机构 * Department of Computer Science(计算机科学系) Center for Machine Vision(机器视觉中心) Signal Analysis, University of Oulu, Finland(信号分析中心,奥卢大学,芬兰) State Key Lab of Blockchain(区块链国家重点实验室) Data Security, Zhejiang University, China(数据安全,浙江大学,中国)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

Comments 35 Pages, 4 figures, 5 Tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14443 2025-08-21 cs.CV 57%

Reconstruction Using the Invisible: Intuition from NIR and Metadata for Enhanced 3D Gaussian Splatting

Gyusam Chang, Tuan-Anh Vu, Vivek Alumootil, Harris Song, Deanna Pham, Sangpil Kim, M. Khalid Jawed

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14104 2025-08-21 cs.SE cs.AI 57%

You Don't Know Until You Click:Automated GUI Testing for Production-Ready Software Evaluation

Yutong Bian, Xianhao Lin, Yupeng Xie, Tianyang Liu, Mingchen Zhuge, Siyuan Lu, Haoming Tang, Jinlin Wang, Jiayi Zhang, Jiaqi Chen, Xiangru Tang, Yongxin Ni, Sirui Hong, Chenglin Wu

专题命中 多模态评测 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14507 2025-08-21 cs.IT math.IT 50%

DeepTelecom: A Digital-Twin Deep Learning Dataset for Channel and MIMO Applications

Bohao Wang, Zehua Jiang, Zhenyu Yang, Chongwen Huang, Yongliang Shen, Siming Jiang, Chen Zhu, Zhaohui Yang, Richeng Jin, Zhaoyang Zhang, Sami Muhaidat, Merouane Debbah

专题命中 多模态评测 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11105 2025-08-21 cs.LG cs.IR 50%

Hybrid-Hierarchical Fashion Graph Attention Network for Compatibility-Oriented and Personalized Outfit Recommendation

Sajjad Saed, Babak Teimourpour

专题命中 多模态评测 :multimodal(abstract)

Comments The corresponding author: Babak Teimourpour

详情

展开后加载摘要…

URL PDF HTML 收藏