arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-09-03 至 2025-09-03 共收录 120 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 13 篇

2506.02353 2025-09-03 cs.RO 50%

SAVOR: Skill Affordance Learning from Visuo-Haptic Perception for Robot-Assisted Bite Acquisition

Zhanxin Wu, Bo Ai, Tom Silver, Tapomayukh Bhattacharjee

机构 * Cornell University(康奈尔大学) UC San Diego(南加州大学)

专题命中 视频多模态 :multi-modal(abstract)

Comments Conference on Robot Learning, Oral

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 跨模态检索 15 篇

2509.01360 2025-09-03 cs.CV cs.LG 83%

M3Ret: Unleashing Zero-shot Multimodal Medical Image Retrieval via Self-Supervision

Che Liu, Zheng Jiang, Chengyu Fang, Heng Guo, Yan-Jie Zhou, Jiaqi Qu, Le Lu, Minfeng Xu

机构 * DAMO Academy, Alibaba Group(达摩院,阿里巴巴集团) Imperial College London(帝国理工学院) Tsinghua University(清华大学) Hupan Lab(华潘实验室)

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.16701 2025-09-03 cs.IR cs.CL 83%

AlzheimerRAG: Multimodal Retrieval Augmented Generation for Clinical Use Cases using PubMed articles

Aritra Kumar Lahiri, Qinmin Vivian Hu

机构 * Department of Computer Science, Toronto Metropolitan University(计算机科学系,多伦多 Metropolitan 大学)

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL

Journal ref Machine Learning and Knowledge Extraction. 2025; 7(3):89

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01341 2025-09-03 cs.CV cs.AI 81%

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation

Yunus Serhat Bicakci, Joseph Shingleton, Anahid Basiri

机构 * Vocational School of Social Sciences, Marmara University(马尔马拉大学社会科学职业学校) Geospatial Data Science Group, School of Geographical & Earth Sciences, University of Glasgow(格拉斯哥大学地理与地球科学学院空间数据科学小组)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00353 2025-09-03 cs.CV cs.AI 81%

AQFusionNet: Multimodal Deep Learning for Air Quality Index Prediction with Imagery and Sensor Data

Koushik Ahmed Kushal, Abdullah Al Mamun

机构 * Department of Computer Science Clarkson University(计算机科学系克拉克逊大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 8 pages, 5 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.10435 2025-09-03 cs.CV cs.AI 81%

RAMer: Reconstruction-based Adversarial Model for Multi-party Multi-modal Multi-label Emotion Recognition

Xudong Yang, Yizhang Zhu, Hanfeng Liu, Zeyi Wen, Nan Tang, Yuyu Luo

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) The Hong Kong University of Science and Technology(香港科学与技术大学)

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02017 2025-09-03 cs.IR cs.AI 79%

Empowering Large Language Model for Sequential Recommendation via Multimodal Embeddings and Semantic IDs

Yuhao Wang, Junwei Pan, Xinhang Li, Maolin Wang, Yuan Wang, Yue Liu, Dapeng Liu, Jie Jiang, Xiangyu Zhao

机构 * City University of Hong Kong(香港城市大学) Tencent Inc.(腾讯公司) Tsinghua University(清华大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

Comments CIKM 2025 Full Research Paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00751 2025-09-03 cs.CV 79%

EVENT-Retriever: Event-Aware Multimodal Image Retrieval for Realistic Captions

Dinh-Khoi Vo, Van-Loc Nguyen, Minh-Triet Tran, Trung-Nghia Le

机构 * University of Science, VNU-HCM(越南国家大学科学学院)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

Comments ACM Multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.21595 2025-09-03 cs.CV 79%

PS-ReID: Advancing Person Re-Identification and Precise Segmentation with Multimodal Retrieval

Jincheng Yan, Yun Wang, Xiaoyan Luo, Yu-Wing Tai

机构 * School of Astronautics, Beihang University(北京航空航天大学航天学院) Department of Computer Science, Dartmouth College(达特茅斯学院计算机科学系)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11452 2025-09-03 cs.AI cs.CL cs.HC 62%

Inclusion Arena: An Open Platform for Evaluating Large Foundation Models with Real-World Apps

Kangyu Wang, Hongliang He, Lin Liu, Ruiqi Liang, Zhenzhong Lan, Jianguo Li

机构 * Shanghai Jiao Tong University(上海交通大学) Zhejiang University(浙江大学) Westlake University(西湖大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL、cs.AI

Comments Our platform is publicly accessible at https://www.tbox.cn/about/model-ranking

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01198 2025-09-03 cs.LG cs.AI 57%

Preserving Vector Space Properties in Dimensionality Reduction: A Relationship Preserving Loss Framework

Eddi Weinwurm, Alexander Kovalenko

机构 * Department of Applied Mathematics, Faculty of Information Technology, Czech Technical University in Prague(应用数学系,信息科技学院,布拉格捷克技术大学)

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.03567 2025-09-03 cs.CV 57%

Uncertainty-Aware Prototype Semantic Decoupling for Text-Based Person Search in Full Images

Zengli Luo, Canlong Zhang, Zhixin Li, Zhiwen Wang, Chunrong Wei

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

Comments 9 pages, 5 figures. Accepted by the 18th International Conference on Knowledge Science, Engineering and Management (KSEM 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.05892 2025-09-03 cs.CR cs.AI 57%

PBI-Attack: Prior-Guided Bimodal Interactive Black-Box Jailbreak Attack for Toxicity Maximization

Ruoxi Cheng, Yizhong Ding, Shuirong Cao, Ranjie Duan, Xiaoshuang Jia, Shaowei Yuan, Simeng Qin, Zhiqiang Wang, Xiaojun Jia

机构 * Alibaba Group(阿里巴巴集团) Beijing Electronic Science and Technology Institute(北京电子科技研究所) Nanjing University(南京大学) Renmin University of China(中国人民大学) Northeastern University(东北大学) BraneMatrix AI Nanyang Technological University(南洋理工大学)

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.AI

Comments Prior-Guided Bimodal Interactive Black-Box Jailbreak Attack for Toxicity Maximization

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.02592 2025-09-03 cs.CV 57%

OCR Hinders RAG: Evaluating the Cascading Impact of OCR on Retrieval-Augmented Generation

Junyuan Zhang, Qintong Zhang, Bin Wang, Linke Ouyang, Zichen Wen, Ying Li, Ka-Ho Chow, Conghui He, Wentao Zhang

机构 * Shanghai AI Laboratory(上海人工智能实验室) Peking University(北京大学) The University of HongKong(香港大学) Shanghai Jiaotong University(上海交通大学) Beihang University(北京航空航天大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01184 2025-09-03 cs.IR 50%

MARS: Modality-Aligned Retrieval for Sequence Augmented CTR Prediction

Yutian Xiao, Shukuan Wang, Binhao Wang, Zhao Zhang, Yanze Zhang, Shanqi Liu, Chao Feng, Xiang Li, Fuzhen Zhuang

专题命中 跨模态检索 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09128 2025-09-03 stat.ML cs.LG 50%

A Generalization Theory for Zero-Shot Prediction

Ronak Mehta, Zaid Harchaoui

机构 * University of Washington(华盛顿大学)

专题命中 跨模态检索 :multimodal(abstract)

Comments Published at ICML '25 (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 多模态生成 18 篇

2501.09012 2025-09-03 cs.CV cs.AI cs.CL cs.MM 83%

Multimodal LLMs Can Reason about Aesthetics in Zero-Shot

Ruixiang Jiang, Changwen Chen

机构 * The Hong Kong Polytechnic University(香港理工大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments ACM MM 2025 Camera Ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.01074 2025-09-03 cs.CV cs.GR 79%

Multimodal Conditional 3D Face Geometry Generation

Christopher Otto, Prashanth Chandran, Sebastian Weiss, Markus Gross, Gaspard Zoss, Derek Bradley

机构 * ETH Zürich(苏黎世联邦理工学院) DisneyResearch | Studios(迪士尼研究室)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments Added more evaluation since the first version. Accepted to SMI 2025. Computers & Graphics

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13602 2025-09-03 cs.CV 79%

PersonaVlog: Personalized Multimodal Vlog Generation with Multi-Agent Collaboration and Iterative Self-Correction

Xiaolu Hou, Bing Ma, Jiaxiang Cheng, Xuhua Ren, Kai Yu, Wenyue Li, Tianxiang Zheng, Qinglin Lu

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments Project Page: https://personavlog-paper.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02518 2025-09-03 cs.LG 78%

AnalogCoder-Pro: Unifying Analog Circuit Generation and Optimization via Multi-modal LLMs

Yao Lai, Souradip Poddar, Sungyoung Lee, Guojin Chen, Mengkang Hu, Bei Yu, Ping Luo, David Z. Pan

专题命中 多模态生成 :multi-modal(title);multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.23149 2025-09-03 eess.IV 71%

Towards Interpretable Counterfactual Generation via Multimodal Autoregression

Chenglong Ma, Yuanfeng Ji, Jin Ye, Lu Zhang, Ying Chen, Tianbin Li, Mingjie Li, Junjun He, Hongming Shan

专题命中 多模态生成 :multimodal(title)

Comments MICCAI'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02278 2025-09-03 cs.GR cs.AI cs.MM 62%

Think2Sing: Orchestrating Structured Motion Subtitles for Singing-Driven 3D Head Animation

Zikai Huang, Yihan Zhou, Xuemiao Xu, Cheng Xu, Xiaofen Xing, Jing Qin, Shengfeng He

机构 * School of Computer Science and Engineering, South China University of Technology(华南理工大学计算机科学与工程学院) South China University of Technology(华南理工大学) Guangdong Engineering Center for Large Model and GenAI Technology(广东省大模型与生成式人工智能技术工程中心) Guangdong Provincial Key Lab of Computational Intelligence and Cyberspace Information(广东省计算智能与网络信息重点实验室) Centre for Smart Health, Hong Kong Polytechnic University(香港理工大学智能健康研究中心) CAS-Hong Kong Joint Laboratory for Multimodal Medical Molecular Imaging(中国科学院-香港联合多模态医学分子成像联合实验室) School of Electronic and Information Engineering, South China University of Technology(华南理工大学电子与信息学院) Pazhou Lab, Guangzhou, Guangdong, China(广州琶洲实验室) School of Computing and Information Systems, Singapore Management University(新加坡国立大学计算机与信息系统学院)

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01656 2025-09-03 cs.CV cs.CL 62%

Reinforced Visual Perception with Tools

Zetong Zhou, Dongping Chen, Zixian Ma, Zhihan Hu, Mingyang Fu, Sinan Wang, Yao Wan, Zhou Zhao, Ranjay Krishna

机构 * ONE Lab, HUST(华中科技大学 ONE 实验室) ONE Lab, HUST University of Maryland(华中科技大学 与 马里兰大学 ONE 实验室) University of Washington(华盛顿大学) Zhejiang University(浙江大学)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.CL

Comments Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.12470 2025-09-03 cs.CV 57%

SC-Diff: 3D Shape Completion with Latent Diffusion Models

Simon Schaefer, Juan D. Galvis, Xingxing Zuo, Stefan Leutengger

机构 * Technical University of Munich(慕尼黑技术大学) Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心) ETH Zurich(苏黎世联邦理工学院) MBZUAI(穆桑比克大学人工智能研究所)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments 13 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00428 2025-09-03 cs.CV 57%

Mixture of Global and Local Experts with Diffusion Transformer for Controllable Face Generation

Xuechao Zou, Shun Zhang, Xing Fu, Yue Li, Kai Li, Yushe Cao, Congyan Lang, Pin Tao, Junliang Xing

机构 * Beijing Jiaotong University(北京交通大学) Ant Group(蚂蚁集团) Qinghai University(青海大学) Tsinghua University(清华大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments 14 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18512 2025-09-03 physics.optics cs.CL 57%

Designing across domains with declarative thinking: Insights from the 96-Eyes ptychographic imager project

Antony C Chan

机构 * Consultant, high-throughput microscopy and hardware-accelerated algorithms(咨询顾问,高通量显微镜和硬件加速算法)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CL

Comments Minor changes: resolve HTML rendering issues of sideways tables; Code listing in dark mode. Cite three more journal articles

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.05304 2025-09-03 cs.LG cs.CV 57%

Gaussian Mixture Flow Matching Models

Hansheng Chen, Kai Zhang, Hao Tan, Zexiang Xu, Fujun Luan, Leonidas Guibas, Gordon Wetzstein, Sai Bi

机构 * Stanford University(斯坦福大学) Adobe Research(Adobe研究)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

Comments ICML 2025. Code: https://github.com/Lakonik/GMFlow

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02551 2025-09-03 cs.NI cs.LG 50%

On Transferring, Merging, and Splitting Task-Oriented Network Digital Twins

Zifan Zhang, Minghong Fang, Mingzhe Chen, Yuchen Liu

机构 * Department of Computer Science, North Carolina State University(计算机科学系,北卡罗来纳州立大学) Department of Computer Science and Engineering, University of Louisville(计算机科学与工程系,路易斯维尔大学) Department of Electrical and Computer Engineering and Frost Institute for Data Science and Computing, University of Miami(电气与计算机工程系及弗罗斯特数据科学与计算研究所,迈阿密大学)

专题命中 多模态生成 :multi-modal(abstract)

Comments Accepted by IEEE MobiWac 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01942 2025-09-03 stat.AP gr-qc 50%

Efficient Bayesian Sampling with Langevin Birth-Death Dynamics

Alex Leviyev, Francesco Iacovelli, Aaron Zimmerman

专题命中 多模态生成 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01819 2025-09-03 cs.RO 50%

ManiFlow: A General Robot Manipulation Policy via Consistency Flow Training

Ge Yan, Jiyue Zhu, Yuquan Deng, Shiqi Yang, Ri-Zhao Qiu, Xuxin Cheng, Marius Memmel, Ranjay Krishna, Ankit Goyal, Xiaolong Wang, Dieter Fox

机构 * University of Washington(华盛顿大学) UC San Diego(圣地亚哥大学) Nvidia(英伟达)

专题命中 多模态生成 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏