arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4979 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 4979 篇

2312.07424 2024-02-27 cs.LG cs.AI cs.CV 73%

How Well Does GPT-4V(ision) Adapt to Distribution Shifts? A Preliminary Investigation

Zhongyi Han, Guanglin Zhou, Rundong He, Jindong Wang, Tailin Wu, Yilong Yin, Salman Khan, Lina Yao, Tongliang Liu, Kun Zhang

专题命中 多模态生成 :multimodal(abstract);multimodal foundation model(abstract);分类 cs.CV、cs.AI

Comments added the investigation of Gemini. 66 pages, 41 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.13208 2024-01-24 cs.CV cs.AI 73%

Iterative Adversarial Attack on Image-guided Story Ending Generation

Youze Wang, Wenbo Hu, Richang Hong

专题命中 多模态生成 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.02147 2024-01-05 cs.CL cs.CV 73%

Exploring Boundary of GPT-4V on Marine Analysis: A Preliminary Case Study

Ziqiang Zheng, Yiwei Chen, Jipeng Zhang, Tuan-Anh Vu, Huimin Zeng, Yue Him Wong Tim, Sai-Kit Yeung

专题命中 多模态生成 :multi-modal(abstract);MLLM(abstract);分类 cs.CV、cs.CL

Comments 51 pages, 36 figures, Repository: https://github.com/hkust-vgd/Marine_GPT-4V_Eval

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.17421 2023-10-12 cs.CV cs.CL 73%

The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision)

Zhengyuan Yang, Linjie Li, Kevin Lin, Jianfeng Wang, Chung-Ching Lin, Zicheng Liu, Lijuan Wang

专题命中 多模态生成 :multimodal(abstract);multimodal foundation model(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.11504 2023-06-21 cs.GR cs.CV cs.SD eess.AS 73%

Align, Adapt and Inject: Sound-guided Unified Image Generation

Yue Yang, Kaipeng Zhang, Yuying Ge, Wenqi Shao, Zeyue Xue, Yu Qiao, Ping Luo

专题命中 多模态生成 :multi-modal(abstract);audio-visual(abstract);分类 cs.CV、eess.AS

Comments Tech Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.18752 2023-05-31 cs.CV cs.CL 73%

GPT4Tools: Teaching Large Language Model to Use Tools via Self-instruction

Rui Yang, Lin Song, Yanwei Li, Sijie Zhao, Yixiao Ge, Xiu Li, Ying Shan

专题命中 多模态生成 :multimodal(abstract);multi-modal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.12208 2023-03-23 cs.CV cs.CL cs.LG 73%

MAGVLT: Masked Generative Vision-and-Language Transformer

Sungwoong Kim, Daejin Jo, Donghoon Lee, Jongmin Kim

专题命中 多模态生成 :multimodal(abstract);image-text(abstract);分类 cs.CV、cs.CL

Comments CVPR 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.14491 2022-11-23 cs.CV cs.AI cs.LG 73%

Re-Imagen: Retrieval-Augmented Text-to-Image Generator

Wenhu Chen, Hexiang Hu, Chitwan Saharia, William W. Cohen

专题命中 多模态生成 :multi-modal(abstract);image-text(abstract);分类 cs.CV、cs.AI

Comments 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.01455 2022-05-05 cs.CV cs.AI cs.GR cs.LG 73%

Zero-Shot Text-Guided Object Generation with Dream Fields

Ajay Jain, Ben Mildenhall, Jonathan T. Barron, Pieter Abbeel, Ben Poole

专题命中 多模态生成 :multi-modal(abstract);image-text(abstract);分类 cs.CV、cs.AI

Comments CVPR 2022. 13 pages. Website: https://ajayj.com/dreamfields

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.08910 2021-04-20 cs.CV cs.MM 73%

Towards Open-World Text-Guided Face Image Generation and Manipulation

Weihao Xia, Yujiu Yang, Jing-Hao Xue, Baoyuan Wu

专题命中 多模态生成 :multimodal(abstract);multi-modal(abstract);分类 cs.CV、cs.MM

Comments arXiv admin note: substantial text overlap with arXiv:2012.03308

详情

展开后加载摘要…

URL PDF HTML 收藏
2004.03212 2021-03-23 cs.CV cs.CL 73%

Text-Guided Neural Image Inpainting

Lisai Zhang, Qingcai Chen, Baotian Hu, Shuoran Jiang

专题命中 多模态生成 :multimodal(abstract);image-text(abstract);分类 cs.CV、cs.CL

Comments ACM MM'2020 (Oral). 9 pages, 4 tables, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.10767 2026-08-04 cond-mat.mes-hall cond-mat.mtrl-sci quant-ph 版本更新 71%

Lambert W Function Framework for Graphene Nanoribbon Quantum Sensing: Theory, Verification, and Multi-Modal Applications

基于拉姆伯特W函数框架的石墨烯纳米带量子传感:理论、验证与多模应用

F. A. Chishtie, K. Roberts, N. Jisrawi, S. R. Valluri, A. Soni, P. C. Deshmukh

专题命中 多模态生成 :multi-modal(title)

AI总结 基于拉姆伯特W函数的石墨烯纳米带量子传感框架,通过理论验证和多模应用,实现了灵敏度增强和性能预测。

Comments 21 pages, 10 figures, published version at Results in Engineering journal

Journal ref Results in Engineering, Volume 32, 2026, 112238

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.21450 2026-06-24 cs.CE q-bio.BM 版本更新 71%

CMADiff: Cross-Modal Aligned Diffusion for Controllable Protein Generation

CMADiff: 用于可控蛋白质生成的跨模态对齐扩散

Changjian Zhou, Yuexi Qiu, Jiafeng Li, Jia Song, Wensheng Xiang

专题命中 多模态生成 :cross-modal(title)

AI总结 提出CMADiff框架,通过条件变分自编码器整合理化特征,并利用对比学习模块BioAligner对齐文本描述与蛋白质特征,实现基于文本驱动的可控蛋白质序列生成。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.16618 2026-06-16 eess.SP 新提交 71%

Acoustic, VOC, and Multimodal Stress Source Localization in the Internet of Plants

植物物联网中的声学、VOC和多模态胁迫源定位

Ahmet B. Kilic, Ozgur B. Akan

专题命中 多模态生成 :multimodal(title)

AI总结 提出一种两阶段粗到细定位流程,结合声学到达时间差多边定位和VOC弥散格林函数模型,实现植物网络中胁迫源的空间定位。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.00535 2026-06-02 cs.LG 71%

DREAM-S: Speculative Decoding with Searchable Drafting and Target-Aware Refinement for Multimodal Generation

DREAM-S: 基于可搜索草稿与目标感知精炼的推测解码用于多模态生成

Zining Liu, Yunhai Hu, Tianhua Xia, Bo Bao, Eric Sather, Vithursan Thangarasa, Sai Qian Zhang

机构 * New York University(纽约大学) Cerebras Systems Inc.(Cerebras Systems公司) University of Pennsylvania(宾夕法尼亚大学)

专题命中 多模态生成 :multimodal(title)

AI总结 提出DREAM-S框架,通过神经架构搜索和目标感知超网训练自动优化草稿模型架构与交互策略,结合注意力熵引导的自适应中间特征蒸馏,实现视觉语言模型的高效推测解码,加速比达3.85倍。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.16473 2026-05-19 stat.ML cs.LG cs.NA math.NA math.PR 71%

Dimension-Uniform Discretization Analysis of Preconditioned Annealed Langevin Dynamics for Multimodal Gaussian Mixtures

预处理退火 Langevin 动力学在多模高斯混合中的维度均匀离散化分析

Lorenzo Baldassari, Josselin Garnier, Knut Solna, Maarten V. de Hoop

机构 * University of Basel(巴塞尔大学) Ecole Polytechnique, IP Paris(巴黎高等理工学院) University of California Irvine(加州大学尔湾分校) Rice University(里德大学)

专题命中 多模态生成 :multimodal(title)

AI总结 本文研究了预处理退火 Langevin 动力学在高斯混合中的稳定性问题,通过 Euler-Maruyama 离散化和指数积分方案,证明了在满足特定谱条件时,KL 散度具有维度均匀的上界。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.13502 2026-05-14 eess.SP 71%

A Multi-Modal Intelligent U2V Channel Model for 6G Sensing-Communication Integration

一种面向6G感知通信一体化的多模智能U2V信道模型

Shuo Wang, Zengrui Han, Lu Bai, Xiang Cheng

专题命中 多模态生成 :multi-modal(title)

AI总结 本文提出一种基于三维散射体预测的新型U2V信道模型,通过构建宽车道场景下的高保真混合感知通信集成U2V仿真数据集,设计了3D-SPADE算法,利用LiDAR点云准确预测散射体分布,提升了动态U2V场景的建模精度与计算效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.04421 2026-04-07 cond-mat.supr-con cond-mat.str-el 71%

Multimodal Terahertz Spectroscopy of the Pairing Symmetry and Normal-State Pseudogap in (La,Pr)$_3$Ni$_2$O$_7$ Films

多模态太赫兹光谱学研究(La,Pr)3Ni2O7薄膜中的配对对称性和正常态伪间隙

Shuxiang Xu, Guangdi Zhou, Hao Wang, Tianyi Wu, Wei Wang, Liyu Shi, Dong Wu, Haoliang Huang, Xinbo Wang, Jinfeng Jia, Qi-Kun Xue, Zhuoyu Chen, Tao Dong, Nanlin Wang

专题命中 多模态生成 :multimodal(title)

AI总结 通过结合体敏感太赫兹时域光谱与太赫兹三次谐波生成,研究(La,Pr)3Ni2O7薄膜的超导配对对称性和正常态伪间隙特性,发现有序态与伪间隙的共存与竞争。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02763 2026-01-13 math.ST cs.NA math.NA math.PR stat.CO stat.TH 71%

Time-complexity of sampling from a multimodal distribution using sequential Monte Carlo

使用序贯蒙特卡罗方法从多模分布中采样的时间复杂度

Ruiyu Han, Gautam Iyer, Dejan Slepčev

专题命中 多模态生成 :multimodal(title)

AI总结 该研究探讨了在低温下使用序贯蒙特卡罗方法从多模分布采样的时间复杂度,分析了几何退火调度与兰格-丹德扩散的效率,并得出了收敛性的数学结果。

Comments 65 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26670 2025-10-31 cs.RO 71%

Hybrid Consistency Policy: Decoupling Multi-Modal Diversity and Real-Time Efficiency in Robotic Manipulation

Qianyou Zhao, Yuliang Shen, Xuanran Zhai, Ce Hao, Duidi Wu, Jin Qi, Jie Hu, Qiaojun Yu

机构 * Shanghai Jiao Tong University(上海交通大学) National University of Singapore(国立新加坡大学) Shanghai AI Lab(上海人工智能实验室)

专题命中 多模态生成 :multi-modal(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01622 2025-10-03 cs.RO cs.LG 71%

VFP: Variational Flow-Matching Policy for Multi-Modal Robot Manipulation

Xuanran Zhai, Qianyou Zhao, Qiaojun Yu, Ce Hao

机构 * National University of Singapore(新加坡国立大学) Shanghai Jiao Tong University(上海交通大学) Shanghai AI Lab(上海人工智能实验室)

专题命中 多模态生成 :multi-modal(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.14837 2025-09-18 cs.RO cs.LG 71%

Learning Multimodal Attention for Manipulating Deformable Objects with Changing States

Namiko Saito, Mayu Tatsumi, Ayuna Kubo, Kanata Suzuki, Hiroshi Ito, Shigeki Sugano, Tetsuya Ogata

机构 * Future Robotics Organization, Waseda University(早稻田大学未来机器人组织) Microsoft Research Asia(微软亚洲研究院) Department of Modern Mechanical Engineering, Waseda University(早稻田大学现代机械工程系) Artificial Intelligence Laboratories, Fujitsu Limited(Fujitsu 人工智能实验室) Center for Technology Innovation - Controls and Robotics, Research & Development Group, Hitachi, Ltd.(富士通技术研发集团技术创新中心 - 控制与机器人) Faculty of Science and Engineering, Waseda University(早稻田大学工学部) National Institute of Advanced Science and Technology(国家先进科学研究院)

专题命中 多模态生成 :multimodal(title)

Comments Humanoids2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16320 2025-09-16 astro-ph.IM cs.LG 71%

Learning novel representations of variable sources from multi-modal $\textit{Gaia}$ data via autoencoders

P. Huijse, J. De Ridder, L. Eyer, L. Rimoldini, B. Holl, N. Chornay, J. Roquette, K. Nienartowicz, G. Jevardat de Fombelle, D. J. Fritzewski, A. Kemp, V. Vanlaer, M. Vanrespaille, H. Wang, M. I. Carnerero, C. M. Raiteri, G. Marton, M. Madarász, G. Clementini, P. Gavras, C. Aerts

机构 * Institute of Astronomy, KU Leuven, Celestijnenlaan 200D, B-3001 Leuven, Belgium Millennium Institute of Astrophysics, Nuncio Monse\ nor Sotero Sanz 100, Of. 104, Providencia, Santiago, Chile Department of Astronomy, University of Geneva, Chemin Pegasi 51, 1290 Versoix, Switzerland Department of Astronomy, University of Geneva, Chemin d’Ecogia 16, 1290 Versoix, Switzerland Sednai S\`arl, Geneva, Switzerland INAF - Osservatorio Astrofisico di Torino, Via Osservatorio 20, I-10025 Pino Torinese, Italy Konkoly Observatory, HUN-REN Research Centre for Astronomy Earth Sciences, Konkoly Thege 15-17, 1121 Budapest, Hungary CSFK, MTA Centre of Excellence, Konkoly Thege 15-17, 1121, Budapest, Hungary INAF - Osservatorio di Astrofisica e Scienza dello Spazio di Bologna, Via Piero Gobetti 93/3, Bologna 40129, Italy Starion for European Space Agency, Camino bajo del Castillo, s/n, Urbanizacion Villafranca del Castillo, Villanueva de la Ca \ n ada, 28692 Madrid, Spain Department of Astrophysics, IMAPP, Radboud University Nijmegen, PO Box 9010, 6500 GL Nijmegen, The Netherlands Max Planck Institute for Astronomy, Koenigstuhl 17, 69117 Heidelberg, Germany

专题命中 多模态生成 :multi-modal(title)

Comments Manuscript accepted on Astronomy & Astrophysics, 20 pages, 20 figures, 2 tables

Journal ref A&A 701, A150 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.23149 2025-09-03 eess.IV 71%

Towards Interpretable Counterfactual Generation via Multimodal Autoregression

Chenglong Ma, Yuanfeng Ji, Jin Ye, Lu Zhang, Ying Chen, Tianbin Li, Mingjie Li, Junjun He, Hongming Shan

专题命中 多模态生成 :multimodal(title)

Comments MICCAI'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.23180 2025-07-01 cs.HC 71%

ImprovMate: Multimodal AI Assistant for Improv Actor Training

Riccardo Drago, Yotam Sechayk, Mustafa Doga Dogan, Andrea Sanna, Takeo Igarashi

专题命中 多模态生成 :multimodal(title)

Comments ACM DIS '25

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.16380 2025-06-16 cs.RO cs.HC cs.LG 71%

Learning Multimodal Latent Dynamics for Human-Robot Interaction

Vignesh Prasad, Lea Heitlinger, Dorothea Koert, Ruth Stock-Homburg, Jan Peters, Georgia Chalvatzaki

专题命中 多模态生成 :multimodal(title)

Comments Preprint version of paper accepted at IEEE T-RO. Project website: https://sites.google.com/view/mild-hri

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.17807 2025-03-31 cs.LG cs.NA math.NA stat.ML 71%

Neural Network Approach to Stochastic Dynamics for Smooth Multimodal Density Estimation

Z. Zarezadeh, N. Zarezadeh

专题命中 多模态生成 :multimodal(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.00734 2025-02-18 hep-lat cond-mat.stat-mech cs.LG 71%

Flow-based sampling for multimodal and extended-mode distributions in lattice field theory

Daniel C. Hackett, Chung-Chun Hsieh, Sahil Pontula, Michael S. Albergo, Denis Boyda, Jiunn-Wei Chen, Kai-Feng Chen, Kyle Cranmer, Gurtej Kanwar, Phiala E. Shanahan

专题命中 多模态生成 :multimodal(title)

Comments 38+3 pages, 39 figures. v2: major revisions including new application to extended modes

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.06897 2024-09-12 eess.IV 71%

RATNUS: Rapid, Automatic Thalamic Nuclei Segmentation using Multimodal MRI inputs

Anqi Feng, Zhangxing Bian, Blake E. Dewey, Alexa Gail Colinco, Jiachen Zhuo, Jerry L. Prince

专题命中 多模态生成 :multimodal(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.09461 2024-08-23 cs.LG cond-mat.mtrl-sci physics.chem-ph q-bio.BM 71%

Advancements in Molecular Property Prediction: A Survey of Single and Multimodal Approaches

Tanya Liyaqat, Tanvir Ahmad, Chandni Saxena

专题命中 多模态生成 :multimodal(title)

Comments Submitted to the journal

详情

展开后加载摘要…

URL PDF HTML 收藏