arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4975 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 4975 篇

2509.17850 2025-09-23 cs.RO 50%

SocialTraj: Two-Stage Socially-Aware Trajectory Prediction for Autonomous Driving via Conditional Diffusion Model

Xiao Zhou, Zengqi Peng, Jun Ma

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

专题命中 多模态生成 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17790 2025-09-23 physics.med-ph eess.IV 50%

Conditional Diffusion Models for CT Image Synthesis from CBCT: A Systematic Review

Alzahra Altalib, Chunhui Li, Alessandro Perelli

专题命中 多模态生成 :multi-modal(abstract)

Comments 36 pages, 8 figures, 3 tables, submitted to Elsevier Computerized Medical Imaging and Graphics

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17080 2025-09-23 cs.RO 50%

CoPlanner: An Interactive Motion Planner with Contingency-Aware Diffusion for Autonomous Driving

Ruiguo Zhong, Ruoyu Yao, Pei Liu, Xiaolong Chen, Rui Yang, Jun Ma

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

专题命中 多模态生成 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14711 2025-09-19 eess.SP 50%

LLM4MG: Adapting Large Language Model for Multipath Generation via Synesthesia of Machines

Ziwei Huang, Shiliang Lu, Lu Bai, Xuesong Cai, Xiang Cheng

专题命中 多模态生成 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08775 2025-09-12 cs.RO 50%

Joint Model-based Model-free Diffusion for Planning with Constraints

Wonsuhk Jung, Utkarsh A. Mishra, Nadun Ranawaka Arachchige, Yongxin Chen, Danfei Xu, Shreyas Kousik

机构 * Georgia Institute of Technology(佐治亚理工学院)

专题命中 多模态生成 :multi-modal(abstract)

Comments The first two authors contributed equally. Last three authors advised equally. Accepted to CoRL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.01312 2025-09-09 cs.LG 50%

Sampling from Energy-based Policies using Diffusion

Vineet Jain, Tara Akhound-Sadegh, Siamak Ravanbakhsh

专题命中 多模态生成 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02551 2025-09-03 cs.NI cs.LG 50%

On Transferring, Merging, and Splitting Task-Oriented Network Digital Twins

Zifan Zhang, Minghong Fang, Mingzhe Chen, Yuchen Liu

机构 * Department of Computer Science, North Carolina State University(计算机科学系,北卡罗来纳州立大学) Department of Computer Science and Engineering, University of Louisville(计算机科学与工程系,路易斯维尔大学) Department of Electrical and Computer Engineering and Frost Institute for Data Science and Computing, University of Miami(电气与计算机工程系及弗罗斯特数据科学与计算研究所,迈阿密大学)

专题命中 多模态生成 :multi-modal(abstract)

Comments Accepted by IEEE MobiWac 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01942 2025-09-03 stat.AP gr-qc 50%

Efficient Bayesian Sampling with Langevin Birth-Death Dynamics

Alex Leviyev, Francesco Iacovelli, Aaron Zimmerman

专题命中 多模态生成 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01819 2025-09-03 cs.RO 50%

ManiFlow: A General Robot Manipulation Policy via Consistency Flow Training

Ge Yan, Jiyue Zhu, Yuquan Deng, Shiqi Yang, Ri-Zhao Qiu, Xuxin Cheng, Marius Memmel, Ranjay Krishna, Ankit Goyal, Xiaolong Wang, Dieter Fox

机构 * University of Washington(华盛顿大学) UC San Diego(圣地亚哥大学) Nvidia(英伟达)

专题命中 多模态生成 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00753 2025-09-03 stat.ME cs.LG stat.AP stat.CO stat.ML 50%

FBMS: An R Package for Flexible Bayesian Model Selection and Model Averaging

Florian Frommlet, Jon Lachmann, Geir Storvik, Aliaksandr Hubin

专题命中 多模态生成 :multi-modal(abstract)

Comments 69 pages, 5 tables, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00098 2025-09-03 physics.ins-det cond-mat.mtrl-sci 50%

Operating advanced scientific instruments with AI agents that learn on the job

Aikaterini Vriza, Michael H. Prince, Tao Zhou, Henry Chan, Mathew J. Cherukara

专题命中 多模态生成 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21660 2025-09-03 cs.LG 50%

PreGenie: An Agentic Framework for High-quality Visual Presentation Generation

Xiaojie Xu, Xinli Xu, Sirui Chen, Haoyu Chen, Fan Zhang, Ying-Cong Chen

机构 * The Hong Kong University of Science and Technology(Guangzhou)(香港科技大学(广州)) The Hong Kong University of Science and Technology(香港科技大学) Shanghai AI Laboratory(上海人工智能实验室)

专题命中 多模态生成 :multimodal(abstract)

Comments Accepted at EMNLP 2025, Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.08034 2025-09-03 cs.OH 50%

Opportunities and Applications of GenAI in Smart Cities: A User-Centric Survey

Ankit Shetgaonkar, Dipen Pradhan, Lakshit Arora, Sanjay Surendranath Girija, Shashank Kapoor, Aman Raj

专题命中 多模态生成 :multimodal(abstract)

Comments Accepted in IEEE COINS 2025

Journal ref 2025 IEEE International Conference on Omni-layer Intelligent Systems (COINS)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12166 2025-08-28 cs.RO cs.LG cs.SY eess.SY 50%

Belief-Conditioned One-Step Diffusion: Real-Time Trajectory Planning with Just-Enough Sensing

Gokul Puthumanaillam, Aditya Penumarti, Manav Vora, Paulo Padrao, Jose Fuentes, Leonardo Bobadilla, Jane Shin, Melkior Ornik

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) University of Florida(佛罗里达大学) Providence College(普罗维登斯学院) Florida International University(佛罗里达国际大学)

专题命中 多模态生成 :multi-modal(abstract)

Comments Accepted to CoRL 2025 (Conference on Robot Learning)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.14207 2025-08-28 cs.RO 50%

A Comprehensive Review on Traffic Datasets and Simulators for Autonomous Vehicles

Supriya Sarker, Brent Maples, Iftekharul Islam, Muyang Fan, Christos Papadopoulos, Weizi Li

专题命中 多模态生成 :multimodal(abstract)

Comments This manuscript has been withdrawn due to the need for substantial updates and revisions

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11639 2025-08-27 cs.LG 50%

Deep Generative Methods and Tire Architecture Design

Fouad Oubari, Raphael Meunier, Rodrigue Décatoire, Mathilde Mougeot

机构 * ENS Paris-Saclay, Centre Borelli(巴黎-萨克雷大学ENS分校,Borelli中心) Université Paris-Saclay, CNRS, ENS Paris-Saclay, Centre Borelli(巴黎-萨克雷大学,法国国家科学研究中心,巴黎-萨克雷大学ENS分校,Borelli中心) ENSIIE(Évry法国的ENSIEE)

专题命中 多模态生成 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.00210 2025-08-26 cs.LG cs.CE cs.SY eess.SY 50%

Generative Machine Learning in Adaptive Control of Dynamic Manufacturing Processes: A Review

Suk Ki Lee, Hyunwoong Ko

机构 * School of Manufacturing Systems and Networks, Arizona State University(制造系统与网络学院,亚利桑那州立大学)

专题命中 多模态生成 :multimodal(abstract)

Comments 12 pages, 1 figure, 1 table. This paper has been accepted for publication in the proceedings of ASME IDETC-CIE 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.13799 2025-08-21 math.ST math.PR stat.ML stat.TH 50%

Non-asymptotic bounds for forward processes in denoising diffusions: Ornstein-Uhlenbeck is hard to beat

Miha Brešar, Aleksandar Mijatović

专题命中 多模态生成 :multi-modal(abstract)

Comments new Subsection 4.1 on Kinetic Langevin diffusion as forward process included; to appear in Annals of Applied Probability; 25 pages, 4 figures; see short YouTube videos https://youtu.be/hQvfpwI0UPk?si=tfL-DrH2EzqCuGSN and https://youtu.be/xjzVPOEkl44?si=fq9l3kZFg8eELYG3 explaining the main results and ideas of proofs

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12701 2025-08-19 eess.SY cs.SY 50%

Deadline-Aware Bandwidth Allocation for Semantic Generative Communication with Diffusion Models

Jinhyuk Choi, Jihong Park, Seungeun Oh, Seong-Lyun Kim

专题命中 多模态生成 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12629 2025-08-19 cs.LG q-bio.BM 50%

FlowMol3: Flow Matching for 3D De Novo Small-Molecule Generation

Ian Dunn, David R. Koes

机构 * Department of Computational and Systems Biology, University of Pittsburgh(计算与系统生物学系,匹兹堡大学)

专题命中 多模态生成 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.19948 2025-08-15 cs.RO 50%

Motion Planning Diffusion: Learning and Adapting Robot Motion Planning with Diffusion Models

J. Carvalho, A. Le, P. Kicki, D. Koert, J. Peters

机构 * Intelligent Autonomous Systems Lab, Computer Science Department, Technical University of Darmstadt, Germany(德意志图林根技术大学计算机科学系智能自主系统实验室) Poznan University of Technology, Poland(波兰波兹南技术大学) IDEAS, Warsaw, Poland(波兰华沙IDEAS) Centre for Cognitive Science, Technical University of Darmstadt, Germany(德意志图林根技术大学认知科学中心) German Research Center for AI (DFKI), Research Department: SAIROL, Darmstadt, Germany(德国人工智能研究中心(DFKI)研究部:SAIROL,德意志图林根,德国) Hessian.AI, Darmstadt, Germany(黑森AI,德意志图林根,德国)

专题命中 多模态生成 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.12724 2025-08-14 cs.RO 50%

Responsive Noise-Relaying Diffusion Policy: Responsive and Efficient Visuomotor Control

Zhuoqun Chen, Xiu Yuan, Tongzhou Mu, Hao Su

机构 * UC San Diego(UC圣地亚哥大学)

专题命中 多模态生成 :multi-modal(abstract)

Comments Project website: https://rnr-dp.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09028 2025-08-13 cs.HC 50%

Envisioning Generative Artificial Intelligence in Cartography and Mapmaking

Yuhao Kang, Chenglong Wang

专题命中 多模态生成 :multimodal(abstract)

Comments 9 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.10397 2025-08-12 cs.CY 50%

Large Model Empowered Metaverse: State-of-the-Art, Challenges and Opportunities

Yuntao Wang, Qinnan Hu, Zhou Su, Linkang Du, Qichao Xu, Weiwei Li

专题命中 多模态生成 :multimodal(abstract)

Comments 9 pages,5 figures, 1 table, accepted by IEEE Network in Aug. 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.08501 2025-08-11 cs.LG 50%

Learning to Match Unpaired Data with Minimum Entropy Coupling

Mustapha Bounoua, Giulio Franzese, Pietro Michiardi

机构 * Department of Data Science, EURECOM, France(数据科学系,EURECOM,法国)

专题命中 多模态生成 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04732 2025-08-08 cs.LG cs.GR 50%

LumiGen: An LVLM-Enhanced Iterative Framework for Fine-Grained Text-to-Image Generation

Xiaoqi Dong, Xiangyu Zhou, Nicholas Evans, Yujia Lin

机构 * Dali University(大理大学) Bandırma Onyedi Eylül University(班迪尔马第十七个九月大学)

专题命中 多模态生成 :cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10576 2025-08-07 cs.GR 50%

Robust Photo-Realistic Hand Gesture Generation: from Single View to Multiple View

Qifan Fu, Xu Chen, Muhammad Asad, Shanxin Yuan, Changjae Oh, Gregory Slabaugh

专题命中 多模态生成 :multi-modal(abstract)

Comments This nine pages paper has been accepted for publication in Proceedings of the 33rd ACM International Conference on Multimedia (ACM MM 2025). This is the author's version which has not been fully edited and content may change prior to final publication. Citation information: DOI https://doi.org/10.1145/3746027.3755828

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.09488 2025-08-07 quant-ph cond-mat.dis-nn cond-mat.str-el 50%

Foundation Neural-Networks Quantum States as a Unified Ansatz for Multiple Hamiltonians

Riccardo Rende, Luciano Loris Viteritti, Federico Becca, Antonello Scardicchio, Alessandro Laio, Giuseppe Carleo

专题命中 多模态生成 :multimodal(abstract)

Comments 10 pages, 6 figures, 1 table

Journal ref Nature Communications 16, 7213 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02117 2025-08-05 eess.SP 50%

Scoring ISAC: Benchmarking Integrated Sensing and Communications via Score-Based Generative Modeling

Lin Chen, Chang Cai, Huiyuan Yang, Xiaojun Yuan, Ying-Jun Angela Zhang

专题命中 多模态生成 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01975 2025-08-05 cs.LG stat.ML 50%

Diffusion models for inverse problems

Hyungjin Chung, Jeongsol Kim, Jong Chul Ye

机构 * EverEx KAIST(韩国科学技术院)

专题命中 多模态生成 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏