arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-08-07 至 2025-08-07 共收录 11 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 11 篇

2412.03859 2025-08-07 cs.CV 83%

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation

Hui Zhang, Dexiang Hong, Yitong Wang, Jie Shao, Xinglong Wu, Zuxuan Wu, Yu-Gang Jiang

机构 * Institute of Trustworthy Embodied AI(可信具身人工智能研究院) Fudan University(复旦大学) Shanghai Collaborative Innovation Center of Intelligent Visual Computing(上海智能视觉计算协同创新中心) Bytedance Intelligent Creation(字节跳动智能创作)

专题命中 多模态生成 :multimodal(title,abstract);image-text(abstract);分类 cs.CV

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21167 2025-08-07 cs.CV cs.AI 81%

ChartM$^3$: Benchmarking Chart Editing with Multimodal Instructions

Donglu Yang, Liang Zhang, Zihao Yue, Liangyu Chen, Yichen Xu, Wenxuan Wang, Qin Jin

机构 * independent researcher(独立研究者)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05683 2025-08-07 cs.LG cs.AI cs.CR cs.MM 81%

Multi-Modal Multi-Task Federated Foundation Models for Next-Generation Extended Reality Systems: Towards Privacy-Preserving Distributed Intelligence in AR/VR/MR

Fardis Nadimi, Payam Abdisarabshali, Kasra Borazjani, Jacob Chakareski, Seyyedali Hosseinalipour

机构 * University at Buffalo–SUNY(布法罗大学-纽约州立大学) Department of Electrical Engineering(电气工程系) New Jersey Institute of Technology (NJIT)(新泽西理工学院)

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.AI、cs.MM

Comments 16 pages, 4 Figures, 8 Tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04229 2025-08-07 cs.CV 79%

Intention Enhanced Diffusion Model for Multimodal Pedestrian Trajectory Prediction

Yu Liu, Zhijie Liu, Xiao Ren, You-Fu Li, He Kong

机构 * Guangdong Provincial Key Laboratory of Fully Actuated System Control Theory and Technology, the Southern University of Science and Technology, Shenzhen 518055, China(广东省全自动化系统控制理论与技术重点实验室,南方科技大学,深圳518055,中国) Department of Mechanical Engineering, City University of Hong Kong, Hong Kong SAR, China(香港城市大学机械工程系,香港特别行政区,中国)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments To be presented at the 28th IEEE International Conference on Intelligent Transportation Systems (ITSC), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.20830 2025-08-07 cs.CV 79%

CMT: A Cascade MAR with Topology Predictor for Multimodal Conditional CAD Generation

Jianyu Wu, Yizhou Wang, Xiangyu Yue, Xinzhu Ma, Jingyang Guo, Dongzhan Zhou, Wanli Ouyang, Shixiang Tang

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Shanghai Jiao Tong University(上海交通大学) The Chinese University of Hong Kong(香港中文大学) Beihang University(北京航空航天大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04271 2025-08-07 cs.DC 78%

S2M3: Split-and-Share Multi-Modal Models for Distributed Multi-Task Inference on the Edge

JinYi Yoon, JiHo Lee, Ting He, Nakjung Choi, Bo Ji

专题命中 多模态生成 :multi-modal(title,abstract)

Comments Accepted at IEEE International Conference on Distributed Computing Systems (ICDCS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.18087 2025-08-07 cs.CV 70%

Disentangle Identity, Cooperate Emotion: Correlation-Aware Emotional Talking Portrait Generation

Weipeng Tan, Chuming Lin, Chengming Xu, FeiFan Xu, Xiaobin Hu, Xiaozhong Ji, Junwei Zhu, Chengjie Wang, Yanwei Fu

机构 * Fudan University(复旦大学) Tencent, YouTu Lab(腾讯、YouTu实验室)

专题命中 多模态生成 :cross-modal(abstract);audio-visual(abstract);分类 cs.CV

Comments Accepted by ACM MM'25. arXiv admin note: text overlap with arXiv:2409.03270

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03696 2025-08-07 cs.CR cs.AI cs.CV 62%

PLA: Prompt Learning Attack against Text-to-Image Generative Models

Xinqi Lyu, Yihao Liu, Yanjie Li, Bin Xiao

机构 * The Hong Kong Polytechnic University(香港理工大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 10 pages, 3 figures, and published to ICCV2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04049 2025-08-07 cs.CV 57%

Motion is the Choreographer: Learning Latent Pose Dynamics for Seamless Sign Language Generation

Jiayi He, Xu Wang, Shengeng Tang, Yaxiong Wang, Lechao Cheng, Dan Guo

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments 9 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10576 2025-08-07 cs.GR 50%

Robust Photo-Realistic Hand Gesture Generation: from Single View to Multiple View

Qifan Fu, Xu Chen, Muhammad Asad, Shanxin Yuan, Changjae Oh, Gregory Slabaugh

专题命中 多模态生成 :multi-modal(abstract)

Comments This nine pages paper has been accepted for publication in Proceedings of the 33rd ACM International Conference on Multimedia (ACM MM 2025). This is the author's version which has not been fully edited and content may change prior to final publication. Citation information: DOI https://doi.org/10.1145/3746027.3755828

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.09488 2025-08-07 quant-ph cond-mat.dis-nn cond-mat.str-el 50%

Foundation Neural-Networks Quantum States as a Unified Ansatz for Multiple Hamiltonians

Riccardo Rende, Luciano Loris Viteritti, Federico Becca, Antonello Scardicchio, Alessandro Laio, Giuseppe Carleo

专题命中 多模态生成 :multimodal(abstract)

Comments 10 pages, 6 figures, 1 table

Journal ref Nature Communications 16, 7213 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏