arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-10-14 至 2025-10-14 共收录 16 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 16 篇

2510.11096 2025-10-14 cs.CV 83%

CoDefend: Cross-Modal Collaborative Defense via Diffusion Purification and Prompt Optimization

Fengling Zhu, Boshi Liu, Jingyu Hua, Sheng Zhong

专题命中 多模态生成 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10037 2025-10-14 cs.CE 82%

Automated Glaucoma Report Generation via Dual-Attention Semantic Parallel-LSTM and Multimodal Clinical Data Integration

Cheng Huang, Weizheng Xie, Zeyu Han, Tsengdar Lee, Karanjit Kooner, Jui-Ka Wang, Ning Zhang, Jia Zhang

专题命中 多模态生成 :multimodal(title,abstract);cross-modal(abstract)

Comments Accepted by IEEE 25th BIBE

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09479 2025-10-14 cs.AI cs.CL 81%

Draw with Thought: Unleashing Multimodal Reasoning for Scientific Diagram Generation

Zhiqing Cui, Jiahao Yuan, Hanqing Wang, Yanshu Li, Chenxu Du, Zhenglong Ding

机构 * Nanjing University of Information Science \& Technology Nanjing China East China Normal University Shanghai China The Hong Kong University of Science Brown University Providence America Southwest Jiaotong University Chengdu China Nanjing University of Information Science \& Technology East China Normal University Brown University Southwest Jiaotong University

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments 10 pages, 5 figures, accepted to appear in the Proceedings of the 33rd ACM International Conference on Multimedia (MM '25)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10914 2025-10-14 eess.SY cs.SY 78%

Optimal Multi-Modal Transportation and Electric Power Flow: The Value of Coordinated Dynamic Operation

Jiajie Qiu, Dakota Thompson, Kamal Youcef-Toumi, Amro M. Farid

专题命中 多模态生成 :multi-modal(title,abstract)

Comments 31 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10633 2025-10-14 cs.AI 70%

Collaborative Text-to-Image Generation via Multi-Agent Reinforcement Learning and Semantic Fusion

Jiabao Shi, Minfeng Qi, Lefeng Zhang, Di Wang, Yingjie Zhao, Ziying Li, Yalong Xing, Ningran Li

机构 * Minzu University of China(民族大学) City University of Macau(澳门城市大学) Key Laboratory of Computing Power Network and Information Security, Ministry of Education, Shandong Computer Science Center (National Supercomputer Center in Jinan), Qilu University of Technology (Shandong Academy of Sciences)(计算能力网络与信息安全重点实验室,教育部,山东计算机科学中心(济南国家超级计算机中心),齐鲁工业大学(山东省科学院)) The University of Adelaide(阿德莱德大学)

专题命中 多模态生成 :multimodal(abstract);cross-modal(abstract);分类 cs.AI

Comments 16 pages, 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07249 2025-10-14 cs.CV 70%

TalkCuts: A Large-Scale Dataset for Multi-Shot Human Speech Video Generation

Jiaben Chen, Zixin Wang, Ailing Zeng, Yang Fu, Xueyang Yu, Siyuan Cen, Julian Tanke, Yihang Chen, Koichi Saito, Yuki Mitsufuji, Chuang Gan

机构 * UMass Amherst(马萨诸塞大学阿姆赫斯特分校) Sony AI(索尼人工智能) UC San Diego(加州大学圣地亚哥分校)

专题命中 多模态生成 :multimodal(abstract);multi-modal(abstract);分类 cs.CV

Comments Project page: https://talkcuts.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10637 2025-10-14 cs.RO 67%

High-Fidelity Simulated Data Generation for Real-World Zero-Shot Robotic Manipulation Learning with Gaussian Splatting

Haoyu Zhao, Cheng Zeng, Linghao Zhuang, Yaxi Zhao, Shengke Xue, Hao Wang, Xingyue Zhao, Zhongyu Li, Kehan Li, Siteng Huang, Mingxiu Chen, Xin Li, Deli Zhao, Hua Zou

机构 * Wuhan University(武汉大学) DAMO Academy, Alibaba Group(达摩院,阿里巴巴集团) Hupan Lab(虎扑实验室) The Chinese University of Hong Kong(香港中文大学) Tsinghua University(清华大学) Huazhong University of Science and Technology(华中科技大学) Zhejiang University(浙江大学)

专题命中 多模态生成 :multi-modal(abstract);MLLM(abstract)

Comments 13 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09757 2025-10-14 q-bio.OT 67%

A path towards AI-scale, interoperable biological data

Brian Aevermann, Andrea Califano, Chi-Li Chiu, Nathan Clack, William M. Clemons, Jonah Cool Florence D. D'Orazi, Joseph L. DeRisi, Joshua E. Elias, Elizabeth Fahsbender, Scott E. Fraser, Carlos G. Gonzalez, Matthias Haury, Theofanis Karaletsos, Shana O. Kelley, Aly A. Khan, Alan R. Lowe, Emma Lundberg, Ryan A. McClure, Stephani Otte, Evan O. Paull, Loïc A. Royer, Dana Sadgat, Sandra L. Schmid, Samantha Scovanner, Cathy Stolitzka, Jason R. Swedlow, Joan Wong, Garabet Yeretssian, Patricia Brennan, Ambrose J. Carr

专题命中 多模态生成 :multimodal(abstract);multi-modal(abstract)

Comments 8 pages, 2 images

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.08980 2025-10-14 cs.LG cs.AI cs.CV 62%

Learning Diffusion Models with Flexible Representation Guidance

Chenyu Wang, Cai Zhou, Sharut Gupta, Zongyu Lin, Stefanie Jegelka, Stephen Bates, Tommi Jaakkola

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

Comments NeurIPS 2025; Also Oral at ICML 2025 FM4LS workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11650 2025-10-14 cs.CV 57%

InfiniHuman: Infinite 3D Human Creation with Precise Control

Yuxuan Xue, Xianghui Xie, Margaret Kostyrko, Gerard Pons-Moll

机构 * University of Tübingen, Tübingen AI Center(图宾根大学,图宾根人工智能中心) University of Tübingen, Tübingen AI Center, MPI for Informatics, SIC(图宾根大学,图宾根人工智能中心,马克斯·普朗克信息研究所,SIC) University of Tübingen(图宾根大学)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

Comments Accepted to ACM SIGGRAPH Asia 2025. Project website: https://yuxuan-xue.com/infini-human

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10158 2025-10-14 cs.NI cs.AI 57%

Multi-Scale Diffusion Transformer for Jointly Simulating User Mobility and Mobile Traffic Pattern

Ziyi Liu, Qingyue Long, Zhiwen Xue, Huandong Wang, Yong Li

机构 * Department of Electronic Engineering, Beijing National Research Center for Information Science and Technology (BNRist), Tsinghua University(电子工程系,信息科学与技术国家研究中心(BNRist),清华大学) International School, Beijing University of Posts and Telecommunications(国际学院,北京邮电大学)

专题命中 多模态生成 :cross-modal(abstract);分类 cs.AI

Comments 9 pages, 4 figures. Code: https://github.com/tsinghua-fib-lab/MSTDiff

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10156 2025-10-14 cs.CV 57%

ReMix: Towards a Unified View of Consistent Character Generation and Editing

Benjia Zhou, Bin Fu, Pei Cheng, Yanru Wang, Jiayuan Fan, Tao Chen

机构 * Tencent GYLab(腾讯GY实验室) Fudan University(复旦大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11346 2025-10-14 cs.CV cs.AI 54%

Uncertainty-Aware ControlNet: Bridging Domain Gaps with Synthetic Image Generation

Joshua Niemeijer, Jan Ehrhardt, Heinz Handels, Hristina Uzunova

机构 * German Aerospace Center (DLR)(德国航空航天中心) University of Lübeck(吕贝克大学) German Research Center for Artificial Intelligence (DFKI)(德国人工智能研究中心)

专题命中 多模态生成 :分类 cs.CV、cs.AI;multimodal(comments);multimodal foundation model(comments)

Comments Accepted for presentation at ICCV Workshops 2025, "The 4th Workshop on What is Next in Multimodal Foundation Models?" (MMFM)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11138 2025-10-14 cs.SE 50%

What Slows Down FMware Development? An Empirical Study of Developer Challenges and Resolution Times

Zitao Wang, Zhimin Zhao, Michael W. Godfrey

专题命中 多模态生成 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.15567 2025-10-14 cs.LG 50%

Towards Unified and Lossless Latent Space for 3D Molecular Latent Diffusion Modeling

Yanchen Luo, Zhiyuan Liu, Yi Zhao, Sihang Li, Hengxing Cai, Kenji Kawaguchi, Tat-Seng Chua, Yang Zhang, Xiang Wang

机构 * University of Science and Technology of China(中国科学技术大学) National University of Singapore(新加坡国立大学) DP Technology(DP技术)

专题命中 多模态生成 :multi-modal(abstract)

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.15775 2025-10-14 cs.RO 50%

Humanoid Robots and Humanoid AI: Review, Perspectives and Directions

Longbing Cao

机构 * Frontier AI Research Centre, Macquarie University(前沿人工智能研究中心,麦考瑞大学)

专题命中 多模态生成 :multimodal(abstract)

Comments 35 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏