arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-09-18 至 2025-09-18 共收录 6 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 6 篇

2509.13642 2025-09-18 cs.LG cs.CV 83%

LLM-I: LLMs are Naturally Interleaved Multimodal Creators

Zirun Guo, Feng Zhang, Kai Jia, Tao Jin

机构 * Zhejiang University(浙江大学)

专题命中 多模态生成 :multimodal(title);MLLM(abstract);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.01086 2025-09-18 cs.CV cs.AI 81%

DPDEdit: Detail-Preserved Diffusion Models for Multimodal Fashion Image Editing

Xiaolong Wang, Zhi-Qi Cheng, Jue Wang, Xiaojiang Peng

机构 * Shenzhen Technology University(深圳科技大学) University of Washington(华盛顿大学) Shenzhen Institute of Advanced Technology, Chinese Academy of Sciences(深圳先进技术研究院,中国科学院)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 13 pages,12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13227 2025-09-18 math.OC cs.AI cs.SY eess.SY 74%

Rich Vehicle Routing Problem in Disaster Management enabling Temporally-causal Transhipments across Multi-Modal Transportation Network

Santanu Banerjee, Goutam Sen, Siddhartha Mukhopadhyay

机构 * Department of Industrial and Systems Engineering (ISE), Indian Institute of Technology (IIT) Kharagpur(工业与系统工程系,印度理工学院Kharagpur分校)

专题命中 多模态生成 :multi-modal(title);分类 cs.AI

Comments Major changes in version II: 1) Supplementary is now a separate document, 2) Algorithm steps have been updated with pseudocode in the Heuristic, 3) Explanation of the MILP formulation construction is further detailed in a supplementary section

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.14837 2025-09-18 cs.RO cs.LG 71%

Learning Multimodal Attention for Manipulating Deformable Objects with Changing States

Namiko Saito, Mayu Tatsumi, Ayuna Kubo, Kanata Suzuki, Hiroshi Ito, Shigeki Sugano, Tetsuya Ogata

机构 * Future Robotics Organization, Waseda University(早稻田大学未来机器人组织) Microsoft Research Asia(微软亚洲研究院) Department of Modern Mechanical Engineering, Waseda University(早稻田大学现代机械工程系) Artificial Intelligence Laboratories, Fujitsu Limited(Fujitsu 人工智能实验室) Center for Technology Innovation - Controls and Robotics, Research & Development Group, Hitachi, Ltd.(富士通技术研发集团技术创新中心 - 控制与机器人) Faculty of Science and Engineering, Waseda University(早稻田大学工学部) National Institute of Advanced Science and Technology(国家先进科学研究院)

专题命中 多模态生成 :multimodal(title)

Comments Humanoids2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13760 2025-09-18 cs.CV 57%

Iterative Prompt Refinement for Safer Text-to-Image Generation

Jinwoo Jeon, JunHyeok Oh, Hayeong Lee, Byung-Jun Lee

机构 * Korea University(韩国大学)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13696 2025-09-18 cs.CL 57%

Integrating Text and Time-Series into (Large) Language Models to Predict Medical Outcomes

Iyadh Ben Cheikh Larbi, Ajay Madhavan Ravichandran, Aljoscha Burchardt, Roland Roller

机构 * German Research Center for Artificial Intelligence (DFKI)(德国人工智能研究中心) Technical University Berlin(柏林技术大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CL

Comments Presented and published at BioCreative IX

详情

展开后加载摘要…

URL PDF HTML 收藏