arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-08-12 至 2025-08-12 共收录 16 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 16 篇

2503.06134 2025-08-12 cs.CV 83%

X2I: Seamless Integration of Multimodal Understanding into Diffusion Transformer via Attention Distillation

Jian Ma, Qirong Peng, Xu Guo, Chen Chen, Haonan Lu, Zhenyu Yang

机构 * OPPO AI Center(OPPO人工智能中心) Tsinghua University(清华大学)

专题命中 多模态生成 :multimodal(title,abstract);image-text(abstract);分类 cs.CV

Comments Accepted to ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05658 2025-08-12 cs.CR cs.CV cs.MM 81%

Universally Unfiltered and Unseen:Input-Agnostic Multimodal Jailbreaks against Text-to-Image Model Safeguards

Song Yan, Hui Wei, Jinlong Fei, Guoliang Yang, Zhengyu Zhao, Zheng Wang

机构 * Information Engineering University Zhengzhou China School of Computer Science, \ University Wuhan China Xi’an Jiaotong University Xi’an China Wuhan University Wuhan China Information Engineering University School of Computer Science, \ University Xi’an Jiaotong University Wuhan University

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.MM

Comments This paper has been accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07519 2025-08-12 cs.CV 79%

Exploring Multimodal Diffusion Transformers for Enhanced Prompt-based Image Editing

Joonghyuk Shin, Alchan Hwang, Yujin Kim, Daneul Kim, Jaesik Park

机构 * Seoul National University(首尔国立大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments ICCV 2025. Project webpage: https://joonghyuk.com/exploring-mmdit-web/

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23186 2025-08-12 cs.CV 79%

HiGarment: Cross-modal Harmony Based Diffusion Model for Flat Sketch to Realistic Garment Image

Junyi Guo, Jingxuan Zhang, Fangyu Wu, Huanda Lu, Qiufeng Wang, Wenmian Yang, Eng Gee Lim, Dongming Lu

机构 * Xi’an Jiaotong Liverpool University(西安交通大学利物浦大学) NingboTech University(宁波科技学院) Beijing Normal University(北京师范大学) Zhejiang University(浙江大学)

专题命中 多模态生成 :cross-modal(title);multi-modal(abstract);分类 cs.CV

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07021 2025-08-12 cs.CV 79%

DocRefine: An Intelligent Framework for Scientific Document Understanding and Content Optimization based on Multimodal Large Model Agents

Kun Qian, Wenjie Li, Tianyu Sun, Wenhong Wang, Wenhan Luo

机构 * Shangqiu University(商丘大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06888 2025-08-12 cs.SE 78%

Multi-Modal Requirements Data-based Acceptance Criteria Generation using LLMs

Fanyu Wang, Chetan Arora, Yonghui Liu, Kaicheng Huang, Chakkrit Tantithamthavorn, Aldeida Aleti, Dishan Sambathkumar, David Lo

专题命中 多模态生成 :multi-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.05148 2025-08-12 cs.CV cs.AI cs.LG 73%

LoRA.rar: Learning to Merge LoRAs via Hypernetworks for Subject-Style Conditioned Image Generation

Donald Shenaj, Ondrej Bohdal, Mete Ozay, Pietro Zanuttigh, Umberto Michieli

专题命中 多模态生成 :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments ICCV 2025. Project page: https://donaldssh.github.io/LoRA.rar

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22792 2025-08-12 cs.CV 70%

Rhetorical Text-to-Image Generation via Two-layer Diffusion Policy Optimization

Yuxi Zhang, Yueting Li, Xinyu Du, Sibo Wang

机构 * The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) University of California, Berkeley(加州大学伯克利分校)

专题命中 多模态生成 :multimodal(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10574 2025-08-12 cs.CV cs.MM cs.SD eess.AS 67%

DanceChat: Large Language Model-Guided Music-to-Dance Generation

Qing Wang, Xiaohang Yang, Yilan Dong, Naveen Raj Govindaraj, Gregory Slabaugh, Shanxin Yuan

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.MM、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07146 2025-08-12 cs.CV cs.AI 62%

Intention-Aware Diffusion Model for Pedestrian Trajectory Prediction

Yu Liu, Zhijie Liu, Xiao Ren, You-Fu Li, He Kong

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08220 2025-08-12 cs.CV 57%

Learning User Preferences for Image Generation Model

Wenyi Mo, Ying Ba, Tianyu Zhang, Yalong Bai, Biye Li

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07540 2025-08-12 cs.CV 57%

CoT-Pose: Chain-of-Thought Reasoning for 3D Pose Generation from Abstract Prompts

Junuk Cha, Jihyeon Kim

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

Comments ICCVW'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07225 2025-08-12 eess.IV cs.CV q-bio.QM 57%

HaDM-ST: Histology-Assisted Differential Modeling for Spatial Transcriptomics Generation

Xuepeng Liu, Zheng Jiang, Pinan Zhu, Hanyu Liu, Chao Li

机构 * University of Cambridge, UK(剑桥大学,英国) Northeastern University, Shenyang, China(东北大学,中国)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments 10 pages, 5 figures, includes comparisons with TESLA, HiStoGene, and iStar; submitted to arXiv 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.16726 2025-08-12 cs.CV cs.LG 57%

EDiT: Efficient Diffusion Transformers with Linear Compressed Attention

Philipp Becker, Abhinav Mehrotra, Ruchika Chavhan, Malcolm Chadwick, Luca Morreale, Mehdi Noroozi, Alberto Gil Ramos, Sourav Bhattacharya

机构 * Samsung, AI Center Cambridge(三星人工智能中心剑桥)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.15173 2025-08-12 cs.IR cs.AI cs.ET 57%

Recommendation with Generative Models

Yashar Deldjoo, Zhankui He, Julian McAuley, Anton Korikov, Scott Sanner, Arnau Ramisa, Rene Vidal, Maheswaran Sathiamoorthy, Atoosa Kasrizadeh, Silvia Milano, Francesco Ricci

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

Comments This submission is a full-length book, expanding significantly on two chapters previously submitted (arXiv:2409.10993v1, arXiv:2408.10946v1). It includes additional chapters, context, analysis, and content, providing a comprehensive presentation of the subject. We have ensured it is appropriately presented as a new, distinct work. arXiv admin note: substantial text overlap with arXiv:2409.10993

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.10397 2025-08-12 cs.CY 50%

Large Model Empowered Metaverse: State-of-the-Art, Challenges and Opportunities

Yuntao Wang, Qinnan Hu, Zhou Su, Linkang Du, Qichao Xu, Weiwei Li

专题命中 多模态生成 :multimodal(abstract)

Comments 9 pages,5 figures, 1 table, accepted by IEEE Network in Aug. 2025

详情

展开后加载摘要…

URL PDF HTML 收藏