arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-11-05 至 2025-11-05 共收录 8 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 8 篇

2510.16888 2025-11-05 cs.CV 83%

Uniworld-V2: Reinforce Image Editing with Diffusion Negative-aware Finetuning and MLLM Implicit Feedback

Zongjian Li, Zheyuan Liu, Qihui Zhang, Bin Lin, Feize Wu, Shenghai Yuan, Zhiyuan Yan, Yang Ye, Wangbo Yu, Yuwei Niu, Shaodong Wang, Xinhua Cheng, Li Yuan

专题命中 多模态生成 :MLLM(title,abstract);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.12973 2025-11-05 eess.IV cs.CV cs.LG q-bio.QM 83%

Cross-modal Diffusion Modelling for Super-resolved Spatial Transcriptomics

Xiaofei Wang, Xingxu Huang, Stephen J. Price, Chao Li

机构 * Department of Clinical Neurosciences, University of Cambridge, UK(剑桥大学临床神经科学系) Department of Applied Mathematics and Theoretical Physics, University of Cambridge, UK(剑桥大学应用数学与理论物理系) School of Science and Engineering, University of Dundee, UK(邓迪大学科学与工程学院) School of Medicine, University of Dundee, UK(邓迪大学医学院) Zhejiang Lab, China(浙江实验室)

专题命中 多模态生成 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02046 2025-11-05 cs.CV cs.AI 76%

Text-VQA Aug: Pipelined Harnessing of Large Multimodal Models for Automated Synthesis

Soham Joshi, Shwet Kamal Mishra, Viswanath Gopalakrishnan

机构 * International Institute of Information Technology Bangalore(国际信息科技学院班加罗尔)

专题命中 多模态生成 :multimodal(title);分类 cs.CV、cs.AI

Comments First two authors contributed equally

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02468 2025-11-05 cs.HC cs.CV 70%

HAGI++: Head-Assisted Gaze Imputation and Generation

Chuhan Jiao, Zhiming Hu, Andreas Bulling

机构 * University of Stuttgart(斯图加特大学) The Hong Kong University of Science(香港科学大学)

专题命中 多模态生成 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

Comments Extended version of our UIST'25 paper "HAGI: Head-Assisted Gaze Imputation for Mobile Eye Trackers"

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02423 2025-11-05 eess.SP 67%

LLM4PG: Adapting Large Language Model for Pathloss Map Generation via Synesthesia of Machines

Mingran Sun, Lu Bai, Xiang Cheng, Jianjun Wu

专题命中 多模态生成 :multi-modal(abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02091 2025-11-05 cs.LG cs.AI 57%

Natural Building Blocks for Structured World Models: Theory, Evidence, and Scaling

Lancelot Da Costa, Sanjeev Namjoshi, Mohammed Abbas Ansari, Bernhard Schölkopf

机构 * VERSES AI Research Lab(VERSES AI研究实验室) University of Tübingen(图宾根大学) ELLIS Institute, Tübingen(图宾根ELLIS研究所) MPI for Intelligent Systems, Tübingen(图宾根智能系统研究所)

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

Comments 13 pages, 3 figures, under review for World Modeling Workshop 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01894 2025-11-05 cs.GR cs.AI cs.LG 57%

LGCC: Enhancing Flow Matching Based Text-Guided Image Editing with Local Gaussian Coupling and Context Consistency

Fangbing Liu, Pengfei Duan, Wen Li, Yi He

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02530 2025-11-05 cs.AR 50%

Implementation and Evaluation of Stable Diffusion on a General-Purpose CGLA Accelerator

Takuto Ando, Yu Eto, Yasuhiko Nakashima

专题命中 多模态生成 :multi-modal(abstract)

Comments This paper is accepted at 2025 IEEE 18th International Symposium on Embedded Multicore/Many-core Systems-on-Chip (MCSoC)

详情

展开后加载摘要…

URL PDF HTML 收藏