arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-08-21 至 2025-08-21 共收录 7 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 7 篇

2504.02906 2025-08-21 cs.CL cs.AI 84%

Boosting Chart-to-Code Generation in MLLM via Dual Preference-Guided Refinement

Zhihan Zhang, Yixin Cao, Lizi Liao

机构 * Singapore Management University(新加坡国立管理学院) Fudan University(复旦大学)

专题命中 多模态生成 :MLLM(title);multimodal(abstract);cross-modal(abstract);分类 cs.CL、cs.AI

Comments Accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.13034 2025-08-21 cs.CL cs.AI cs.CV cs.LG 67%

Natural Language Generation from Visual Events: State-of-the-Art and Key Open Questions

Aditya K Surikuchi, Raquel Fernández, Sandro Pezzelle

机构 * Institute for Logic, Language and Computation (ILLC), University of Amsterdam(逻辑、语言与计算研究所(ILLC),阿姆斯特丹大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14718 2025-08-21 cs.CL 57%

The Digital Sous Chef -- A Comparative Study on Fine-Tuning Language Models for Recipe Generation

Shubham Pundhir, Ganesh Bagler

机构 * Indraprastha Institute of Information Technology(印度理工学院信息技术研究所)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CL

Comments 8 pages, 4 figures. Code is available at: https://github.com/shubh-iiit/RecipeGPT2-Your-Own-AI-Chef

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14405 2025-08-21 cs.CV 57%

CTA-Flux: Integrating Chinese Cultural Semantics into High-Quality English Text-to-Image Communities

Yue Gong, Shanyuan Liu, Liuzhuozheng Li, Jian Zhu, Bo Cheng, Liebucha Wu, Xiaoyu Wu, Yuhang Ma, Dawei Leng, Yuhui Yin

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14393 2025-08-21 cs.CV 57%

Img2ST-Net: Efficient High-Resolution Spatial Omics Prediction from Whole Slide Histology Images via Fully Convolutional Image-to-Image Learning

Junchao Zhu, Ruining Deng, Junlin Guo, Tianyuan Yao, Juming Xiong, Chongyu Qu, Mengmeng Yin, Yu Wang, Shilin Zhao, Haichun Yang, Daguang Xu, Yucheng Tang, Yuankai Huo

机构 * Department of Computer Science, Vanderbilt University(计算机科学系,范德比尔特大学) Weill Cornell Medicine(韦尔·科恩医学中心) Department of Electrical and Computer Engineering, Vanderbilt University(电气与计算机工程系,范德比尔特大学) Department of Biostatistics, Vanderbilt University Medical Center(生物统计学系,范德比尔特大学医学中心) NVIDIA

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14359 2025-08-21 cs.CV 57%

Taming Transformer for Emotion-Controllable Talking Face Generation

Ziqi Zhang, Cheng Deng

机构 * Xidian University(西安电子科技大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.13799 2025-08-21 math.ST math.PR stat.ML stat.TH 50%

Non-asymptotic bounds for forward processes in denoising diffusions: Ornstein-Uhlenbeck is hard to beat

Miha Brešar, Aleksandar Mijatović

专题命中 多模态生成 :multi-modal(abstract)

Comments new Subsection 4.1 on Kinetic Langevin diffusion as forward process included; to appear in Annals of Applied Probability; 25 pages, 4 figures; see short YouTube videos https://youtu.be/hQvfpwI0UPk?si=tfL-DrH2EzqCuGSN and https://youtu.be/xjzVPOEkl44?si=fq9l3kZFg8eELYG3 explaining the main results and ideas of proofs

详情

展开后加载摘要…

URL PDF HTML 收藏