arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-08-13 至 2025-08-13 共收录 8 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 8 篇

2508.08821 2025-08-13 cs.CV 79%

3DFroMLLM: 3D Prototype Generation only from Pretrained Multimodal LLMs

Noor Ahmed, Cameron Braunstein, Steffen Eger, Eddy Ilg

专题命中 多模态生成 :multimodal(title);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.11676 2025-08-13 cs.RO cs.AI cs.LG cs.MA 79%

Hypergraph-based Motion Generation with Multi-modal Interaction Relational Reasoning

Keshu Wu, Yang Zhou, Haotian Shi, Dominique Lord, Bin Ran, Xinyue Ye

机构 * organization= Center for Geospatial Sciences, Applications Department of Landscape of Architecture Urban Planning, Texas A\&M University , addressline= 788 Ross St , city= College Station , postcode= 77840 , state= TX , country= United States organization= Zachry Department of Civil Environmental Engineering, Texas A\&M University , addressline= 201 Dwight Look Engineering Building , city= College Station , postcode= 77843 , state= TX , country= United States organization= Department of Civil Environmental Engineering, University of Wisconsin-Madison , addressline= 1415 Engineering Dr , city= Madison , postcode= 53706 , state= WI , country= United States

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21349 2025-08-13 cs.CE 78%

Out of the Past: An AI-Enabled Pipeline for Traffic Simulation from Noisy, Multimodal Detector Data and Stakeholder Feedback

Rex Chen, Karen Wu, John McCartney, Norman Sadeh, Fei Fang

专题命中 多模态生成 :multimodal(title,abstract)

Comments 17 pages; 1 table; 6 figures; extended version of accepted version, published at the 2025 Winter Simulation Conference (WSC '25)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08987 2025-08-13 cs.CV cs.HC 74%

ColorGPT: Leveraging Large Language Models for Multimodal Color Recommendation

Ding Xia, Naoto Inoue, Qianru Qiu, Kotaro Kikuchi

机构 * The University of Tokyo(东京大学) CyberAgent AI Lab(CyberAgent AI实验室)

专题命中 多模态生成 :multimodal(title);分类 cs.CV

Comments Accepted to ICDAR2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.00043 2025-08-13 cs.CL cs.AI cs.CV 67%

CrossWordBench: Evaluating the Reasoning Capabilities of LLMs and LVLMs with Controllable Puzzle Generation

Jixuan Leng, Chengsong Huang, Langlin Huang, Bill Yuchen Lin, William W. Cohen, Haohan Wang, Jiaxin Huang

机构 * CMU(卡内基梅隆大学) WUSTL(华盛顿大学) UIUC(伊利诺伊大学香槟分校)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.12899 2025-08-13 cs.CV cs.AI cs.MM 67%

DreamStory: Open-Domain Story Visualization by LLM-Guided Multi-Subject Consistent Diffusion

Huiguo He, Huan Yang, Zixi Tuo, Yuan Zhou, Qiuyue Wang, Yuhang Zhang, Zeyu Liu, Wenhao Huang, Hongyang Chao, Jian Yin

机构 * School of Computer Science and Engineering, Sun Yat-sen University(中山大学计算机科学与工程学院) School of Artificial Intelligence, Sun Yat-sen University(中山大学人工智能学院)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI、cs.MM

Comments Accepted by TPAMI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08891 2025-08-13 cs.CV 57%

Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos

Chaoyi Wang, Yifan Yang, Jun Pei, Lijie Xia, Jianpo Liu, Xiaobing Yuan, Xinhan Di

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

Comments This paper has been accepted by ICCV 2025 Workshop MMFM4

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09028 2025-08-13 cs.HC 50%

Envisioning Generative Artificial Intelligence in Cartography and Mapmaking

Yuhao Kang, Chenglong Wang

专题命中 多模态生成 :multimodal(abstract)

Comments 9 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏