arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-08-06 至 2025-08-06 共收录 8 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 8 篇

2411.04954 2025-08-06 cs.CV 84%

CAD-MLLM: Unifying Multimodality-Conditioned CAD Generation With MLLM

Jingwei Xu, Chenyu Wang, Zibo Zhao, Wen Liu, Yi Ma, Shenghua Gao

机构 * School of Information Science and Technology, ShanghaiTech University(信息科学与技术学院,上海科技大学) Transcengram DeepSeek AI University of Hong Kong(香港大学)

专题命中 多模态生成 :MLLM(title,abstract);multimodal(abstract);分类 cs.CV

Comments Project page: https://cad-mllm.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03426 2025-08-06 cs.CV cs.AI cs.LG 81%

R2GenKG: Hierarchical Multi-modal Knowledge Graph for LLM-based Radiology Report Generation

Futian Wang, Yuhan Qiao, Xiao Wang, Fuling Wang, Yuxiang Zhang, Dengdi Sun

机构 * School of Computer Science and Technology, Anhui University(安徽大学计算机科学与技术学院)

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21893 2025-08-06 cs.CV 79%

Aether Weaver: Multimodal Affective Narrative Co-Generation with Dynamic Scene Graphs

Saeed Ghorbani

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03690 2025-08-06 cs.CV cs.RO 57%

Veila: Panoramic LiDAR Generation from a Monocular RGB Image

Youquan Liu, Lingdong Kong, Weidong Yang, Ao Liang, Jianxiong Gao, Yang Wu, Xiang Xu, Xin Li, Linfeng Li, Runnan Chen, Ben Fei

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV

Comments Preprint; 10 pages, 6 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03669 2025-08-06 cs.CV cs.RO 57%

OmniShape: Zero-Shot Multi-Hypothesis Shape and Pose Estimation in the Real World

Katherine Liu, Sergey Zakharov, Dian Chen, Takuya Ikeda, Greg Shakhnarovich, Adrien Gaidon, Rares Ambrus

机构 * Toyota Research Institute(丰田研究院) Woven by Toyota(丰田编织) Toyota Technological Institute at Chicago(芝加哥丰田技术研究所)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

Comments 8 pages, 5 figures. This version has typo fixes on top of the version published at ICRA 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03539 2025-08-06 cs.CV 57%

Quality-Aware Language-Conditioned Local Auto-Regressive Anomaly Synthesis and Detection

Long Qian, Bingke Zhu, Yingying Chen, Ming Tang, Jinqiao Wang

机构 * Foundation Model Research Center, Institute of Automation, Chinese Academy of Sciences(自动化研究所基础模型研究中心,中国科学院) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Objecteye Inc.(Objecteye公司)

专题命中 多模态生成 :image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03535 2025-08-06 cs.CV 57%

CoEmoGen: Towards Semantically-Coherent and Scalable Emotional Image Content Generation

Kaishen Yuan, Yuting Zhang, Shang Gao, Yijie Zhu, Wenshuo Chen, Yutao Yue

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments 10 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03320 2025-08-06 cs.CV 57%

Skywork UniPic: Unified Autoregressive Modeling for Visual Understanding and Generation

Peiyu Wang, Yi Peng, Yimeng Gan, Liang Hu, Tianyidan Xie, Xiaokun Wang, Yichen Wei, Chuanxin Tang, Bo Zhu, Changshi Li, Hongyang Wei, Eric Li, Xuchen Song, Yang Liu, Yahui Zhou

机构 * Multimodality Team, Skywork AI(Skywork AI 多模态团队)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏