arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-07-29 至 2025-07-29 共收录 13 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 13 篇

2507.20368 2025-07-29 cs.CV cs.MM 84%

MagicAnime: A Hierarchically Annotated, Multimodal and Multitasking Dataset with Benchmarks for Cartoon Animation Generation

Shuolin Xu, Bingyuan Wang, Zeyu Cai, Fangteng Fu, Yue Ma, Tongyi Lee, Hongchuan Yu, Zeyu Wang

机构 * National Centre for Computer Animation, Bournemouth University(伯恩茅斯大学计算机动画国家中心) Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Hong Kong University of Science and Technology(香港科技大学) Department of Computer Science and Information Engineering, National Cheng Kung University(国立成功大学计算机科学与信息工程系)

专题命中 多模态生成 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV、cs.MM

Comments 8 pages,6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.12789 2025-07-29 cs.CV 83%

Efficient Physics Simulation for 3D Scenes via MLLM-Guided Gaussian Splatting

Haoyu Zhao, Hao Wang, Xingyue Zhao, Hao Fei, Hongqiu Wang, Chengjiang Long, Hua Zou

机构 * School of Computer Science, Wuhan University(武汉大学计算机学院) Wuhan National Laboratory for Optoelectronics, Huazhong University of Science and Technology(华中科技大学光电研究院) Meta Reality Lab(Meta现实实验室) Xi’an Jiao Tong University(西安交通大学) National University of Singapore(新加坡国立大学) The Department of Systems Hub, Hong Kong University of Science and Technology (Guangzhou)(香港科技大学系统枢纽部门(广州))

专题命中 多模态生成 :MLLM(title,abstract);multi-modal(abstract);分类 cs.CV

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19939 2025-07-29 cs.CV 79%

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs

Jiaze Wang, Rui Chen, Haowang Cui

机构 * Tianjin University(天津大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02984 2025-07-29 cs.CL 79%

From Answers to Rationales: Self-Aligning Multimodal Reasoning with Answer-Oriented Chain-of-Thought

Wentao Tan, Qiong Cao, Yibing Zhan, Chao Xue, Changxing Ding

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.21291 2025-07-29 cs.CV 79%

MIGE: Mutually Enhanced Multimodal Instruction-Based Image Generation and Editing

Xueyun Tian, Wei Li, Bingbing Xu, Yige Yuan, Yuanzhuo Wang, Huawei Shen

机构 * State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences(人工智能安全国家重点实验室,计算技术研究所,中国科学院) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments This paper have been accepted by ACM MM25

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.03225 2025-07-29 cs.CV 79%

MaterialPicker: Multi-Modal DiT-Based Material Generation

Xiaohe Ma, Valentin Deschaintre, Miloš Hašan, Fujun Luan, Kun Zhou, Hongzhi Wu, Yiwei Hu

机构 * State Key Lab of CAD\&CG, Zhejiang University(浙江大学CAD与CG国家重点实验室) Adobe Research(Adobe研究) State Key Lab of CAD\&CG, Zhejiang University and ZJU-FaceUnity Joint Lab of Intelligent Graphics(浙江大学CAD与CG国家重点实验室) ZJU-FaceUnity Joint Lab of Intelligent Graphics(浙大-面 unity 智能图形联合实验室)

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.17046 2025-07-29 cs.CV 70%

Text-to-Image Generation Via Energy-Based CLIP

Roy Ganz, Michael Elad

机构 * Electrical Engineering Department Technion(技术学院电子工程系) Computer Science Department Technion(技术学院计算机科学系)

专题命中 多模态生成 :multimodal(abstract);image-text(abstract);分类 cs.CV

Comments Accepted to TMLR

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19492 2025-07-29 cs.HC cs.AI cs.CV 62%

ChartGen: Scaling Chart Understanding Via Code-Guided Synthetic Chart Generation

Jovana Kondic, Pengyuan Li, Dhiraj Joshi, Zexue He, Shafiq Abedin, Jennifer Sun, Ben Wiesel, Eli Schwartz, Ahmed Nassar, Bo Wu, Assaf Arbelle, Aude Oliva, Dan Gutfreund, Leonid Karlinsky, Rogerio Feris

机构 * MIT(麻省理工学院) MIT-IBM Watson AI Labs(麻省理工-IBM沃森人工智能实验室) IBM Research(IBM研究院)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.00045 2025-07-29 cs.MM cs.AI cs.LG 62%

Detecting Multimedia Generated by Large AI Models: A Survey

Li Lin, Neeraj Gupta, Yue Zhang, Hainan Ren, Chun-Hao Liu, Feng Ding, Xin Wang, Xin Li, Luisa Verdoliva, Shu Hu

机构 * Department of Computer and Information Technology, Purdue University(普渡大学计算机与信息科技系) School of Software, Nanchang University(南昌大学软件学院) Amazon Prime Video(亚马逊Prime视频) Department of Epidemiology and Biostatistics, School of Public Health(公共卫生学院流行病学与生物统计学系) Department of Computer Science, College of Nanotechnology, Science, and Engineering(纳米技术、科学与工程学院计算机科学系) University at Albany, SUNY(阿尔巴尼大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20976 2025-07-29 cs.CV 57%

Adapting Vehicle Detectors for Aerial Imagery to Unseen Domains with Weak Supervision

Xiao Fang, Minhyek Jeon, Zheyang Qin, Stanislav Panev, Celso de Melo, Shuowen Hu, Shayok Chakraborty, Fernando De la Torre

机构 * Carnegie Mellon University(卡内基梅隆大学) DEVCOM Army Research Laboratory(陆军研究实验室) Florida State University(佛罗里达州立大学)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19882 2025-07-29 cs.AI 57%

Causality-aligned Prompt Learning via Diffusion-based Counterfactual Generation

Xinshu Li, Ruoyu Wang, Erdun Gao, Mingming Gong, Lina Yao

机构 * The University of New South Wales(新南威尔士大学) The University of Adelaide(阿德莱德大学) The University of Melbourne(墨尔本大学) CSIRO’s Data 61(CSIRO数据61)

专题命中 多模态生成 :image-text(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.21771 2025-07-29 cs.CV 57%

A Unified Image-Dense Annotation Generation Model for Underwater Scenes

Hongkai Lin, Dingkang Liang, Zhenghao Qi, Xiang Bai

机构 * Huazhong University of Science and Technology(华中科技大学)

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV

Comments Accepted by CVPR 2025. The code is available at https://github.com/HongkLin/TIDE

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.11824 2025-07-29 cs.CV 57%

KITTEN: A Knowledge-Intensive Evaluation of Image Generation on Visual Entities

Hsin-Ping Huang, Xinyi Wang, Yonatan Bitton, Hagai Taitelbaum, Gaurav Singh Tomar, Ming-Wei Chang, Xuhui Jia, Kelvin C. K. Chan, Hexiang Hu, Yu-Chuan Su, Ming-Hsuan Yang

专题命中 多模态生成 :MLLM(abstract);分类 cs.CV

Comments Project page: https://kitten-project.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏