arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4975 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 4975 篇

2512.10376 2025-12-12 cs.CV 70%

RaLiFlow: Scene Flow Estimation with 4D Radar and LiDAR Point Clouds

RaLiFlow: 基于4D雷达和激光雷达点云的场景流估计

Jingyun Fu, Zhiyu Xiang, Na Zhao

专题命中 多模态生成 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 RaLiFlow是首个联合4D雷达和激光雷达的场景流学习框架,通过动态感知双向跨模态融合模块和精心设计的损失函数实现高效的雷达-激光雷达融合。

Comments Accepted by AAAI

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05991 2025-12-12 cs.CV 70%

EmoDiffTalk:Emotion-aware Diffusion for Editable 3D Gaussian Talking Head

EmoDiffTalk:面向可编辑3D高斯说话头的情绪感知扩散

Chang Liu, Tianjiao Jing, Chengcheng Ma, Xuanqi Zhou, Zhengxuan Lian, Qin Jin, Hongliang Yuan, Shi-Sheng Huang

机构 * Beijing Normal University(北京师范大学) Renmin University of China(中国人民大学) Tencent AI Lab(腾讯AI实验室)

专题命中 多模态生成 :multimodal(abstract);multi-modal(abstract);分类 cs.CV

AI总结 EmoDiffTalk通过情绪感知高斯扩散实现细粒度面部动画与多模态情绪编辑,提升3D说话头的情绪表达精度和可控性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.09271 2025-12-11 cs.CV 70%

LongT2IBench: A Benchmark for Evaluating Long Text-to-Image Generation with Graph-structured Annotations

LongT2IBench: 一个用于评估长文本到图像生成的基准,具有图结构注释

Zhichao Yang, Tianjiao Gu, Jianjie Wang, Feiyu Lin, Xiangfei Sheng, Pengfei Chen, Leida Li

专题命中 多模态生成 :multi-modal(abstract);image-text(abstract);分类 cs.CV

AI总结 LongT2IBench通过图结构注释和LongT2IExpert提出,用于评估长文本到图像生成的对齐和解释能力。

Comments The paper has been accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04563 2025-12-08 cs.CV 70%

COOPER: A Unified Model for Cooperative Perception and Reasoning in Spatial Intelligence

COOPER:一种用于空间智能中协作感知与推理的统一模型

Zefeng Zhang, Xiangzhao Hao, Hengzhu Tang, Zhenyu Zhang, Jiawei Sheng, Xiaodong Li, Zhenyang Li, Li Gao, Daiting Shi, Dawei Yin, Tingwen Liu

机构 * Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络安全学院) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Baidu Inc.(百度公司)

专题命中 多模态生成 :multimodal(abstract);MLLM(abstract);分类 cs.CV

AI总结 COOPER是一种统一的多模态大语言模型,通过整合深度和分割等辅助模态,提升空间感知与推理能力,实现空间智能的增强。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03623 2025-12-04 cs.LG cs.AI physics.ao-ph 70%

The promising potential of vision language models for the generation of textual weather forecasts

视觉语言模型在生成文本天气预报中的巨大潜力

Edward C. C. Steele, Dinesh Mane, Emilio Monti, Luis Orus, Rebecca Chantrill-Cheyette, Matthew Couch, Kirstine I. Dale, Simon Eaton, Govindarajan Rangarajan, Amir Majlesi, Steven Ramsdale, Michael Sharpe, Craig Smith, Jonathan Smith, Rebecca Yates, Holly Ellis, Charles Ewen

机构 * Met Office(英国气象局) Amazon Web Services(亚马逊网络服务) University of East Anglia(东安格利亚大学)

专题命中 多模态生成 :multimodal(abstract);multimodal foundation model(abstract);分类 cs.AI

AI总结 本文研究了视觉语言模型在直接生成文本天气预报中的潜力,通过视频编码的格网天气数据提升气象服务的生产效率与创新。

Comments 7 pages, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22625 2025-12-02 cs.CV 70%

ReasonEdit: Towards Reasoning-Enhanced Image Editing Models

ReasonEdit: 向推理增强的图像编辑模型迈进

Fukun Yin, Shiyu Liu, Yucheng Han, Zhibo Wang, Peng Xing, Rui Wang, Wei Cheng, Yingming Wang, Aojie Li, Zixin Yin, Pengtao Chen, Xiangyu Zhang, Daxin Jiang, Xianfang Zeng, Gang Yu

专题命中 多模态生成 :multimodal(abstract);MLLM(abstract);分类 cs.CV

AI总结 ReasonEdit通过引入思考和反思机制,提升图像编辑模型的推理能力,实现更准确的编辑效果。

Comments code: https://github.com/stepfun-ai/Step1X-Edit

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00084 2025-12-02 cs.CV cs.LG 70%

A Fast and Efficient Modern BERT based Text-Conditioned Diffusion Model for Medical Image Segmentation

一种快速且高效的基于现代BERT的文本条件扩散模型用于医学图像分割

Venkata Siddharth Dhara, Pawan Kumar

机构 * International Institute of Information Technology, Hyderabad, 500032, India(国际信息科技学院,海得拉巴)

专题命中 多模态生成 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出FastTextDiff,利用ModernBERT提升医学图像分割的效率和准确性,通过整合文本注释和多模态注意力机制改进传统扩散模型。

Comments 15 pages, 3 figures, Accepted in Slide 3 10th International Conference on Computer Vision & Image Processing (CVIP 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.23469 2025-12-01 cs.CV 70%

Visual Generation Tuning

视觉生成微调

Jiahao Guo, Sinan Du, Jingfeng Yao, Wenyu Liu, Bo Li, Haoxiang Cao, Kun Gai, Chun Yuan, Kai Wu, Xinggang Wang

机构 * Huazhong University of Science and Technology(华中科技大学) Tsinghua University(清华大学) School of Artificial Intelligence, South China Normal University(华南师范大学人工智能学院) Kolors Team, Kuaishou Technology(快手科技Kolors团队)

专题命中 多模态生成 :multimodal(abstract);multimodal foundation model(abstract);分类 cs.CV

AI总结 本文提出VGT,通过视觉生成微调提升视觉语言模型的视觉生成能力,在图像重建和生成任务中均取得优异成果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22896 2025-12-01 cs.CV 70%

DM$^3$T: Harmonizing Modalities via Diffusion for Multi-Object Tracking

DM$^3$T: 通过扩散和谐多模态以实现多目标跟踪

Weiran Li, Yeqiang Liu, Yijie Wei, Mina Han, Qiannan Guo, Zhenbo Li

机构 * China Agricultural University(中国农业大学) Beijing Normal University(北京师范大学)

专题命中 多模态生成 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 DM$^3$T通过扩散模型实现多模态特征对齐,提升多目标跟踪的准确性和鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21579 2025-12-01 cs.CV 70%

Harmony: Harmonizing Audio and Video Generation through Cross-Task Synergy

Harmony: 通过跨任务协同实现音频和视频生成

Teng Hu, Zhentao Yu, Guozhen Zhang, Zihan Su, Zhengguang Zhou, Youliang Zhang, Yuan Zhou, Qinglin Lu, Ran Yi

机构 * Shanghai Jiao Tong University(上海交通大学) Tencent Hunyuan Project(腾讯文心项目)

专题命中 多模态生成 :cross-modal(abstract);audio-visual(abstract);分类 cs.CV

AI总结 Harmony通过跨任务协同机制和同步增强CFG,实现高效音频视频同步生成。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14760 2025-11-19 cs.CV 70%

UniGen-1.5: Enhancing Image Generation and Editing through Reward Unification in Reinforcement Learning

Rui Tian, Mingfei Gao, Haiming Gang, Jiasen Lu, Zhe Gan, Yinfei Yang, Zuxuan Wu, Afshin Dehghan

机构 * Institute of Trustworthy Embodied AI, Fudan University(可信具身人工智能研究院,复旦大学) Apple(苹果公司)

专题命中 多模态生成 :multimodal(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01119 2025-11-19 cs.CV cs.LG 70%

The Promise of RL for Autoregressive Image Editing

Saba Ahmadi, Rabiul Awal, Ankur Sikarwar, Amirhossein Kazemnejad, Ge Ya Luo, Juan A. Rodriguez, Sai Rajeswar, Siva Reddy, Christopher Pal, Benno Krojer, Aishwarya Agrawal

机构 * Mila – Quebec AI Institute(魁北克AI研究所) Université de Montréal(蒙特利尔大学) McGill University(麦吉尔大学) École de Technologie Supérieure (ETS)(高等技术学院) Polytechnique Montréal(蒙特利尔理工学院) ServiceNow(ServiceNow公司) Canada CIFAR AI Chair(加拿大CIFAR人工智能主席)

专题命中 多模态生成 :multimodal(abstract);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12363 2025-11-18 cs.CV 70%

Explainable AI-Generated Image Detection RewardBench

Michael Yang, Shijian Deng, William T. Doan, Kai Wang, Tianyu Yang, Harsh Singh, Yapeng Tian

机构 * The University of Texas at Dallas(德克萨斯大学达拉斯分校) University of Toronto(多伦多大学) University of Notre Dame(诺特大学) Stony Brook University(石溪大学)

专题命中 多模态生成 :multimodal(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11434 2025-11-17 cs.CV 70%

WEAVE: Unleashing and Benchmarking the In-context Interleaved Comprehension and Generation

Wei Chow, Jiachun Pan, Yongyuan Liang, Mingze Zhou, Xue Song, Liyu Jia, Saining Zhang, Siliang Tang, Juncheng Li, Fengda Zhang, Weijia Wu, Hanwang Zhang, Tat-Seng Chua

机构 * National University of Singapore(新加坡国立大学) Nanyang Technological University(南洋理工大学) University of Maryland, College Park(马里兰大学学院公园分校) Zhejiang University(浙江大学)

专题命中 多模态生成 :multimodal(abstract);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00596 2025-11-11 cs.CV 70%

Seg2Any: Open-set Segmentation-Mask-to-Image Generation with Precise Shape and Semantic Control

Danfeng Li, Hui Zhang, Sheng Wang, Jiacheng Li, Zuxuan Wu

机构 * Shanghai Key Lab of Intell. Info. Processing, School of CS, Fudan University(上海智能信息处理关键实验室,复旦大学计算机学院) Shanghai Collaborative Innovation Center of Intelligent Visual Computing(上海智能视觉计算协同创新中心) HiThink Research(HiThink研究机构)

专题命中 多模态生成 :multimodal(abstract);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03757 2025-11-07 cs.LG cs.AI 70%

Laugh, Relate, Engage: Stylized Comment Generation for Short Videos

Xuan Ouyang, Senan Wang, Bouzhou Wang, Siyuan Xiahou, Jinrong Zhou, Yuekang Li

机构 * University of New South Wales(新南威尔士大学) University of Sydney(悉尼大学) The University of Hong Kong(香港大学) University of Southern California(南加州大学)

专题命中 多模态生成 :multimodal(abstract);MLLM(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16495 2025-11-06 cs.CV 70%

ALTo: Adaptive-Length Tokenizer for Autoregressive Mask Generation

Lingfeng Wang, Hualing Lin, Senda Chen, Tao Wang, Changxu Cheng, Yangyang Zhong, Dong Zheng, Wuyue Zhao

机构 * Uni-Ubi Zhejiang University(浙江大学) Tongji University(同济大学)

专题命中 多模态生成 :multimodal(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02468 2025-11-05 cs.HC cs.CV 70%

HAGI++: Head-Assisted Gaze Imputation and Generation

Chuhan Jiao, Zhiming Hu, Andreas Bulling

机构 * University of Stuttgart(斯图加特大学) The Hong Kong University of Science(香港科学大学)

专题命中 多模态生成 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

Comments Extended version of our UIST'25 paper "HAGI: Head-Assisted Gaze Imputation for Mobile Eye Trackers"

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01593 2025-11-04 cs.CV 70%

Wave-Particle (Continuous-Discrete) Dualistic Visual Tokenization for Unified Understanding and Generation

Yizhu Chen, Chen Ju, Zhicheng Wang, Shuai Xiao, Xu Chen, Jinsong Lan, Xiaoyong Zhu, Ying Chen

机构 * Zhejiang University(浙江大学) Alibaba Group(阿里巴巴集团) Peking University(北京大学)

专题命中 多模态生成 :multi-modal(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.13267 2025-10-31 cs.CV 70%

Dynamic Traceback Learning for Medical Report Generation

Shuchang Ye, Mingyuan Meng, Mingjian Li, Dagan Feng, Usman Naseem, Jinman Kim

专题命中 多模态生成 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted to IEEE Transactions on Multimedia (TMM)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20803 2025-10-24 cs.CV 70%

ARGenSeg: Image Segmentation with Autoregressive Image Generation Model

Xiaolong Wang, Lixiang Ru, Ziyuan Huang, Kaixiang Ji, Dandan Zheng, Jingdong Chen, Jun Zhou

机构 * Ant Group(蚂蚁集团)

专题命中 多模态生成 :multimodal(abstract);MLLM(abstract);分类 cs.CV

Comments Accepted to NeurIPS 2025, 18 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01480 2025-10-22 cs.CV 70%

Janus-Pro-R1: Advancing Collaborative Visual Comprehension and Generation via Reinforcement Learning

Kaihang Pan, Yang Wu, Wendong Bu, Kai Shen, Juncheng Li, Yingting Wang, Yunfei Li, Siliang Tang, Jun Xiao, Fei Wu, Hang Zhao, Yueting Zhuang

机构 * Zhejiang University(浙江大学) Ant Group(蚂蚁集团)

专题命中 多模态生成 :multimodal(abstract);MLLM(abstract);分类 cs.CV

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12325 2025-10-15 cs.IR cs.AI 70%

Causal Inspired Multi Modal Recommendation

Jie Yang, Chenyang Gu, Zixuan Liu

机构 * National University of Singapore Master of Industrial and Systems Engineering(新加坡国立大学工业与系统工程硕士) East China Normal University Master of Library and Information Science(华东师范大学图书馆与信息科学硕士) Shandong Normal University(山东师范大学)

专题命中 多模态生成 :multimodal(abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12000 2025-10-15 cs.SD cs.CL cs.LG 70%

UALM: Unified Audio Language Model for Understanding, Generation and Reasoning

Jinchuan Tian, Sang-gil Lee, Zhifeng Kong, Sreyan Ghosh, Arushi Goel, Chao-Han Huck Yang, Wenliang Dai, Zihan Liu, Hanrong Ye, Shinji Watanabe, Mohammad Shoeybi, Bryan Catanzaro, Rafael Valle, Wei Ping

机构 * CMU(卡内基梅隆大学) NVIDIA(英伟达) UMD(马里兰大学)

专题命中 多模态生成 :multimodal(abstract);cross-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10633 2025-10-14 cs.AI 70%

Collaborative Text-to-Image Generation via Multi-Agent Reinforcement Learning and Semantic Fusion

Jiabao Shi, Minfeng Qi, Lefeng Zhang, Di Wang, Yingjie Zhao, Ziying Li, Yalong Xing, Ningran Li

机构 * Minzu University of China(民族大学) City University of Macau(澳门城市大学) Key Laboratory of Computing Power Network and Information Security, Ministry of Education, Shandong Computer Science Center (National Supercomputer Center in Jinan), Qilu University of Technology (Shandong Academy of Sciences)(计算能力网络与信息安全重点实验室,教育部,山东计算机科学中心(济南国家超级计算机中心),齐鲁工业大学(山东省科学院)) The University of Adelaide(阿德莱德大学)

专题命中 多模态生成 :multimodal(abstract);cross-modal(abstract);分类 cs.AI

Comments 16 pages, 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07249 2025-10-14 cs.CV 70%

TalkCuts: A Large-Scale Dataset for Multi-Shot Human Speech Video Generation

Jiaben Chen, Zixin Wang, Ailing Zeng, Yang Fu, Xueyang Yu, Siyuan Cen, Julian Tanke, Yihang Chen, Koichi Saito, Yuki Mitsufuji, Chuang Gan

机构 * UMass Amherst(马萨诸塞大学阿姆赫斯特分校) Sony AI(索尼人工智能) UC San Diego(加州大学圣地亚哥分校)

专题命中 多模态生成 :multimodal(abstract);multi-modal(abstract);分类 cs.CV

Comments Project page: https://talkcuts.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05096 2025-10-10 cs.CV cs.AI cs.CL cs.MA cs.MM 70%

Paper2Video: Automatic Video Generation from Scientific Papers

Zeyu Zhu, Kevin Qinghong Lin, Mike Zheng Shou

机构 * Show Lab, National University of Singapore(展示实验室,新加坡国立大学)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Project Page: https://showlab.github.io/Paper2Video/

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21653 2025-10-08 cs.CV 70%

Think Before You Diffuse: Infusing Physical Rules into Video Diffusion

Ke Zhang, Cihan Xiao, Jiacong Xu, Yiqun Mei, Vishal M. Patel

专题命中 多模态生成 :multimodal(abstract);MLLM(abstract);分类 cs.CV

Comments 19 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03341 2025-10-07 cs.CV 70%

OpusAnimation: Code-Based Dynamic Chart Generation

Bozheng Li, Miao Yang, Zhenhan Chen, Jiawang Cao, Mushui Liu, Yi Lu, Yongliang Wu, Bin Zhang, Yangguang Ji, Licheng Tang, Jay Wu, Wenbo Zhu

机构 * Opus AI Research(Opus人工智能研究机构) Brown University(布朗大学) Zhejiang University(浙江大学) University of Toronto(多伦多大学)

专题命中 多模态生成 :multi-modal(abstract);MLLM(abstract);分类 cs.CV

Comments working in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25047 2025-09-30 cs.AI 70%

Scaling Synthetic Task Generation for Agents via Exploration

Ram Ramrakhya, Andrew Szot, Omar Attia, Yuhao Yang, Anh Nguyen, Bogdan Mazoure, Zhe Gan, Harsh Agrawal, Alexander Toshev

机构 * Apple(苹果公司)

专题命中 多模态生成 :multimodal(abstract);MLLM(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏