arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4975 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 4975 篇

2511.19835 2025-11-26 cs.CV cs.AI 62%

Rectified SpaAttn: Revisiting Attention Sparsity for Efficient Video Generation

校正SpaAttn:重新审视注意力稀疏性以实现高效的视频生成

Xuewen Liu, Zhikai Li, Jing Zhang, Mengjuan Chen, Qingyi Gu

机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出Rectified SpaAttn,通过校正注意力分配提升视频生成效率,实现显著速度提升且保持生成质量。

Comments Code at https://github.com/BienLuky/Rectified-SpaAttn

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.06250 2025-11-26 cs.CV cs.AI cs.HC 62%

Generative AI for Cel-Animation: A Survey

生成式AI用于动画:一种调查

Yolo Y. Tang, Junjia Guo, Pinxin Liu, Zhiyuan Wang, Hang Hua, Jia-Xing Zhong, Yunzhong Xiao, Chao Huang, Luchuan Song, Susan Liang, Yizhi Song, Liu He, Jing Bi, Mingqian Feng, Xinyang Li, Zeliang Zhang, Chenliang Xu

机构 * University of Rochester(罗切斯特大学) UCSB University of Oxford(牛津大学) CMU(卡内基梅隆大学) Purdue University(普渡大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 本文调查生成式AI如何通过自动化任务革新传统动画流程,降低技术门槛,扩大创作者群体,并促进艺术创新。

Comments Accepted by ICCV 2025 AISTORY Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17986 2025-11-25 cs.CV cs.AI 62%

Plan-X: Instruct Video Generation via Semantic Planning

Plan-X: 通过语义规划指导视频生成

Lun Huang, You Xie, Hongyi Xu, Tianpei Gu, Chenxu Zhang, Guoxian Song, Zenan Li, Xiaochen Zhao, Linjie Luo, Guillermo Sapiro

机构 * Duke University(杜克大学) Princeton University(普林斯顿大学) ByteDance Intelligent Creation(字节跳动智能创作) Apple(苹果公司)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 Plan-X通过语义规划框架减少视频生成中的视觉幻觉,实现与多模态上下文一致的精细指令对齐生成。

Comments The project page is at https://byteaigc.github.io/Plan-X

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12851 2025-11-18 cs.CL cs.AI 62%

NeuroLex: A Lightweight Domain Language Model for EEG Report Understanding and Generation

Kang Yin, Hye-Bin Shin

机构 * Dept. of Artificial Intelligence Korea University Seoul, Republic of Korea(人工智能系韩国大学首尔共和国韩国)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03457 2025-11-18 cs.GR cs.CV cs.SD eess.AS 62%

READ: Real-time and Efficient Asynchronous Diffusion for Audio-driven Talking Head Generation

Haotian Wang, Yuzhe Weng, Jun Du, Haoran Xu, Xiaoyan Wu, Shan He, Bing Yin, Cong Liu, Jianqing Gao, Qingfeng Liu

专题命中 多模态生成 :audio-visual(abstract);分类 cs.CV、eess.AS

Comments Project page: https://readportrait.github.io/READ/

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.15217 2025-11-18 cs.SD cs.AI cs.LG cs.MM 62%

DRAGON: Distributional Rewards Optimize Diffusion Generative Models

Yatong Bai, Jonah Casebeer, Somayeh Sojoudi, Nicholas J. Bryan

机构 * University of California, Berkeley(加州大学伯克利分校) Adobe Research(Adobe研究)

专题命中 多模态生成 :cross-modal(abstract);分类 cs.AI、cs.MM

Comments Accepted to TMLR

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.00998 2025-11-18 cs.CV cs.AI 62%

FBSDiff: Plug-and-Play Frequency Band Substitution of Diffusion Features for Highly Controllable Text-Driven Image Translation

Xiang Gao, Jiaying Liu

机构 * Wangxuan Institute of Computer Technology, Peking University(王轩计算机技术研究所,北京大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Accepted conference paper of ACM MM 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10154 2025-11-14 cs.CV cs.AI 62%

GEA: Generation-Enhanced Alignment for Text-to-Image Person Retrieval

Hao Zou, Runqing Zhang, Xue Zhou, Jianxiao Zou

机构 * School of Automation Engineering, University of Electronic Science and Technology of China(自动化工程学院,电子科学与技术大学) Shenzhen Institute for Advanced Study, University of Electronic Science and Technology of China(深圳高级研究学院,电子科学与技术大学)

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV、cs.AI

Comments 8pages,3figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10020 2025-11-14 cs.CV cs.AI 62%

Anomagic: Crossmodal Prompt-driven Zero-shot Anomaly Generation

Yuxin Jiang, Wei Luo, Hui Zhang, Qiyu Chen, Haiming Yao, Weiming Shen, Yunkang Cao

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.18638 2025-11-14 cs.CR cs.AI cs.CL 62%

Graph of Attacks with Pruning: Optimizing Stealthy Jailbreak Prompt Generation for Enhanced LLM Content Moderation

Daniel Schwartz, Dmitriy Bespalov, Zhe Wang, Ninad Kulkarni, Yanjun Qi

机构 * Amazon Bedrock Science(亚马逊Bedrock科学) Drexel University(德雷塞尔大学) University of Virginia(弗吉尼亚大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CL、cs.AI

Comments 14 pages, 5 figures; published in EMNLP 2025 ; Code at: https://github.com/dsbuddy/GAP-LLM-Safety

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.14905 2025-11-04 eess.IV cs.AI cs.CL 62%

Structural Entities Extraction and Patient Indications Incorporation for Chest X-ray Report Generation

Kang Liu, Zhuoqi Ma, Xiaolu Kang, Zhusi Zhong, Zhicheng Jiao, Grayson Baird, Harrison Bai, Qiguang Miao

机构 * School of Computer Science and Technology, Xidian University(西安电子科技大学计算机科学与技术学院) Xi'an Key Laboratory of Big Data and Intelligent Vision(西安大数据与智能视觉重点实验室) Key Laboratory of Collaborative Intelligence Systems, Ministry of Education, Xidian University(教育部协同智能系统重点实验室) Warren Alpert Medical School, Brown University(布朗大学沃伦·阿尔珀特医学院) School of Electronic Engineering, Xidian University(西安电子科技大学电子工程学院) Department of Radiology and Radiological Sciences, Johns Hopkins University School of Medicine(约翰霍普金斯大学医学院放射科)

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CL、cs.AI

Comments The code is available at https://github.com/mk-runner/SEI-Temp or https://github.com/mk-runner/SEI

Journal ref Medical Image Computing and Computer Assisted Intervention (MICCAI 2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00362 2025-11-04 cs.CV cs.AI cs.GR 62%

Oitijjo-3D: Generative AI Framework for Rapid 3D Heritage Reconstruction from Street View Imagery

Momen Khandoker Ope, Akif Islam, Mohd Ruhul Ameen, Abu Saleh Musa Miah, Md Rashedul Islam, Jungpil Shin

机构 * University of Rajshahi(拉贾沙希大学) Marshall University(马歇尔大学) University of Aizu(御所大学) University of Asia Pacific(亚洲太平洋大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 6 Pages, 4 figures, 2 Tables, Submitted to ICECTE 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00107 2025-11-04 cs.CV cs.AI cs.IR 62%

AI Powered High Quality Text to Video Generation with Enhanced Temporal Consistency

Piyushkumar Patel

机构 * Microsoft(微软)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23020 2025-10-28 cs.CV cs.CL 62%

M$^{3}$T2IBench: A Large-Scale Multi-Category, Multi-Instance, Multi-Relation Text-to-Image Benchmark

Huixuan Zhang, Xiaojun Wan

机构 * Wangxuan Institute of Computer Technology, Peking University(计算机技术研究所,北京大学)

专题命中 多模态生成 :image-text(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.06771 2025-10-27 cs.AI cs.CV cs.LG 62%

Proactive Agents for Multi-Turn Text-to-Image Generation Under Uncertainty

Meera Hahn, Wenjun Zeng, Nithish Kannen, Rich Galt, Kartikeya Badola, Been Kim, Zi Wang

机构 * Google DeepMind(谷歌DeepMind)

专题命中 多模态生成 :image-text(abstract);分类 cs.CV、cs.AI

Journal ref International Conference on Machine Learning, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19641 2025-10-23 cs.CL cs.AI 62%

Style Attack Disguise: When Fonts Become a Camouflage for Adversarial Intent

Yangshijie Zhang, Xinda Wang, Jialin Liu, Wenqiang Wang, Zhicong Ma, Xingxing Jia

机构 * Lanzhou University(兰州大学) Peking University(北京大学) Sun Yat-sen University(中山大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17519 2025-10-23 cs.CV cs.AI 62%

MUG-V 10B: High-efficiency Training Pipeline for Large Video Generation Models

Yongshun Zhang, Zhongyi Fan, Yonghang Zhang, Zhangzikang Li, Weifeng Chen, Zhongwei Feng, Chaoyue Wang, Peng Hou, Anxiang Zeng

机构 * LLM Team, Shopee Pte. Ltd.(Shopee 股份有限公司语言模型团队)

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV、cs.AI

Comments Technical Report; Project Page: https://github.com/Shopee-MUG/MUG-V

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.00939 2025-10-23 cs.CV cs.CL 62%

WikiVideo: Article Generation from Multiple Videos

Alexander Martin, Reno Kriz, William Gantt Walden, Kate Sanders, Hannah Recknor, Eugene Yang, Francis Ferraro, Benjamin Van Durme

机构 * Johns Hopkins University(约翰霍普金斯大学) Human Language Technology Center of Excellence(人机语言技术卓越中心) University of Maryland Baltimore County(马里兰大学巴尔的摩县分校)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.CL

Comments Repo can be found here: https://github.com/alexmartin1722/wikivideo

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.02048 2025-10-22 eess.IV cs.AI cs.CV 62%

Regression is all you need for medical image translation

Sebastian Rassmann, David Kügler, Christian Ewert, Martin Reuter

机构 * German Center for Neurodegenerative Diseases (DZNE)(德国神经退行性疾病研究中心)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16844 2025-10-21 cs.CL cs.AI cs.CE 62%

FinSight: Towards Real-World Financial Deep Research

Jiajie Jin, Yuyao Zhang, Yimeng Xu, Hongjin Qian, Yutao Zhu, Zhicheng Dou

机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学耿丽人工智能学院) BAAI(北京人工智能研究院)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CL、cs.AI

Comments Working in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.15176 2025-10-20 cs.CV cs.AI 62%

Methods and Trends in Detecting AI-Generated Images: A Comprehensive Review

Arpan Mahara, Naphtali Rishe

机构 * Knight Foundation School of Computing and Information Sciences, Florida International University(骑士基金会计算与信息科学学院,佛罗里达国际大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 34 pages, 4 Figures, 10 Tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18668 2025-10-17 cs.CV cs.CL 62%

ChartGalaxy: A Dataset for Infographic Chart Understanding and Generation

Zhen Li, Duan Li, Yukai Guo, Xinyuan Guo, Bowen Li, Lanxi Xiao, Shenyu Qiao, Jiashu Chen, Zijian Wu, Hui Zhang, Xinhuan Shu, Shixia Liu

机构 * Tsinghua University(清华大学) Newcastle University(新castle大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.CL

Comments 58 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.08980 2025-10-14 cs.LG cs.AI cs.CV 62%

Learning Diffusion Models with Flexible Representation Guidance

Chenyu Wang, Cai Zhou, Sharut Gupta, Zongyu Lin, Stefanie Jegelka, Stephen Bates, Tommi Jaakkola

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

Comments NeurIPS 2025; Also Oral at ICML 2025 FM4LS workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05976 2025-10-08 cs.CV cs.AI cs.LG 62%

Diffusion Models for Low-Light Image Enhancement: A Multi-Perspective Taxonomy and Performance Analysis

Eashan Adhikarla, Yixin Liu, Brian D. Davison

机构 * Lehigh University(莱维大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15046 2025-10-08 cs.CL cs.AI 62%

ChartCards: A Chart-Metadata Generation Framework for Multi-Task Chart Understanding

Yifan Wu, Lutao Yan, Leixian Shen, Yinan Mei, Jiannan Wang, Yuyu Luo

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CL、cs.AI

Comments Need to be revised

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04577 2025-10-07 cs.SD cs.LG cs.MM eess.AS 62%

Language Model Based Text-to-Audio Generation: Anti-Causally Aligned Collaborative Residual Transformers

Juncheng Wang, Chao Xu, Cheng Yu, Zhe Hu, Haoyu Xie, Guoqi Yu, Lei Shang, Shujun Wang

机构 * The Hong Kong Polytechnic University(香港理工大学) Alibaba Group(阿里巴巴集团)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.MM、eess.AS

Comments Accepted to EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04498 2025-10-07 cs.CL cs.AI 62%

GenQuest: An LLM-based Text Adventure Game for Language Learners

Qiao Wang, Adnan Labib, Robert Swier, Michael Hofmeyr, Zheng Yuan

机构 * Hosei University(立命馆大学) King’s College London(伦敦大学国王学院) Kindai University(_kindai大学) Tokyo Uni. of Science(东京科学大学) University of Sheffield(谢菲尔德大学)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CL、cs.AI

Comments Workshop on Wordplay: When Language Meets Games, EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04201 2025-10-07 cs.CV cs.AI 62%

World-To-Image: Grounding Text-to-Image Generation with Agent-Driven World Knowledge

Moo Hyun Son, Jintaek Oh, Sun Bin Mun, Jaechul Roh, Sehyun Choi

机构 * The Hong Kong University of Science and Technology(香港科学与技术大学) Georgia Institute of Technology(佐治亚理工学院) University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校) TwelveLabs

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24251 2025-10-07 cs.CV cs.CL 62%

Latent Visual Reasoning

Bangzheng Li, Ximeng Sun, Jiang Liu, Ze Wang, Jialian Wu, Xiaodong Yu, Hao Chen, Emad Barsoum, Muhao Chen, Zicheng Liu

机构 * University of California, Davis(加州大学戴维斯分校) Advanced Micro Devices, Inc.(先进微器件公司)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16612 2025-10-07 cs.HC cs.AI cs.CV cs.CY 62%

Negative Shanshui: Real-time Interactive Ink Painting Synthesis

Aven-Le Zhou

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港理工大学(广州))

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏