arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4959 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 4959 篇

2503.08096 2025-12-03 cs.CV 57%

MegaSR: Mining Customized Semantics and Expressive Guidance for Real-World Image Super-Resolution

MegaSR: 为真实世界图像超分辨率挖掘定制语义和表达引导

Xinrui Li, Jinrong Zhang, Jianlong Wu, Chong Chen, Liqiang Nie, Zhouchen Lin

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

AI总结 MegaSR通过定制语义和表达引导提升真实世界图像超分辨率的重建质量与结构一致性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.06263 2025-12-02 cs.CV 57%

OmniSVG: A Unified Scalable Vector Graphics Generation Model

OmniSVG: 一种统一的可扩展矢量图形生成模型

Yiying Yang, Wei Cheng, Sijin Chen, Xianfang Zeng, Fukun Yin, Jiaxu Zhang, Liao Wang, Gang Yu, Xingjun Ma, Yu-Gang Jiang

机构 * Fudan University(复旦大学) StepFun Project(StepFun项目)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

AI总结 OmniSVG通过统一框架和预训练视觉-语言模型,实现高效多模态SVG生成,提升复杂结构的表达能力,并引入大规模数据集推动SVG合成发展。

Comments 20 pages; Project Page: https://omnisvg.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22643 2025-12-02 cs.CV 57%

SPIRAL: Semantic-Aware Progressive LiDAR Scene Generation and Understanding

SPIRAL: 基于语义的渐进式LiDAR场景生成与理解

Dekai Zhu, Yixuan Hu, Youquan Liu, Dongyue Lu, Lingdong Kong, Slobodan Ilic

机构 * Technical University of Munich(慕尼黑技术大学) Fudan University(复旦大学) National University of Singapore(新加坡国立大学) Munich Center for Machine Learning(慕尼黑机器学习中心)

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV

AI总结 Spiral提出了一种基于范围视图的LiDAR扩散模型,能够同时生成深度、反射图像和语义地图,实现高效的3D场景生成与理解。

Comments NeurIPS 2025; 24 pages, 10 figures, 9 tables; Code at https://github.com/worldbench/SPIRAL

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.10349 2025-12-02 cs.RO cs.CV 57%

Ensuring Force Safety in Vision-Guided Robotic Manipulation via Implicit Tactile Calibration

通过隐式触觉校准确保视觉引导的机器人操作中的力安全

Lai Wei, Jiahua Ma, Yibo Hu, Ruimao Zhang

机构 * Sun Yat-sen University(中山大学) UC San Diego(加州大学圣地亚哥分校) Zhejiang University(浙江大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

AI总结 本文提出SafeDiff框架,通过隐式触觉校准提升机器人操作中的力安全,结合视觉和触觉数据优化状态规划,减少有害力风险。

Comments Website URL: see https://i-am-future.github.io/safediff/

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.07507 2025-12-02 cs.CL 57%

Thought2Text: Text Generation from EEG Signal using Large Language Models (LLMs)

Thought2Text: 使用大型语言模型(LLMs)从EEG信号生成文本

Abhijit Mishra, Shreya Shukla, Jose Torres, Jacek Gwizdka, Shounak Roychowdhury

机构 * School of Information University of Texas at Austin(信息学院 德克萨斯大学奥斯汀分校)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CL

AI总结 Thought2Text利用大型语言模型从EEG信号生成文本,通过多阶段微调实现脑活动的可理解表达。

Comments Accepted to Findings of NAACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.23377 2025-12-01 cs.CV 57%

DEAL-300K: Diffusion-based Editing Area Localization with a 300K-Scale Dataset and Frequency-Prompted Baseline

DEAL-300K:基于扩散的编辑区域定位与30万规模数据集及频率提示基线

Rui Zhang, Hongxia Wang, Hangqing Liu, Yang Zhou, Qiang Zeng

机构 * School of Cyber Science and Engineering, Sichuan University(四川大学计算机科学与工程学院) Key Laboratory of Data Protection and Intelligent Management, Ministry of Education, Sichuan University(教育部数据保护与智能管理重点实验室)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

AI总结 DEAL-300K通过多模态大语言模型生成编辑指令,结合频率提示微调方法,实现了基于扩散的图像编辑区域定位,提供高精度的基线和研究基础。

Comments 13pages,12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22863 2025-12-01 cs.CV 57%

CoordSpeaker: Exploiting Gesture Captioning for Coordinated Caption-Empowered Co-Speech Gesture Generation

CoordSpeaker: 利用手势描述生成协调的 caption-驱动的语音同步手势生成

Fengyi Fang, Sicheng Yang, Wenming Yang

机构 * Tsinghua University(清华大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

AI总结 CoordSpeaker通过引入手势描述框架和条件潜在扩散模型,实现了协调的caption驱动的语音同步手势生成,解决了手势生成中的语义先验间隙和多模态控制难题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22773 2025-12-01 cs.RO cs.AI 57%

CAPE: Context-Aware Diffusion Policy Via Proximal Mode Expansion for Collision Avoidance

CAPE:通过近端模式扩展实现上下文感知的扩散策略用于避障

Rui Heng Yang, Xuan Zhao, Leo Maxime Brunswic, Montgomery Alban, Mateo Clemente, Tongtong Cao, Jun Jin, Amir Rasouli

机构 * Huawei Technologies Canada(华为加拿大技术有限公司) Huawei(华为) Department of Electrical and Computer Engineering, University of Alberta(阿尔伯塔大学电气与计算机工程系)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.AI

AI总结 CAPE通过近端模式扩展实现上下文感知的扩散策略,提升避障任务中轨迹生成的泛化能力与碰撞规避性能。

Comments 4 tables, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23240 2025-12-01 cs.CV 57%

Autoregressive Styled Text Image Generation, but Make it Reliable

自回归风格文本图像生成,但使其更可靠

Carmine Zaccagnino, Fabio Quattrini, Vittorio Pippi, Silvia Cascianelli, Alessio Tonioni, Rita Cucchiara

机构 * University of Modena and Reggio Emilia(摩德纳和雷吉奥艾米利亚大学) Google(谷歌)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

AI总结 本文提出Eruku方法,通过多模态提示条件生成任务和分类器自由引导策略,改进自回归模型以生成更可靠、更忠实于文本提示的风格化文本图像。

Comments Accepted at WACV2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00802 2025-12-01 cs.CV 57%

TRACE: Temporally Reliable Anatomically-Conditioned 3D CT Generation with Enhanced Efficiency

TRACE: 基于时间可靠性的解剖条件3D CT生成与增强效率

Minye Shao, Xingyu Miao, Haoran Duan, Zeyu Wang, Jingkun Chen, Yawen Huang, Xian Wu, Jingjing Deng, Yang Long, Yefeng Zheng

机构 * Department of Computer Science, Durham University(杜伦大学计算机科学系) Department of Automation, Tsinghua University(清华大学自动化系) College of Computer Science and Engineering, Dalian Minzu University(大连民族大学计算机科学与工程学院) Department of Engineering Science, University of Oxford(牛津大学工程科学系) Jarvis Research Center, Tencent YouTu Lab(腾讯YouTu实验室 Jarvis 研究中心) School of Engineering Mathematics and Technology, University of Bristol(布里斯托大学工程数学与技术学院) Medical Artificial Intelligence Laboratory, School of Engineering, Westlake University(西湖大学工程学院医学人工智能实验室)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

AI总结 TRACE通过2D多模态条件扩散方法生成具有时空对齐的3D CT图像,提升生成效率和解剖保真度。

Comments Accepted to MICCAI 2025 (this version is not peer-reviewed; it is the extended version)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21780 2025-12-01 cs.MM cs.SD 57%

3MDiT: Unified Tri-Modal Diffusion Transformer for Text-Driven Synchronized Audio-Video Generation

3MDiT:用于文本驱动同步音频视频生成的统一三模态扩散变换器

Yaoru Li, Heyu Si, Federico Landi, Pilar Oplustil Gallegos, Ioannis Koutsoumpas, O. Ricardo Cortez Vazquez, Ruiju Fu, Qi Guo, Xin Jin, Shunyu Liu, Mingli Song

机构 * Zhejiang University(浙江大学) Huawei(华为) Nanyang Technological University(南洋理工大学)

专题命中 多模态生成 :cross-modal(abstract);分类 cs.MM

AI总结 3MDiT提出了一种统一三模态扩散变换器,通过联合演变流实现文本驱动的同步音频视频生成,提升多模态对齐与同步性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21746 2025-12-01 cs.CL 57%

DELTA: Language Diffusion-based EEG-to-Text Architecture

DELTA: 基于语言扩散的EEG到文本架构

Mingyu Jeon, Hyobin Kim

机构 * Sungkyunkwan University(成均馆大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CL

AI总结 DELTA通过结合残差向量量化和语言扩散模型,提升了EEG到文本的语义对齐和生成质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10382 2025-11-27 cs.CV 57%

Towards Spatially Consistent Image Generation: On Incorporating Intrinsic Scene Properties into Diffusion Models

迈向空间一致的图像生成:将内在场景属性纳入扩散模型

Hyundo Lee, Suhyung Choi, Inwoo Hwang, Byoung-Tak Zhang

专题命中 多模态生成 :image-text(abstract);分类 cs.CV

AI总结 本文提出一种基于内在场景属性的图像生成方法,通过同时生成图像和其内在属性,提升生成图像的空间一致性和真实感。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.20280 2025-11-26 cs.CV 57%

Bootstrapping Physics-Grounded Video Generation through VLM-Guided Iterative Self-Refinement

通过VLM引导的迭代自优化提升物理导向的视频生成

Yang Liu, Xilin Zhao, Peisong Wen, Siran Dai, Qingming Huang

机构 * School of Computer Science and Technology, University of Chinese Academy of Sciences(中国科学院大学计算机科学与技术学院) School of Computer Science and Technology, Beijing Institute of Technology(北京理工大学计算机科学与技术学院) Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络安全学院) Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

AI总结 本文提出一种通过VLM引导的迭代自优化方法,提升视频生成的物理一致性,实验显示在PhyIQ基准上得分提升明显。

Comments ICCV 2025 Physics-IQ Challenge Third Place Solution

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.16655 2025-11-26 cs.CV 57%

MovieDreamer: Hierarchical Generation for Coherent Long Visual Sequence

MovieDreamer:用于生成连贯长视觉序列的分层生成

Canyu Zhao, Mingyu Liu, Wen Wang, Weihua Chen, Fan Wang, Hao Chen, Bo Zhang, Chunhua Shen

机构 * Zhejiang University(浙江大学) Alibaba Group(阿里巴巴集团)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

AI总结 MovieDreamer通过结合自回归模型与扩散渲染,实现长时长视频生成,提升叙事连贯性和视觉质量。

Comments 30 pages, 22 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18516 2025-11-25 cs.CV 57%

Breaking Forgetting: Training-Free Few-Shot Class-Incremental Learning via Conditional Diffusion

打破遗忘:通过条件扩散实现无训练的少样本类增量学习

Haidong Kang, Ketong Qian, Yi Lu

机构 * Northeastern University(东北大学) School of Information and Intelligent Science(信息与智能科学学院) Whiting School of Engineering(工程学院)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

AI总结 本文提出无训练的少样本类增量学习方法,通过条件扩散过程替代梯度优化,缓解灾难性遗忘并提升泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10353 2025-11-25 cs.CV 57%

Motion-R1: Enhancing Motion Generation with Decomposed Chain-of-Thought and RL Binding

Motion-R1: 通过分解的链式推理与RL绑定增强运动生成

Runqi Ouyang, Haoyun Li, Zhenyuan Zhang, Xiaofeng Wang, Zeyu Zhang, Zheng Zhu, Guan Huang, Sirui Han, Xingang Wang

机构 * GigaAI CASIA HKUST(香港科技大学)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

AI总结 Motion-R1通过结合分解的链式推理与强化学习,提升运动生成的质量和可解释性,实现多项指标的性能提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03596 2025-11-25 cs.CV 57%

ControlThinker: Unveiling Latent Semantics for Controllable Image Generation through Visual Reasoning

ControlThinker: 通过视觉推理揭示潜在语义以实现可控图像生成

Feng Han, Yang Jiao, Shaoxiang Chen, Junhao Xu, Jingjing Chen, Yu-Gang Jiang

机构 * Shanghai Key Lab of Intell. Info. Processing, School of CS, Fudan University(上海智能信息处理关键实验室,复旦大学计算机学院) Shanghai Collaborative Innovation Center on Intelligent Visual Computing(上海智能视觉计算协同创新中心) MiniMax

专题命中 多模态生成 :MLLM(abstract);分类 cs.CV

AI总结 ControlThinker通过视觉推理挖掘潜在语义,提升可控图像生成的语义一致性和视觉质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.03931 2025-11-25 cs.CV 57%

MagicMirror: ID-Preserved Video Generation in Video Diffusion Transformers

MagicMirror: 视频扩散变换器中的身份保留视频生成

Yuechen Zhang, Yaoyang Liu, Bin Xia, Bohao Peng, Zexin Yan, Eric Lo, Jiaya Jia

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV

AI总结 MagicMirror通过双分支特征提取器和两阶段训练策略,在视频扩散变换器中实现高质量身份保留视频生成,优于现有方法。

Comments ICCV 2025, It is best viewed in Acrobat. Project Page: https://julianjuaner.github.io/projects/MagicMirror/

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17501 2025-11-24 cs.CV cs.GR 57%

Native 3D Editing with Full Attention

原生3D编辑与全关注

Weiwei Cai, Shuangkang Fang, Weicai Ye, Xin Dong, Yunhan Yang, Xuanyang Zhang, Wei Cheng, Yanpei Cao, Gang Yu, Tao Chen

机构 * Fudan University(复旦大学) StepFun, Inc.(StepFun公司) Zhejiang University(浙江大学) Tsinghua University(清华大学) VAST

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

AI总结 本文提出一种高效的原生3D编辑框架,通过多模态数据集和3D标记连接方法,提升3D编辑的效率和一致性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00737 2025-11-24 cs.HC cs.AI 57%

How LLMs are Shaping the Future of Virtual Reality

大型语言模型如何塑造虚拟现实的未来

Süeda Özkaya, Santiago Berrezueta-Guzman, Stefan Wagner

机构 * Technical University of Munich(慕尼黑技术大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

AI总结 本文探讨了大型语言模型如何通过增强叙事生成、NPC互动和个性化来塑造虚拟现实的未来,并提出了多模态AI和伦理保障等未来研究方向。

Comments Pre-print

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23584 2025-11-21 cs.CV 57%

VividFace: High-Quality and Efficient One-Step Diffusion For Video Face Enhancement

VividFace: 高质量和高效的一步扩散用于视频面部增强

Shulian Zhang, Yong Guo, Long Peng, Ziyang Wang, Ye Chen, Wenbo Li, Xiao Zhang, Yulun Zhang, Jian Chen

机构 * South China University of Technology(华南理工大学) Max Planck Institute for Informatics(马克斯·普朗克研究所(信息学)) University of Science and Technology of China(中国科学技术大学) The Chinese University of Hong Kong(香港中文大学) Nanjing University of Science and Technology(南京理工大学) Shanghai Jiao Tong University(上海交通大学)

专题命中 多模态生成 :MLLM(abstract);分类 cs.CV

AI总结 VividFace通过单步扩散框架和联合训练策略,高效提升视频面部增强的高质量与效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.14023 2025-11-21 cs.CL 57%

Synthetic Data Generation Using Large Language Models: Advances in Text and Code

利用大语言模型生成合成数据:文本与代码领域的进展

Mihai Nadas, Laura Diosan, Andreea Tomescu

机构 * Faculty of Mathematics and Computer Science, Babeş-Bolyai University(巴贝什-博耶亚大学数学与计算机科学系) KlusAI Research Lab(KlusAI研究实验室)

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CL

AI总结 本文探讨了利用大语言模型生成合成数据在文本和代码领域的新进展,分析了其在低资源任务和代码应用中的潜力及挑战。

Comments 24 pages, 6 tables, 1 figure, 64 references

Journal ref IEEE Access 13, 134615-134633 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15520 2025-11-20 cs.RO cs.AI 57%

Theoretical Closed-loop Stability Bounds for Dynamical System Coupled with Diffusion Policies

动态系统与扩散策略耦合下的理论闭环稳定性边界

Gabriel Lauzier, Alexandre Girard, François Ferland

机构 * Department of Mechanical Engineering, Universite de Sherbrooke(机械工程系, Sherbrooke 大学) Department of Electronic and Computer Engineering, Universite de Sherbrooke(电子与计算机工程系, Sherbrooke 大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

AI总结 本文研究了动态系统与扩散策略耦合下的闭环稳定性边界,提出了一种更快的模仿学习框架和基于演示方差的稳定性判断指标。

Comments 5 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.02249 2025-11-20 cs.RO cs.AI 57%

Natural Selection via Foundation Models for Soft Robot Evolution

Changhe Chen, Xiaohao Xu, Xiangdong Wang, Xiaonan Huang

机构 * University of Michigan-Ann Arbor(密歇根大学安娜堡分校)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.20105 2025-11-19 cs.CV 57%

FreeSeg-Diff: Training-Free Open-Vocabulary Segmentation with Diffusion Models

Barbara Toniella Corradini, Mustafa Shukor, Paul Couairon, Guillaume Couairon, Franco Scarselli, Matthieu Cord

机构 * DIISM University of Siena(DIISM锡耶纳大学) CNRS, ISIR Sorbonne University(CNRS,ISIR索邦大学) Inria, ARCHES Sorbonne University(Inria,ARCHES索邦大学) Valeo.ai Sorbonne University(Valeo.ai索邦大学)

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV

Journal ref Proceedings of the 2025 International Joint Conference on Neural Networks (IJCNN 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13309 2025-11-18 cs.CV 57%

DriveLiDAR4D: Sequential and Controllable LiDAR Scene Generation for Autonomous Driving

Kaiwen Cai, Xinze Liu, Xia Zhou, Hengtong Hu, Jie Xiang, Luyao Zhang, Xueyang Zhang, Kun Zhan, Yifei Zhan, Xianpeng Lang

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments AAAI2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12428 2025-11-18 cs.CV 57%

RedVTP: Training-Free Acceleration of Diffusion Vision-Language Models Inference via Masked Token-Guided Visual Token Pruning

Jingqi Xu, Jingxi Lu, Chenghao Li, Sreetama Sarkar, Souvik Kundu, Peter A. Beerel

机构 * University of Southern California(南加州大学) Intel Labs(英特尔实验室)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12040 2025-11-18 cs.CV 57%

SRSplat: Feed-Forward Super-Resolution Gaussian Splatting from Sparse Multi-View Images

Xinyuan Hu, Changyue Shi, Chuxiao Yang, Minghao Chen, Jiajun Ding, Tao Wei, Chen Wei, Zhou Yu, Min Tan

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments AAAI2026-Oral. Project Page: https://xinyuanhu66.github.io/SRSplat/

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.07334 2025-11-17 cs.CV 57%

Unleashing the Potential of Large Language Models for Text-to-Image Generation through Autoregressive Representation Alignment

Xing Xie, Jiawei Liu, Ziyue Lin, Huijie Fan, Zhi Han, Yandong Tang, Liangqiong Qu

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV

Comments Accepted by AAAI 2026 Oral

详情

展开后加载摘要…

URL PDF HTML 收藏