arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

共收录 86872 信号源:cs.CV, cs.GR, cs.MM

1. 扩散模型 70277 篇

2511.19319 2026-03-09 cs.CV 79%

SyncMV4D: Synchronized Multi-view Joint Diffusion of Appearance and Motion for Hand-Object Interaction Synthesis

SyncMV4D: 同步多视角联合生成外观与运动的 hand-object 交互合成

Lingwei Dang, Zonghan Li, Juntong Li, Hongwen Zhang, Liang An, Yebin Liu, Qingyao Wu

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 SyncMV4D通过联合生成多视角手-物体交互视频和4D运动,提升了视觉真实性和运动合理性。

Comments The structure and logic of writing will undergo a complete revision

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15335 2026-03-09 cs.CV cs.AI 79%

ExDD: Explicit Dual Distribution Learning for Surface Defect Detection via Diffusion Synthesis

ExDD:通过扩散合成实现表面缺陷检测的显式双分布学习

Muhammad Aqeel, Federico Leonardi, Francesco Setti

机构 * Dept. of Engineering for Innovation Medicine, University of Verona(创新医学工程系,威尼斯大学) Qualyco S.r.l.(Qualyco公司)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 ExDD通过显式建模双分布和潜在扩散模型生成合成缺陷,提升工业表面缺陷检测的准确率和鲁棒性。

Comments Accepted to ICIAP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18663 2026-03-09 cs.CV 79%

DVD-Quant: Data-free Video Diffusion Transformers Quantization

DVD-Quant: 数据无依赖的视频扩散变换器量化

Zhiteng Li, Hanxuan Li, Junyi Wu, Kai Liu, Haotong Qin, Linghe Kong, Guihai Chen, Yulun Zhang, Xiaokang Yang

机构 * Shanghai Jiao Tong University(上海交通大学) Zhejiang University(浙江大学) ETH Zürich(苏黎世联邦理工学院)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 DVD-Quant通过无数据量化框架实现视频DiT模型的高效量化,提升运行速度并保持视觉质量。

Comments Code and models will be available at https://github.com/lhxcs/DVD-Quant

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.19509 2026-03-09 cs.CV cs.LG cs.SD eess.AS 79%

Ditto: Motion-Space Diffusion for Controllable Realtime Talking Head Synthesis

Ditto:用于可控实时说话头合成的运动空间扩散

Tianqi Li, Ruobing Zheng, Minghui Yang, Jingdong Chen, Ming Yang

机构 * Ant Group(蚂蚁集团)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 Ditto通过运动空间扩散模型实现可控实时说话头合成,优化架构与训练策略以提升生成质量与实时性能。

Comments Project Page: https://digital-avatar.github.io/ai/Ditto/

Journal ref ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.05315 2026-03-06 cs.CV 79%

Frequency-Aware Error-Bounded Caching for Accelerating Diffusion Transformers

面向频率的误差有界缓存用于加速扩散变换器

Guandong Li

机构 * iFLYTEK

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 SpectralCache通过统一缓存框架提升扩散变换器推理速度,实现2.46倍加速,质量与TeaCache相当。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.04638 2026-03-06 cs.CV cs.LG q-bio.QM 79%

Spinverse: Differentiable Physics for Permeability-Aware Microstructure Reconstruction from Diffusion MRI

Spinverse: 可微物理用于从扩散MRI中感知渗透性的微结构重建

Prathamesh Pradeep Khole, Mario M. Brenes, Zahra Kais Petiwala, Ehsan Mirafzali, Utkarsh Gupta, Jing-Rebecca Li, Andrada Ianus, Razvan Marinescu

机构 * University of California Santa Cruz(加州大学圣克鲁兹分校) University of California San Diego(加州大学圣地亚哥分校) Inria-Saclay(Inria-萨克莱实验室) Kings College London(伦敦国王学院)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 Spinverse通过可微物理方法从扩散MRI中感知渗透性,重建微结构并优化渗透性参数以提高边界准确性和结构有效性。

Comments 10 Pages, 5 Figures, 2 Tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23589 2026-03-06 cs.AI cs.CV cs.LG 79%

BridgeDrive: Diffusion Bridge Policy for Closed-Loop Trajectory Planning in Autonomous Driving

BridgeDrive: 基于扩散桥的闭环轨迹规划扩散策略

Shu Liu, Wenlin Chen, Weihao Li, Zheng Wang, Lijin Yang, Jianing Huang, Yipin Zhang, Zhongzhan Huang, Ze Cheng, Hao Yang

机构 * Bosch (China) Investment Ltd(博世(中国)投资有限公司)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 BridgeDrive提出一种基于锚点的扩散桥策略,通过实时闭环轨迹规划提升自动驾驶的安全性和效率。

Comments Accepted for publication at ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18534 2026-03-06 cs.CV cs.LG 79%

Elucidating the Design Space of Arbitrary-Noise-Based Diffusion Models

阐明基于任意噪声的扩散模型的设计空间

Xingyu Qiu, Mengying Yang, Xinghua Ma, Dong Liang, Fanding Li, Gongning Luo, Wei Wang, Kuanquan Wang, Shuo Li

机构 * Harbin Institute of Technology(哈尔滨工业大学) Case Western Reserve University(凯斯西储大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 EDA通过扩展噪声模式灵活性并保持模块化,提升图像恢复任务的泛化能力。

Comments 16 pages, 4 figures, accepted by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.04307 2026-03-05 cs.CV 79%

Dual Diffusion Models for Multi-modal Guided 3D Avatar Generation

多模态引导的3D人像生成双扩散模型

Hong Li, Yutang Feng, Minqi Meng, Yichen Yang, Xuhui Liu, Baochang Zhang

机构 * Beihang University(北航大学) KAUST(卡塔尔科技大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出PromptAvatar,通过双扩散模型实现多模态引导的3D人像生成,提升生成质量与效率。

Comments 18 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.03792 2026-03-05 cs.CV cs.LG 79%

TAP: A Token-Adaptive Predictor Framework for Training-Free Diffusion Acceleration

TAP:一种用于无训练扩散加速的令牌自适应预测框架

Haowei Zhu, Tingxuan Huang, Xing Wang, Tianyu Zhao, Jiexi Wang, Weifeng Chen, Xurui Peng, Fangmin Chen, Junhai Yong, Bin Wang

机构 * Tsinghua University(清华大学) ByteDance(字节跳动)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 TAP通过为每个令牌自适应选择预测器,实现无训练的扩散模型加速,提升生成效率并减少感知质量损失。

Comments Accepted by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.03602 2026-03-05 cs.CV 79%

DM-CFO: A Diffusion Model for Compositional 3D Tooth Generation with Collision-Free Optimization

DM-CFO: 一种用于无碰撞优化的组合3D牙齿生成扩散模型

Yan Tian, Pengcheng Xue, Weiping Ding, Mahmoud Hassaballah, Karen Egiazarian, Aura Conci, Abdulkadir Sengur, Leszek Rutkowski

机构 * School of Computer Science and Technology, Zhejiang Gongshang University(浙江工商大学计算机科学与技术学院) AGH University of Krakow(克拉科夫AGH大学) Polish Ministry of Science and Higher Education(波兰教育部) Tongxiang Institute of General Artificial Intelligence(同祥通用人工智能研究院) State Key Laboratory of Advanced Medical Materials and Devices(先进医学材料与器件国家重点实验室) School of Artificial Intelligence and Computer Science, Nantong University(南通大学人工智能与计算机科学学院) Department of Computer Science, College of Computer Engineering and Sciences, Prince Sattam Bin Abdulaziz University(普莱斯·本·阿卜杜勒阿齐兹大学计算机科学系) Department of Computer Science, Qena University(基纳大学计算机科学系) Department of Computing Sciences, Tampere University(塔尔库大学计算科学系) Department of Computer Science, Universidade Federal Fluminense(里约热内卢联邦大学计算机科学系) Shining3D Tech Co., Ltd.(Shining3D科技有限公司)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 DM-CFO通过扩散模型和无碰撞优化生成组合3D牙齿,提升生成牙齿的多视图一致性和真实性。

Comments Received by IEEE Transactions on Visualization and Computer Graphics

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21574 2026-03-05 cs.CV cs.LG 79%

Do We Need All the Synthetic Data? Targeted Image Augmentation via Diffusion Models

我们是否需要所有合成数据?通过扩散模型的目标图像增强

Dang Nguyen, Jiping Li, Jinghao Zheng, Baharan Mirzasoleiman

机构 * University of California Los Angeles (UCLA)(加州大学洛杉矶分校) Ecole Polytechnique Federale de Lausanne (EPFL)(瑞士联邦理工学院洛桑分校)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 TADA通过选择性增强训练初期未被学习的示例,利用扩散模型提高图像分类的泛化能力,实验表明仅增强30-40%的数据即可提升性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.19289 2026-03-05 cs.CV 79%

FINE: Factorizing Knowledge for Initialization of Variable-sized Diffusion Models

FINE: 用于变量大小扩散模型初始化的知识因子化

Yucheng Xie, Fu Feng, Ruixiao Shi, Jianlu Shen, Jing Wang, Yong Rui, Xin Geng

机构 * School of Computer Science and Engineering, Southeast University, Nanjing, China(计算机科学与工程学院,东南大学,南京,中国) Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications (Southeast University), Ministry of Education, China(新一代人工智能技术及其交叉应用关键实验室(东南大学),教育部,中国)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 FINE通过分解知识为learngenes,实现变量大小扩散模型的高效初始化,无需重复预训练,提升资源受限部署下的性能与适应性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07093 2026-03-05 cs.CV cs.AI 79%

3D Wavelet-Based Structural Priors for Controlled Diffusion in Whole-Body Low-Dose PET Denoising

基于3D小波的结构先验用于控制扩散的全身体低剂量PET去噪

Peiyuan Jing, Yue Yang, Chun-Wun Cheng, Zhenxuan Zhang, Liutao Yang, Thiago V. Lima, Klaus Strobel, Antoine Leimgruber, Angelica Aviles-Rivero, Guang Yang, Javier A. Montoya-Zegarra

机构 * School of Engineering, Zurich University of Applied Sciences, CH Bioengineering Department Imperial-X, Imperial College London, UK DAMTP, University of Cambridge, UK Lucerne University Teaching Research Hospital, CH Lung Institute, Imperial College London, UK Cardiovascular Research Centre, Royal Brompton Hospital, UK School of Biomedical Engineering \& Imaging Sciences, King's College London, UK Yau Mathematical Sciences Center, Tsinghua University, CN

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 WCC-Net通过引入3D小波结构先验,提升低剂量PET去噪效果,实现更稳定的解剖结构与噪声分离。

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.03265 2026-03-04 cs.CV 79%

DuoMo: Dual Motion Diffusion for World-Space Human Reconstruction

DuoMo:用于世界空间人体重建的双动差分

Yufu Wang, Evonne Ng, Soyong Shin, Rawal Khirodkar, Yuan Dong, Zhaoen Su, Jinhyung Park, Kris Kitani, Alexander Richard, Fabian Prada, Michael Zollhofer

机构 * Meta Reality Labs University of Pennsylvania(宾夕法尼亚大学) Carnegie Mellon University(卡内基梅隆大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 DuoMo通过双扩散模型实现世界空间人体运动重建,有效处理噪声和不完整输入,提升重建精度。

Comments CVPR 2026. Project page: https://yufu-wang.github.io/duomo/

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02882 2026-03-04 cs.CV 79%

SIGMark: Scalable In-Generation Watermark with Blind Extraction for Video Diffusion

SIGMark: 用于视频扩散模型的可扩展生成内水印与盲提取

Xinjie Zhu, Zijing Zhao, Hui Jin, Qingxiao Guo, Yilong Ma, Yunhao Wang, Xiaobing Guo, Weifeng Zhang

机构 * Lenovo Research(联想研究院)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 SIGMark是一种用于视频扩散模型的可扩展生成内水印框架,通过盲提取技术实现无损水印并提升鲁棒性。

Comments Accepted to ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02710 2026-03-04 cs.CV 79%

MiM-DiT: MoE in MoE with Diffusion Transformers for All-in-One Image Restoration

MiM-DiT:基于扩散变换器的MoE在MoE中实现全功能图像修复

Lingshun Kong, Jiawei Zhang, Zhengpeng Duan, Xiaohe Wu, Yueqi Yang, Xiaotao Wang, Dongqing Zou, Lei Lei, Jinshan Pan

机构 * Nanjing University of Science and Technology(南京理工大学) Nankai University(南开大学) Harbin Institute of Technology(哈尔滨工业大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 MiM-DiT通过双层MoE架构与预训练扩散模型,实现对多种图像退化类型的统一修复,提升复杂真实场景下的修复效果。

Comments Project website: https://github.com/kkkls/MIM-DiT

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02691 2026-03-04 cs.CV 79%

ReCo-Diff: Residual-Conditioned Deterministic Sampling for Cold Diffusion in Sparse-View CT

ReCo-Diff:基于残差的确定性采样用于稀疏视角CT中的冷扩散

Yong Eun Choi, Hyoung Suk Park, Kiwan Jeon, Hyun-Cheol Park, Sung Ho Kang

机构 * National Institute for Mathematical Sciences, Daejeon, 34047, Republic of Korea(韩国全南数学研究所) Department of Computer Engineering, Korea National University of Transportation, Chungju, 27469, Republic of Korea(韩国交通国立大学计算机工程系)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 ReCo-Diff通过残差条件化的确定性采样提升稀疏视角CT重建的精度与稳定性。

Comments 10 pages, 4 figures. Submitted to MICCAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17561 2026-03-04 cs.CV cs.AI 79%

Model Already Knows the Best Noise: Bayesian Active Noise Selection via Attention in Video Diffusion Model

模型已知最佳噪声:通过注意力的贝叶斯主动噪声选择在视频扩散模型中

Kwanyoung Kim, Sanghyun Kim

机构 * Department of AI Convergence, Gwangju Institute of Science and Technology (GIST)(人工智能融合系,光州科学技术院) Samsung Research(三星研究所)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 通过注意力机制的贝叶斯主动噪声选择方法,提升视频扩散模型的生成质量与时间一致性。

Comments Cam ready version of ICLR 26

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02035 2026-03-03 cs.RO cs.CV 79%

LAD-Drive: Bridging Language and Trajectory with Action-Aware Diffusion Transformers

LAD-Drive: 通过动作感知扩散变换器弥合语言与轨迹

Fabian Schmidt, Karol Fedurko, Markus Enzweiler, Abhinav Valada

机构 * Institute for Intelligent Systems, Esslingen University of Applied Sciences(智能系统研究所,埃斯林根应用科学大学) Department of Computer Science, University of Freiburg(计算机科学系,弗赖堡大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 LAD-Drive通过动作感知扩散变换器,将语言模型的意图与连续轨迹生成结合,提升自动驾驶中的多模态规划能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02012 2026-03-03 cs.CV cs.AI 79%

MAP-Diff: Multi-Anchor Guided Diffusion for Progressive 3D Whole-Body Low-Dose PET Denoising

MAP-Diff: 多锚点引导的扩散模型用于渐进式三维全身低剂量PET去噪

Peiyuan Jing, Chun-Wun Cheng, Liutao Yang, Zhenxuan Zhang, Thiago V. Lima, Klaus Strobel, Antoine Leimgruber, Angelica Aviles-Rivero, Guang Yang, Javier A. Montoya-Zegarra

机构 * School of Engineering, Zurich University of Applied Sciences, CH Bioengineering Department Imperial-X, Imperial College London, UK DAMTP, University of Cambridge, UK Lucerne University Teaching Research Hospital, CH Lung Institute, Imperial College London, UK Cardiovascular Research Centre, Royal Brompton Hospital, UK School of Biomedical Engineering \& Imaging Sciences, King's College London, UK Yau Mathematical Sciences Center, Tsinghua University, CN

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 MAP-Diff通过多锚点引导的扩散模型实现低剂量PET图像的渐进式去噪,提升PSNR和SSIM,降低NMAE,优于多种基线方法。

Comments 8 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01926 2026-03-03 cs.IR cs.CV 79%

MealRec: Multi-granularity Sequential Modeling via Hierarchical Diffusion Models for Micro-Video Recommendation

MealRec: 基于分层扩散模型的多粒度序列建模用于微视频推荐

Xinxin Dong, Haokai Ma, Yuze Zheng, Yongfu Zha, Yonghui Yang, Xiaodong Wang

机构 * National University of Defense Technology(国防科技大学) National University of Singapore(新加坡国立大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 MealRec通过分层扩散模型实现多粒度序列建模,解决微视频推荐中的偏好无关表示和模态冲突问题,提升推荐效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01913 2026-03-03 cs.CV 79%

Zero-shot Low-Field MRI Enhancement via Diffusion-Based Adaptive Contrast Transport

零场MRI增强通过扩散基于自适应对比传输

Muyu Liu, Chenhe Du, Xuanyu Tian, Qing Wu, Xiao Wang, Haonan Zhang, Hongjiang Wei, Yuyao Zhang

机构 * School of Information Science and Technology, ShanghaiTech University, Shanghai, China(信息科学与技术学院,上海科技大学) School of Biomedical Engineering, Shanghai Jiao Tong University, Shanghai, China(生物医学工程学院,上海交通大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出DACT框架,通过扩散模型和自适应对比传输技术,在无配对监督的情况下实现低场MRI到高场MRI的高质量图像重建,提升组织对比和结构细节。

Comments 11 pages, 4 figures, conference paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01688 2026-03-03 cs.CV 79%

CoopDiff: A Diffusion-Guided Approach for Cooperation under Corruptions

CoopDiff: 一种基于扩散的协作方法以应对腐蚀

Gong Chen, Chaokun Zhang, Pengcheng Lv

机构 * School of Computer Science and Technology, Tianjin University(天津大学计算机科学与技术学院) School of Cybersecurity, Tianjin University(天津大学网络安全学院) School of Future Technology, Tianjin University(天津大学未来技术学院)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 CoopDiff是一种基于扩散的协作感知方法,通过教师-学生范式和去噪机制有效应对腐蚀,提升了鲁棒性和泛化能力。

Comments Accepted by CVPR26

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01686 2026-03-03 cs.CV 79%

DiffusionXRay: A Diffusion and GAN-Based Approach for Enhancing Digitally Reconstructed Chest Radiographs

DiffusionXRay: 一种基于扩散和GAN的方法用于增强数字重建的胸部X光影像

Aryan Goyal, Ashish Mittal, Pranav Rao, Manoj Tadepalli, Preetham Putha

机构 * Indian Institute of Technology Bombay, India(印度理工学院班加罗尔学院) Qure.ai, India

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 DiffusionXRay通过结合扩散模型和GAN解决胸部X光影像质量退化问题,提升诊断价值。

Comments Published at MICCAI 2025

Journal ref Data Engineering in Medical Imaging: Third MICCAI Workshop, DEMI 2025, Held in Conjunction with MICCAI 2025, Daejeon, South Korea, September 27, 2025, Proceedings

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01659 2026-03-03 cs.CV 79%

A Diffusion-Driven Fine-Grained Nodule Synthesis Framework for Enhanced Lung Nodule Detection from Chest Radiographs

一种基于扩散的细粒度结节合成框架,用于增强胸部X光片上结节检测

Aryan Goyal, Shreshtha Singh, Ashish Mittal, Manoj Tadepalli, Piyush Kumar, Preetham Putha

机构 * Indian Institute of Technology, Bombay(印度理工学院,孟买)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出了一种基于扩散的细粒度结节合成框架,通过低秩适应适配器实现对结节特征的精确控制,提升了胸部X光片上结节检测的性能。

Comments Accepted at MIDL 2026 (Poster). Published on OpenReview on February 14, 2026. Proceedings version pending. OpenReview: https://openreview.net/forum?id=7DL7cu8Ui8

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01253 2026-03-03 cs.CV 79%

Cross-Modal Guidance for Fast Diffusion-Based Computed Tomography

跨模态引导用于快速扩散基于计算机断层扫描

Timofey Efimov, Singanallur Venkatakrishnan, Maliha Hossain, Haley Duba-Sullivan, Amirkoushyar Ziabari

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出一种无需重新训练扩散模型的跨模态引导方法,用于提升稀疏视图中子CT的重建质量。

Comments Accepted at the IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01103 2026-03-03 cs.CV 79%

Data-Efficient Brushstroke Generation with Diffusion Models for Oil Painting

基于扩散模型的数据高效笔触生成用于油画

Dantong Qin, Alessandro Bozzon, Xian Yang, Xun Zhang, Yike Guo, Pan Wang

机构 * Delft University of Technology(代尔夫特理工大学) The University of Manchester(曼彻斯特大学) The Hong Kong University of Science(香港科学大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出StrokeDiff模型,通过平滑正则化技术实现基于扩散模型的数据高效笔触生成,用于油画创作。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26818 2026-03-03 cs.SD cs.AI cs.MM eess.AS 79%

GACA-DiT: Diffusion-based Dance-to-Music Generation with Genre-Adaptive Rhythm and Context-Aware Alignment

GACA-DiT:基于扩散的舞蹈到音乐生成,具有风格自适应节奏和上下文感知对齐

Jinting Wang, Chenxing Li, Li Liu

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) Tencent AI Lab(腾讯AI实验室)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.MM

AI总结 GACA-DiT通过风格自适应节奏提取和上下文感知时间对齐模块,实现了舞蹈到音乐生成的节奏一致性和时间对齐,提升了生成音乐与舞蹈动作的同步精度。

Comments 5 pages, 4 figures, submitted to Interspeech2026

Journal ref sumbitted to Interspeech2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15301 2026-03-03 cs.CV cs.AI 79%

Latent Diffusion Model without Variational Autoencoder

无变分自编码器的潜在扩散模型

Minglei Shi, Haolin Wang, Wenzhao Zheng, Ziyang Yuan, Xiaoshi Wu, Xintao Wang, Pengfei Wan, Jie Zhou, Jiwen Lu

机构 * Department of Automation, Tsinghua University(自动化系,清华大学) Kling Team, Kuaishou Technology(快手科技 Kling 团队)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 SVG提出了一种无需变分自编码器的潜在扩散模型,通过自监督表示提升视觉生成的效率和质量。

Comments Accepted by ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏