Prompt-SID: Learning Structural Representation Prompt via Latent Diffusion for Single-Image Denoising
Prompt-SID: 通过潜在扩散学习结构表示提示进行单图像去噪
专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV
AI总结 Prompt-SID通过潜在扩散学习结构表示提示,提升单图像去噪效果,有效保留结构细节。
视觉与机器人
图像生成、文生图、图像编辑、扩散模型和可控生成。
Prompt-SID: 通过潜在扩散学习结构表示提示进行单图像去噪
专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV
AI总结 Prompt-SID通过潜在扩散学习结构表示提示,提升单图像去噪效果,有效保留结构细节。
RDM:用于人类动作生成的递归扩散模型
机构 * Department of Computer Science, University College London(计算机科学系,伦敦大学学院)
专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV
AI总结 RDM通过递归扩散模型生成长序列人类动作,利用归一化流建模循环连接,提升效率并保持概率性质。
Comments v2: Major revision with extensive text polishing and structural updates. Added new experiments on the rollout effect, specifically analyzing the trade-offs between compute time and sequence length. Includes several new visualizations (Figures 6, 9, 10) and an expanded discussion in Section 4
HARP: 仅使用假体数据进行体内扩散磁共振成像的协调
机构 * Department of Psychiatry(精神医学系) ; Brigham and Women's Hospital(布里奇沃特医院) ; Harvard Medical School(哈佛医学院) ; College of Engineering(工程学院) ; Northeastern University(东北大学) ; National Institute of Standards and Technology(国家标准技术研究院) ; University of Utah School of Medicine(犹他大学医学院) ; George E. Wahlen Veterans Affairs Medical Center(乔治·E·瓦伦的退伍军人事务医疗中心) ; University of Pittsburgh(匹兹堡大学) ; Department of Radiology(放射医学系) ; Harvard-MIT Health Sciences and Technology(哈佛-麻省理工健康科学与技术)
专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV
AI总结 HARP通过仅使用假体数据训练深度学习模型,实现了无需多站点活体数据的扩散磁共振成像协调,有效降低了扫描仪间变异性。
剪枝之下:揭示基于剪枝的去学习中概念复兴的风险
机构 * University of Georgia(佐治亚大学) ; Carnegie Mellon University(卡内基梅隆大学) ; Northeastern University(东北大学) ; Stevens Institute of Technology(史蒂文斯理工学院) ; University of Arizona(亚利桑那大学)
专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV
AI总结 本文揭示基于剪枝的去学习中概念复兴的风险,提出攻击框架可无数据恢复被擦除概念,并探讨安全剪枝机制。
Comments Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2026
S2DiT:用于移动流媒体视频生成的 Sandwich Diffusion Transformer
机构 * Snap Inc. ; Northeastern University(东北大学)
专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV
AI总结 S2DiT 通过高效注意力机制和 2-in-1 深度学习框架,在移动设备上实现高质量、高速度的流式视频生成。
LD-RPS:通过潜势扩散递归后验采样实现零样本统一图像修复
机构 * Tsinghua University(清华大学) ; AMAP, Alibaba Group(阿里云(AMAP))
专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV
AI总结 LD-RPS通过潜势扩散递归后验采样实现零样本统一图像修复,结合多模态模型和轻量模块提升修复效果。
从二维对齐到三维合理性:统一异构二维先验和无穿透扩散以实现抗遮挡的双手重建
机构 * AgiBot ; Mohamed bin Zayed University of Artificial Intelligence ; The University of Sydney ; La Trobe University
专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV
AI总结 本文提出统一异构二维先验和无穿透扩散模型,以实现抗遮挡的双手重建,提升交互对齐和穿透抑制性能。
Comments Accepted by CVPR 2026 Main, Project: https://gaogehan.github.io/A2P/
ExpGest:利用扩散模型和混合音频-文本引导的表达性说话生成
机构 * Northwest A&F University(西北农林科技大学) ; University of Technology Sydney(悉尼大学) ; Tencent AILab(腾讯AI实验室) ; City University of Hong Kong(香港城市大学)
专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV
AI总结 ExpGest通过结合文本和音频信息,利用扩散模型生成更自然、可控的全身表达性手势。
Comments Accepted by ICME 2024
探索扩散模型在少样本微调中的腐蚀阶段并利用贝叶斯神经网络缓解
机构 * Shanghai Jiao Tong University(上海交通大学) ; Queen’s University Belfast(女王学院 Belfast) ; Tsinghua University(清华大学) ; Stevens Institute of Technology(史蒂文斯理工学院)
专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV
AI总结 本研究通过贝叶斯神经网络缓解扩散模型在少样本微调中的腐蚀阶段问题,提升生成图像质量。
Comments Accepted by KDD' 26
DiffInf: 基于影响的扩散法用于面部属性学习中的监督对齐
机构 * Departement of Electrical and Computer Engineering(电气与计算机工程系) ; Departement of Biomedical Engineering(生物医学工程系)
专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV
AI总结 DiffInf通过自影响引导的扩散方法,修复面部属性学习中的标注不一致,提升分类泛化能力。
TempoSyncDiff: 压缩时间一致扩散用于低延迟音频驱动说话头生成
机构 * Computer and Informatics Group, Variable Energy Cyclotron Centre(计算机与信息组,变能循环中心)
专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV
AI总结 TempoSyncDiff通过压缩扩散模型实现低延迟音频驱动说话头生成,结合教师-学生压缩和时间正则化以提高生成稳定性。
去噪作为路径规划:通过DPCache实现扩散模型的训练免费加速
机构 * Alibaba Group(阿里巴巴集团)
专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV
AI总结 DPCache通过全局路径规划框架实现扩散模型的高效加速,减少计算开销并提升生成质量。
Comments Accepted by CVPR 2026
SRA 2: 变分自编码器自表示对齐用于高效的扩散训练
机构 * Zhejiang University of Technology(浙江工业大学) ; SGIT AI Lab, State Grid Corporation of China(国网SGIT人工智能实验室) ; Zhejiang University(浙江大学) ; Baidu(百度)
专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV
AI总结 SRA 2通过利用预训练变分自编码器特征,提供一种轻量级的内在引导框架,提升扩散模型的生成质量和训练效率,仅增加4%的计算开销。
SyncMV4D: 同步多视角联合生成外观与运动的 hand-object 交互合成
专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV
AI总结 SyncMV4D通过联合生成多视角手-物体交互视频和4D运动,提升了视觉真实性和运动合理性。
Comments The structure and logic of writing will undergo a complete revision
ExDD:通过扩散合成实现表面缺陷检测的显式双分布学习
机构 * Dept. of Engineering for Innovation Medicine, University of Verona(创新医学工程系,威尼斯大学) ; Qualyco S.r.l.(Qualyco公司)
专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV
AI总结 ExDD通过显式建模双分布和潜在扩散模型生成合成缺陷,提升工业表面缺陷检测的准确率和鲁棒性。
Comments Accepted to ICIAP 2025
DVD-Quant: 数据无依赖的视频扩散变换器量化
机构 * Shanghai Jiao Tong University(上海交通大学) ; Zhejiang University(浙江大学) ; ETH Zürich(苏黎世联邦理工学院)
专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV
AI总结 DVD-Quant通过无数据量化框架实现视频DiT模型的高效量化,提升运行速度并保持视觉质量。
Comments Code and models will be available at https://github.com/lhxcs/DVD-Quant
Ditto:用于可控实时说话头合成的运动空间扩散
机构 * Ant Group(蚂蚁集团)
专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV
AI总结 Ditto通过运动空间扩散模型实现可控实时说话头合成,优化架构与训练策略以提升生成质量与实时性能。
Comments Project Page: https://digital-avatar.github.io/ai/Ditto/
Journal ref ACM MM 2025
面向频率的误差有界缓存用于加速扩散变换器
机构 * iFLYTEK
专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV
AI总结 SpectralCache通过统一缓存框架提升扩散变换器推理速度,实现2.46倍加速,质量与TeaCache相当。
Spinverse: 可微物理用于从扩散MRI中感知渗透性的微结构重建
机构 * University of California Santa Cruz(加州大学圣克鲁兹分校) ; University of California San Diego(加州大学圣地亚哥分校) ; Inria-Saclay(Inria-萨克莱实验室) ; Kings College London(伦敦国王学院)
专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV
AI总结 Spinverse通过可微物理方法从扩散MRI中感知渗透性,重建微结构并优化渗透性参数以提高边界准确性和结构有效性。
Comments 10 Pages, 5 Figures, 2 Tables
BridgeDrive: 基于扩散桥的闭环轨迹规划扩散策略
机构 * Bosch (China) Investment Ltd(博世(中国)投资有限公司)
专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV
AI总结 BridgeDrive提出一种基于锚点的扩散桥策略,通过实时闭环轨迹规划提升自动驾驶的安全性和效率。
Comments Accepted for publication at ICLR 2026
阐明基于任意噪声的扩散模型的设计空间
机构 * Harbin Institute of Technology(哈尔滨工业大学) ; Case Western Reserve University(凯斯西储大学)
专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV
AI总结 EDA通过扩展噪声模式灵活性并保持模块化,提升图像恢复任务的泛化能力。
Comments 16 pages, 4 figures, accepted by CVPR 2026
多模态引导的3D人像生成双扩散模型
机构 * Beihang University(北航大学) ; KAUST(卡塔尔科技大学)
专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV
AI总结 本文提出PromptAvatar,通过双扩散模型实现多模态引导的3D人像生成,提升生成质量与效率。
Comments 18 pages, 10 figures
TAP:一种用于无训练扩散加速的令牌自适应预测框架
机构 * Tsinghua University(清华大学) ; ByteDance(字节跳动)
专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV
AI总结 TAP通过为每个令牌自适应选择预测器,实现无训练的扩散模型加速,提升生成效率并减少感知质量损失。
Comments Accepted by CVPR 2026
DM-CFO: 一种用于无碰撞优化的组合3D牙齿生成扩散模型
机构 * School of Computer Science and Technology, Zhejiang Gongshang University(浙江工商大学计算机科学与技术学院) ; AGH University of Krakow(克拉科夫AGH大学) ; Polish Ministry of Science and Higher Education(波兰教育部) ; Tongxiang Institute of General Artificial Intelligence(同祥通用人工智能研究院) ; State Key Laboratory of Advanced Medical Materials and Devices(先进医学材料与器件国家重点实验室) ; School of Artificial Intelligence and Computer Science, Nantong University(南通大学人工智能与计算机科学学院) ; Department of Computer Science, College of Computer Engineering and Sciences, Prince Sattam Bin Abdulaziz University(普莱斯·本·阿卜杜勒阿齐兹大学计算机科学系) ; Department of Computer Science, Qena University(基纳大学计算机科学系) ; Department of Computing Sciences, Tampere University(塔尔库大学计算科学系) ; Department of Computer Science, Universidade Federal Fluminense(里约热内卢联邦大学计算机科学系) ; Shining3D Tech Co., Ltd.(Shining3D科技有限公司)
专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV
AI总结 DM-CFO通过扩散模型和无碰撞优化生成组合3D牙齿,提升生成牙齿的多视图一致性和真实性。
Comments Received by IEEE Transactions on Visualization and Computer Graphics
我们是否需要所有合成数据?通过扩散模型的目标图像增强
机构 * University of California Los Angeles (UCLA)(加州大学洛杉矶分校) ; Ecole Polytechnique Federale de Lausanne (EPFL)(瑞士联邦理工学院洛桑分校)
专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV
AI总结 TADA通过选择性增强训练初期未被学习的示例,利用扩散模型提高图像分类的泛化能力,实验表明仅增强30-40%的数据即可提升性能。
FINE: 用于变量大小扩散模型初始化的知识因子化
机构 * School of Computer Science and Engineering, Southeast University, Nanjing, China(计算机科学与工程学院,东南大学,南京,中国) ; Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications (Southeast University), Ministry of Education, China(新一代人工智能技术及其交叉应用关键实验室(东南大学),教育部,中国)
专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV
AI总结 FINE通过分解知识为learngenes,实现变量大小扩散模型的高效初始化,无需重复预训练,提升资源受限部署下的性能与适应性。
基于3D小波的结构先验用于控制扩散的全身体低剂量PET去噪
机构 * School of Engineering, Zurich University of Applied Sciences, CH ; Bioengineering Department ; Imperial-X, Imperial College London, UK ; DAMTP, University of Cambridge, UK ; Lucerne University Teaching ; Research Hospital, CH ; Lung Institute, Imperial College London, UK ; Cardiovascular Research Centre, Royal Brompton Hospital, UK ; School of Biomedical Engineering \& Imaging Sciences, King's College London, UK ; Yau Mathematical Sciences Center, Tsinghua University, CN
专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV
AI总结 WCC-Net通过引入3D小波结构先验,提升低剂量PET去噪效果,实现更稳定的解剖结构与噪声分离。
Comments 10 pages
DuoMo:用于世界空间人体重建的双动差分
机构 * Meta Reality Labs ; University of Pennsylvania(宾夕法尼亚大学) ; Carnegie Mellon University(卡内基梅隆大学)
专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV
AI总结 DuoMo通过双扩散模型实现世界空间人体运动重建,有效处理噪声和不完整输入,提升重建精度。
Comments CVPR 2026. Project page: https://yufu-wang.github.io/duomo/
SIGMark: 用于视频扩散模型的可扩展生成内水印与盲提取
机构 * Lenovo Research(联想研究院)
专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV
AI总结 SIGMark是一种用于视频扩散模型的可扩展生成内水印框架,通过盲提取技术实现无损水印并提升鲁棒性。
Comments Accepted to ICLR 2026
MiM-DiT:基于扩散变换器的MoE在MoE中实现全功能图像修复
机构 * Nanjing University of Science and Technology(南京理工大学) ; Nankai University(南开大学) ; Harbin Institute of Technology(哈尔滨工业大学)
专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV
AI总结 MiM-DiT通过双层MoE架构与预训练扩散模型,实现对多种图像退化类型的统一修复,提升复杂真实场景下的修复效果。
Comments Project website: https://github.com/kkkls/MIM-DiT