arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

共收录 86872 信号源:cs.CV, cs.GR, cs.MM

1. 扩散模型 70277 篇

2604.21422 2026-04-24 cs.CV 79%

Pre-process for segmentation task with nonlinear diffusion filters

分割任务前的非线性扩散滤波预处理

Javier Sanguino, Carlos Platero, Olga Velasco

机构 * Health Science Technology Group, Technical University of Madrid(健康科学技术组,马德里技术大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出非线性扩散滤波用于分割前的预处理,设计新型扩散函数以实现清晰轮廓和无模糊边缘,证明滤波器满足尺度空间要求,通过半隐式方案高效生成分段常数图像。

Comments Manuscript from 2017, previously unpublished, 37 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.21279 2026-04-24 cs.CV 79%

LatRef-Diff: Latent and Reference-Guided Diffusion for Facial Attribute Editing and Style Manipulation

LatRef-Diff: 基于潜在和参考引导的扩散模型用于面部属性编辑和风格操控

Wenmin Huang, Weiqi Luo, Xiaochun Cao, Jiwu Huang

机构 * GuangDong Province Key Lab of Information Security Technology and School of Computer Science and Engineering(广东信息安全技术重点实验室和计算机科学与工程学院) School of Cyber Science and Technology, Shenzhen Campus(深圳校区网络科学与技术学院) Guangdong Laboratory of Machine Perception and Intelligent Computing, Faculty of Engineering(广东机器感知与智能计算实验室,工程学院)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出LatRef-Diff,通过潜在和参考引导方法改进面部属性编辑和风格操控,解决传统扩散模型在属性控制和风格操控中的不足,采用风格代码和模块化设计提升精度和图像质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.21221 2026-04-24 cs.CV cs.LG 79%

Sparse Forcing: Native Trainable Sparse Attention for Real-time Autoregressive Diffusion Video Generation

稀疏强制:用于实时自回归扩散视频生成的原生可训练稀疏注意力

Boxun Xu, Yuming Du, Zichang Liu, Siyu Yang, Ziyang Jiang, Siqi Yan, Rajasi Saha, Albert Pumarola, Wenchen Wang, Peng Li

机构 * Meta Superintelligence Labs(Meta超智能实验室) University of California, Santa Barbara(加州大学圣芭芭拉分校)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出稀疏强制方法,通过在训练和推理中提升长时程生成质量并降低解码延迟,采用可训练的原生稀疏机制和高效的GPU内核PBSA来加速稀疏注意力和内存更新。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.17656 2026-04-24 cs.SD cs.AI cs.CL cs.CV cs.LG 79%

Video-Robin: Autoregressive Diffusion Planning for Intent-Grounded Video-to-Music Generation

Video-Robin:基于意图的视频到音乐生成的自回归扩散规划

Vaibhavi Lokegaonkar, Aryan Vijay Bhosale, Vishnu Raj, Gouthaman KV, Ramani Duraiswami, Lie Lu, Sreyan Ghosh, Dinesh Manocha

机构 * University of Maryland College Park(马里兰大学College Park分校) Dolby Laboratories(杜比实验室) NVIDIA(英伟达)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 Video-Robin通过结合自回归规划与扩散合成,实现高质量的语义对齐音乐生成,相比传统方法在速度和质量上均有显著提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18457 2026-04-24 cs.CV cs.LG 79%

VFM-VAE: Vision Foundation Models Can Be Good Tokenizers for Latent Diffusion Models

VFM-VAE:视觉基础模型可以作为潜在扩散模型的良好分词器

Tianci Bi, Xiaoyi Zhang, Yan Lu, Nanning Zheng

机构 * State Key Laboratory of Human-Machine Hybrid Augmented Intelligence, Institute of Artificial Intelligence and Robotics, Xi’an Jiaotong University(人机混合增强智能国家重点实验室,人工智能与机器人研究院,西安交通大学) Microsoft Research Asia(微软亚洲研究院)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出VFM-VAE,利用冻结的视觉基础模型作为潜在扩散模型的分词器,通过设计新解码器提升图像重建能力,实现分词器与扩散模型的协同优化,提升训练效率与性能。

Comments Accepted at CVPR 2026. Code and models available at: https://github.com/tianciB/VFM-VAE

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13576 2026-04-24 eess.IV cs.CV 79%

Cross-Distribution Diffusion Priors-Driven Iterative Reconstruction for Sparse-View CT

跨分布扩散先验驱动的稀疏视图CT迭代重建

Haodong Li, Shuo Han, Haiyang Mao, Yu Shi, Changsheng Fang, Jianjia Zhang, Weiwen Wu, Hengyong Yu

机构 * Department of Electrical and Computer Engineering, University of Massachusetts Lowell(马萨诸塞大学洛厄尔分校电子与计算机工程系) School of Biomedical Engineering, Sun Yat-Sen University(中山大学生物医学工程学院)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出CDPIR框架,通过整合跨分布扩散先验与模型驱动迭代重建方法,解决稀疏视图CT在分布外场景下的重建问题,提升图像质量与泛化能力。

Comments 17 pages, 15 figures, accepted by IEEE Transactions on Medical Imaging

Journal ref IEEE Transactions on Medical Imaging, 2026 (early access)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.20470 2026-04-23 cs.CV 79%

DynamicRad: Content-Adaptive Sparse Attention for Long Video Diffusion

动态Rad:面向长视频扩散的内容自适应稀疏注意力

Yongji Long, Shijun Liang, Jintao Li, Yun Li

机构 * University of Electronic Science and Technology of China(电子科学与技术大学) Shenzhen Institute for Advanced Study, UESTC(深圳先进研究院) Computational Mathematics, Science, & Engineering at Michigan State University (MSU)(密歇根州立大学计算数学、科学与工程)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出DynamicRad,通过引入双模式策略和语义运动路由器,实现长视频扩散中的高效稀疏注意力机制,提升推理速度并保持高质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.20715 2026-04-23 cs.CV 79%

GeoRelight: Learning Joint Geometrical Relighting and Reconstruction with Flexible Multi-Modal Diffusion Transformers

GeoRelight: 基于灵活多模态扩散变换器的联合几何重照明与重建

Yuxuan Xue, Ruofan Liang, Egor Zakharov, Timur Bagautdinov, Chen Cao, Giljoo Nam, Shunsuke Saito, Gerard Pons-Moll, Javier Romero

机构 * Codec Avatars Lab, Meta(Meta编码器动画实验室) University of Tübingen(图宾根大学) Max Planck Institute for Informatics, Saarland Informatics Campus(马克斯·普朗克信息学院,萨尔兰信息校园)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出GeoRelight,通过联合解决几何重建与重照明问题,利用iNOD和混合数据训练方法提升性能,优于传统顺序模型和忽略几何的系统。

Comments CVPR 2026 Highlight; Project page: https://yuxuan-xue.com

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.20594 2026-04-23 cs.CV 79%

Physics-Informed Conditional Diffusion for Motion-Robust Retinal Temporal Laser Speckle Contrast Imaging

具有物理信息的条件扩散用于运动鲁棒的视网膜时间激光散斑对比成像

Qian Chen, Yuehao Chen, Qiang Wang, Lei Zhu, Yanye Lu, Qiushi Ren

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出RetinaDiff框架,通过相位相关注册和条件扩散模型实现运动鲁棒的视网膜时间激光散斑对比成像,提升有限帧下的结构连续性和统计稳定性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.19747 2026-04-22 cs.CV 79%

AnyRecon: Arbitrary-View 3D Reconstruction with Video Diffusion Model

AnyRecon:基于视频扩散模型的任意视角3D重建

Yutian Chen, Shi Guo, Renbiao Jin, Tianshuo Yang, Xin Cai, Yawen Luo, Mingxin Yang, Mulin Yu, Linning Xu, Tianfan Xue

机构 * AnyRecon: Arbitrary-View 3D Reconstruction with Video Diffusion Model(AnyRecon:任意视角3D重建与视频扩散模型)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 AnyRecon通过视频扩散模型实现任意视角3D重建,保留显式几何控制并支持灵活条件设置,通过全局场景记忆和几何感知条件策略提升大场景重建效率与鲁棒性。

Comments Webpage: https://yutian10.github.io/AnyRecon/

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.19392 2026-04-22 cs.CV 79%

HarmoniDiff-RS: Training-Free Diffusion Harmonization for Satellite Image Composition

HarmoniDiff-RS:无训练扩散谐调用于卫星图像合成

Xiaoqi Zhuang, Jefersson A. Dos Santos, Jungong Han

机构 * The University of Sheffield(谢菲尔德大学) Tsinghua University(清华大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出HarmoniDiff-RS,一种无训练的扩散框架,用于在不同领域条件下谐调合成卫星图像。通过潜在均值偏移操作对齐源域和目标域,结合时间步的潜在融合策略生成候选合成图像,并利用轻量级和谐分类器选择最一致的结果。

Comments 8 pages, 6 figures, CVPR 2026 findings. Code is available at https://github.com/XiaoqiZhuang/HarmoniDiff-RS

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04818 2026-04-22 cs.CV eess.IV stat.ML 79%

Single-Step Reconstruction-Free Anomaly Detection and Segmentation via Diffusion Models

基于扩散模型的实时无重建异常检测与分割

Mehrdad Moradi, Marco Grasso, Bianca Maria Colosimo, Kamran Paynabar

机构 * H. Milton Stewart School of Industrial and Systems Engineering(H. Milton Stewart工业与系统工程学院) Georgia Institute of Technology(佐治亚理工学院) Department of Mechanical Engineering(机械工程系) Polytechnic University of Milan(米兰理工学院)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出RADAR方法,通过注意力机制的扩散模型直接生成异常图,提升检测精度和效率,实验证明在MVTec-AD和3D打印材料数据集上均优于现有方法。

Comments 9 pages, 8 figures, 1 table. Accepted to 2025 International Conference on Machine Learning and Applications (ICMLA)

Journal ref Proc. 2025 International Conference on Machine Learning and Applications (ICMLA), Boca Raton, FL, USA, 2025, pp. 663-670

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.19105 2026-04-22 cs.CV 79%

EgoMotion: Hierarchical Reasoning and Diffusion for Egocentric Vision-Language Motion Generation

视角运动:层次推理与扩散用于中心视觉-语言运动生成

Ruibing Hou, Mingyue Zhou, Yuwei Gui, Mingshuang Luo, Bingpeng Ma, Hong Chang, Shiguang Shan, Xilin Chen

机构 * Key Laboratory of Intelligent Information Processing, Institute of Computing Technology (ICT), Chinese Academy of Sciences (CAS)(智能信息处理重点实验室,计算技术研究所(ICT),中国科学院(CAS)) Jilin University (JLU)(吉林大学(JLU)) Beijing University of Posts and Telecommunications (BUPT)(北京邮电大学(BUPT)) University of the Chinese Academy of Sciences(中国科学院大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出EgoMotion框架,通过层次推理和扩散模型生成基于第一视角视觉和语言指令的3D人类运动,解决语义推理与运动建模的协同优化问题,提升多模态接地和运动质量。

Comments 12 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.18632 2026-04-22 cs.CV stat.AP 79%

StomaD2: An All-in-One System for Intelligent Stomatal Phenotype Analysis via Diffusion-Based Restoration Detection Network

StomaD2:一种基于扩散恢复检测网络的智能气孔表型分析一体化系统

Quanling Zhao, Meng'en Qin, Yanfeng Sun, Yuan Miao, Xiaohui Yang

机构 * Henan Engineering Research Center for Artificial Intelligence Theory and Algorithms(河南人工智能理论与算法工程研究中心) School of Mathematics and Statistics(数学与统计学学院) International Joint Research Laboratory for Global Change Ecology(全球变化生态学联合研究实验室) School of Life Sciences(生命科学学院) State Key Laboratory of Cotton Biology(棉花生物学国家重点实验室)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 StomaD2通过扩散恢复模块和专用旋转目标检测网络,实现复杂成像条件下高精度快速气孔表型分析,其柱状结构、上下文感知重采样机制和特征重组模块提升了特征表示能力,实验表明其在玉米和小麦数据集上准确率达99.4%和99.2%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.09806 2026-04-22 cs.CV cs.AI cs.CL cs.LG 79%

A Generalist Model for Diverse Text-Guided Medical Image Synthesis

一种通用模型用于多样化文本引导的医学图像合成

Joseph Cho, Mrudang Mathur, Cyril Zakka, Dhamanpreet Kaur, Matthew Leipzig, Alex Dalal, Aravind Krishnan, Eubee Koo, Karen Wai, Cindy S. Zhao, Akshay Chaudhari, Matthew Duda, Ashley Choi, Ehsan Rahimy, Lyna Azzouz, Robyn Fong, Rohan Shad, William Hiesinger

机构 * Department of Cardiothoracic Surgery(心脏外科部门) Stanford Medicine(斯坦福医学) Department of Ophthalmology(眼科部门) Division of Cardiovascular Surgery(心血管外科分会) Penn Medicine(宾夕法尼亚医学)

专题命中 扩散模型 :image synthesis(title);diffusion(abstract);分类 cs.CV

AI总结 本文提出MediSyn模型,通过公开数据训练,生成六种医学领域和十种成像模态的合成图像,验证了通用模型在效率、质量和医学应用中的优势。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.18393 2026-04-21 cs.CV 79%

One-Step Diffusion with Inverse Residual Fields for Unsupervised Industrial Anomaly Detection

一步扩散与逆残差场用于无监督工业异常检测

Boan Zhang, Wen Li, Guanhua Yu, Xiyang Liu, Wenchao Chen, Long Tian

机构 * Department of Computer Science and Technology(计算机科学与技术系) Xidian University(西电大学) Department of Electronic Engineering(电子工程系)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出OSD-IRF方法,通过逆残差场实现无监督工业异常检测,利用单步扩散提升推理效率,并在多个基准测试中取得最优或竞争性性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.18313 2026-04-21 cs.CV 79%

Denoise and Align: Diffusion-Driven Foreground Knowledge Prompting for Open-Vocabulary Temporal Action Detection

去噪与对齐:基于扩散的前景知识提示用于开放词汇时序动作检测

Sa Zhu, Wanqian Zhang, Lin Wang, Jinchao Zhang, Cong Wang, Bo Li

机构 * Institute of Information Engineering, Chinese Academy of Sciences School of Cyber Security, University of Chinese Academy of Sciences State Key Laboratory of Cyberspace Security Defense Beijing China Institute of Information Engineering, Chinese Academy of Sciences Beijing China Hangzhou Dianzi University Hangzhou China Institute of Information Engineering, Chinese Academy of Sciences\ Key Laboratory of Cyberspace Security Defense Beijing China Engineering, Zhejiang University Hangzhou China Institute of Information Engineering, Chinese Academy of Sciences State Key Laboratory of Cyberspace Security Defense Beijing China Institute of Information Engineering, Chinese Academy of Sciences School of Cyber Security, University of Chinese Academy of Sciences State Key Laboratory of Cyberspace Security Defense Institute of Information Engineering, Chinese Academy of Sciences Hangzhou Dianzi University Institute of Information Engineering, Chinese Academy of Sciences\ Key Laboratory of Cyberspace Security Defense Engineering, Zhejiang University Institute of Information Engineering, Chinese Academy of Sciences State Key Laboratory of Cyberspace Security Defense

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出DFAlign框架,通过扩散去噪生成前景知识,解决开放词汇时序动作检测中语义不平衡问题,提升动作相关片段的判别性。

Comments Accepted by SIGIR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.13182 2026-04-21 cs.CV 79%

Diffusion-Based Feature Denoising and Using NNMF for Robust Brain Tumor Classification

基于扩散的特征去噪和使用NNMF进行鲁棒性脑肿瘤分类

Hiba Adil Al-kharsan, Róbert Rajkó

机构 * Doctoral School of Computer Science, University of Szeged(计算机科学博士学院,塞格德大学) Academic Staff, Doctoral School of Computer Science, University of Szeged(学术人员,计算机科学博士学院,塞格德大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出结合NNMF、轻量CNN和扩散去噪的鲁棒脑肿瘤分类框架,通过预处理MRI图像并提取可解释特征,提升对抗扰动下的分类鲁棒性。

Comments 30 pages, 29 figures

Journal ref Mach. Learn. Knowl. Extr. 2026, 8(4), 105

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26010 2026-04-21 cs.CV 79%

New Fourth-Order Grayscale Indicator-Based Telegraph Diffusion Model for Image Despeckling

新的第四阶灰度指标基于电报扩散模型用于图像去斑

Rajendra K. Ray, Manish Kumar

机构 * School of Mathematical and Statistical Sciences, Indian Institute of Technology Mandi(印度理工学院曼迪数学与统计科学学院) Department of Mathematics, Indian Institute of Technology Delhi(印度理工学院德里数学系)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出一种第四阶非线性PDE模型,结合扩散和波特性,以改善去斑效果,通过PSNR、MSSIM和SI指标验证其有效性,并扩展至彩色图像处理。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.08013 2026-04-21 cs.CV cs.AI cs.LG 79%

StableMTL: Repurposing Latent Diffusion Models for Multi-Task Learning from Partially Annotated Synthetic Datasets

StableMTL: 利用潜在扩散模型重新利用多任务学习于部分标注的合成数据集

Anh-Quan Cao, Ivan Lopes, Raoul de Charette

机构 * Inria(法国国家科研机构)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出StableMTL,通过利用扩散模型的泛化能力,在部分标注的合成数据集上进行多任务学习,采用统一的潜在损失和多流模型促进任务间协同,优于基线模型。

Comments Accepted at CVPR 2026. Code is at https://github.com/astra-vision/StableMTL

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.17818 2026-04-21 cs.CV 79%

AnyLift: Scaling Motion Reconstruction from Internet Videos via 2D Diffusion

AnyLift: 通过2D扩散模型从互联网视频中扩展运动重建

Hongjie Li, Heng Yu, Jiaman Li, Hong-Xing Yu, Ehsan Adeli, C. Karen Liu, Jiajun Wu

机构 * Stanford University(斯坦福大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出AnyLift框架,利用2D扩散模型从互联网视频中重建3D人体运动和人-物交互,通过合成多视角2D运动数据和训练相机条件多视角2D运动扩散模型,提升动态摄像下运动重建的准确性和鲁棒性。

Comments CVPR 2026. Project website: https://awfuact.github.io/anylift/ The first two authors contribute equally

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.03692 2026-04-21 cs.CV cs.AI 79%

Error as Signal: Stiffness-Aware Diffusion Sampling via Embedded Runge-Kutta Guidance

误差作为信号:通过嵌入式龙格-库塔引导的刚性感知扩散采样

Inho Kong, Sojin Lee, Youngjoon Hong, Hyunwoo J. Kim

机构 * Korea University(韩国大学) KAIST(韩国科学技术院) Seoul National University(首尔国立大学) Korea Institute for Advanced Study(韩国高级研究院)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出嵌入式龙格-库塔引导方法,通过识别刚性区域减少局部截断误差,提升扩散模型采样稳定性与质量。

Comments ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.20033 2026-04-21 cs.CV 79%

FlashLips: 100-FPS Mask-Free Latent Lip-Sync using Reconstruction Instead of Diffusion or GANs

FlashLips: 100-FPS 无掩码的潜在唇部同步使用重建而非扩散或GANs

Andreas Zinonos, Michał Stypułkowski, Antoni Bigata, Stavros Petridis, Maja Pantic, Nikita Drobyshev

机构 * Imperial College London(帝国理工学院伦敦分校) Cantina Labs(Cantina实验室) NatWest AI Research(NatWest人工智能研究)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 FlashLips是一种无掩码的唇部同步系统,通过重建而非扩散或GANs实现实时性能,其U-Net变体在单GPU上达100FPS,视觉质量媲美先进模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19979 2026-04-21 cs.CV 79%

CamPVG: Camera-Controlled Panoramic Video Generation with Epipolar-Aware Diffusion

CamPVG:基于视图控制的全景视频生成与视图感知扩散

Chenhao Ji, Chaohui Yu, Junyao Gao, Fan Wang, Cairong Zhao

机构 * Tongji University(同济大学) DAMO Academy, Alibaba Group(阿里达摩院)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出CamPVG,首个基于视图控制的全景视频生成框架,通过全景Plücker嵌入和视图感知模块提升生成视频的质量与一致性。

Comments SIGGRAPH Asia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14461 2026-04-21 cs.CV 79%

Ouroboros: Single-step Diffusion Models for Cycle-consistent Forward and Inverse Rendering

Ouroboros: 单步扩散模型用于循环一致的正向与反向渲染

Shanlin Sun, Yifan Wang, Hanwen Zhang, Yifeng Xiong, Qin Ren, Ruogu Fang, Xiaohui Xie, Chenyu You

机构 * University of California, Irvine(加州大学伊维特分校) Stony Brook University(石溪大学) Huazhong University of Science and Technology(华中科技大学) University of Florida(佛罗里达大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 Ouroboros通过双单步扩散模型实现正反向渲染的互促,扩展了内在分解到室内外场景,并引入循环一致性机制,实验显示其在多样场景中表现优异且推理速度显著提升。

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07826 2026-04-21 cs.CV cs.LG cs.RO 79%

R3D2: Realistic 3D Asset Insertion via Diffusion for Autonomous Driving Simulation

R3D2:通过扩散实现自动驾驶模拟中的真实3D资产插入

William Ljungbergh, Bernardo Taveira, Wenzhao Zheng, Adam Tonderski, Chensheng Peng, Fredrik Kahl, Christoffer Petersson, Michael Felsberg, Kurt Keutzer, Masayoshi Tomizuka, Wei Zhan

机构 * Zenseact Linköping University(林哈尔大学) Chalmers University(查尔姆斯理工大学) UC Berkeley(加州大学伯克利分校)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 R3D2是一种轻量级单步扩散模型,用于在现有场景中实时生成真实3D资产,通过训练学习真实集成,提升自动驾驶模拟的真实性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.16919 2026-04-21 cs.LG cs.AI cs.CV 79%

Noise-Adaptive Diffusion Sampling for Inverse Problems Without Task-Specific Tuning

具有任务无关调优的噪声自适应扩散采样用于逆问题

Yingzhi Xia, Setthakorn Tanomkiattikun, Liangli Zhen, Zaiwang Gu

机构 * Institute of High Performance Computing, Agency for Science, Technology and Research, Singapore(高性能计算研究所,科技研究局,新加坡) Institute for Infocomm Research, Agency for Science, Technology and Research, Singapore(信息通信研究所,科技研究局,新加坡) Johns Hopkins University(约翰·霍普金斯大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出噪声空间Hamilton蒙特卡洛方法,用于逆问题求解,避免局部极值并提升鲁棒性,通过噪声空间探索提升重建质量。

Comments Accepted by ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.16841 2026-04-21 cs.CV cs.LG 79%

When Earth Foundation Models Meet Diffusion: An Application to Land Surface Temperature Super-Resolution

当地球基础模型遇见扩散:一种用于地表温度超分辨率的应用

Yiheng Chen, Zihui Ma, Peishi Jiang, Yilong Dai, Qikai Hu, Xinyue Ye, Lingyao Li, Rita Sousa, Runlong Yu

机构 * University of Alabama(阿拉巴马大学) Emory University(埃默里大学) University of Michigan(密歇根大学) University of South Florida(佛罗里达州立大学) New York University(纽约大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出EFDiff框架,利用地球基础模型指导扩散模型进行极端空间退化下的超分辨率重建,通过交叉注意力机制提升细尺度重建效果,优于基线方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.16479 2026-04-21 cs.CV cs.AI 79%

Latent-Compressed Variational Autoencoder for Video Diffusion Models

视频扩散模型的潜在压缩变分自编码器

Jiarui Guan, Wenshuai Zhao, Zhengtao Zou, Juho Kannala, Arno Solin

机构 * Aalto University(阿alto大学) ELLIS Institute Finland(ELLIS研究所芬兰) University of Oulu(奥卢大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出一种潜在压缩方法,通过去除视频潜在表示中的高频成分而非直接减少通道数,提升视频生成质量与压缩效率。

Comments Accepted to CVPR 2026 findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.09721 2026-04-21 cs.CV 79%

FrameDiT: Diffusion Transformer with Matrix Attention for Efficient Video Generation

FrameDiT:基于矩阵注意力的扩散变换器用于高效视频生成

Minh Khoa Le, Kien Do, Duc Thanh Nguyen, Truyen Tran

机构 * Applied Artificial Intelligence Initiative, Deakin University, Australia(德肯大学应用人工智能倡议,澳大利亚) FPT Smart Cloud, Vietnam(越南FPT智能云) Deakin University, Australia(德肯大学,澳大利亚)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出Matrix Attention,通过将帧视为矩阵来处理视频生成中的时空动态,结合FrameDiT-G和FrameDiT-H模型,在保持效率的同时提升了视频质量和时间一致性。

Comments Code: https://github.com/minhkhoale/FrameDiT Accepted at CVPR 2026 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏