arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

共收录 86872 信号源:cs.CV, cs.GR, cs.MM

1. 扩散模型 70277 篇

2606.03903 2026-06-03 cs.CV 79%

An Attention-Based Denoising Model for Diffusion Weighted Imaging

一种基于注意力的扩散加权成像去噪模型

Prithviraj Verma, Pawan Kumar, Chandan Deshani, Prasun Chandra Tripathi

机构 * Institute of Infrastructure Technology Research and Management (IITRAM)(基础设施技术研究与管理研究所) University of Sheffield(谢菲尔德大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 提出一种结合Swin Transformer窗口注意力和多维门控精化的噪声感知注意力驱动去噪框架,用于解决DWI中信号依赖的Rician噪声问题,在1%至15%噪声水平下实现平均PSNR 33.69 dB和SSIM 0.8539。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09708 2026-06-03 cs.LG cs.AI cs.CV cs.NA math.NA 79%

Physics-informed diffusion models in spectral space

谱空间中的物理信息扩散模型

Davide Gallon, Philippe von Wurstemberger, Patrick Cheridito, Arnulf Jentzen

机构 * ETH Zürich(苏黎世联邦理工学院)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 提出物理信息谱扩散(PISD)方法,结合生成式潜扩散模型与物理信息机器学习,在谱表示潜空间中对偏微分方程参数和解进行扩散建模,通过扩散后验采样施加物理约束和测量条件,在泊松、亥姆霍兹和不可压缩纳维-斯托克斯方程上展现出比现有扩散求解器更高的精度和计算效率。

Comments 18 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09105 2026-06-03 cs.CV 79%

Hybrid Autoregressive-Diffusion Model for Real-Time Sign Language Production

混合自回归-扩散模型用于实时手语生成

Maoxiao Ye, Xinfeng Ye, Mano Manoharan

机构 * University of Auckland(奥克兰大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 提出HybridSign混合自回归-扩散模型,结合因果帧生成与流式扩散精炼,实现低延迟高质量手语生成,在PHOENIX14T和How2Sign上取得最佳质量-效率权衡。

Comments Accepted at ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
1202.5414 2026-06-03 math.AP cs.CV cs.NA math.NA math.RT 79%

Left-Invariant Diffusion on the Motion Group in terms of the Irreducible Representations of SO(3)

基于SO(3)不可约表示的运动群上的左不变扩散

Marco Reisert, Henrik Skibbe

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 利用SO(3)不可约表示将SE(3)上的左不变向量场表示为平移坐标的微分形式和旋转的代数形式,避免了对SO(3)或S2的显式离散化,并应用于扩散加权磁共振成像和目标检测。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.02532 2026-06-02 cs.CV 79%

Improving Combined Detection and Classification of TEM Defects via Mask-Conditioned Latent Diffusion Augmentation

通过掩码条件潜在扩散增强改善TEM缺陷的联合检测与分类

Ni Li, Nuohao Liu, Ryan Jacobs, Ajay Annamareddy, Maciej P. Polak, Kevin Field, Izabela Szlufarska, Dane Morgan

机构 * University of Wisconsin-Madison(威斯康星大学麦迪逊分校) University of Michigan-Ann Arbor(密歇根大学安娜堡分校)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 提出一种基于掩码条件潜在扩散模型(LDM)的生成式数据增强方法,用于合成可控、自动标注的多类缺陷掩码的TEM图像,以提升小样本下Mask R-CNN模型的缺陷检测与分类性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.02129 2026-06-02 cs.CV 79%

Equilibrated Diffusion: Frequency-aware Textual Embedding for Equilibrated Image Customization

均衡扩散:面向均衡图像定制的频率感知文本嵌入

Liyuan Ma, Xueji Fang, Guo-Jun Qi

机构 * Westlake University(西湖大学) Zhejiang University(浙江大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 提出均衡扩散方法,通过频率空间分解概念特征并独立优化嵌入,实现风格与主体解耦,提升定制图像的保真度和文本对齐。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.02105 2026-06-02 cs.CV 79%

Multimodal Action Diffusion for Robust End-to-End Autonomous Driving

多模态动作扩散用于鲁棒的端到端自动驾驶

Jorge Daniel Rodríguez-Vidal, Diego Porres, Gabriel Villalonga Pineda, Antonio M. López Peña

机构 * Computer Vision Center (CVC)(计算机视觉中心) Universitat Autònoma de Barcelona (UAB)(巴塞罗那自治大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 提出动作扩散变换器(ADT),通过多模态动作建模和最近邻匹配,在闭环Bench2Drive基准上超越先前最优方法,同时延迟降低十倍。

Comments Preprint. June 1st, 2026. Corresponding author: Jorge Daniel Rodríguez-Vidal

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.02000 2026-06-02 cs.CV cs.AI eess.IV 79%

Towards 3D-Aware Video Diffusion Models: Render-Free Human Motion Control with Mesh Tokenization

迈向3D感知视频扩散模型:基于网格标记化的无渲染人体运动控制

Jingyun Liang, Min Wei, Shikai Li, Yizeng Han, Hangjie Yuan, Lei Sun, Weihua Chen, Fan Wang

机构 * DAMO Academy, Alibaba Group(阿里巴巴集团大模型实验室) Hupan Lab(虎盘实验室) Zhejiang University(浙江大学) INSAIT

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 提出一种无渲染框架,通过压缩的3D人体网格标记直接条件化视频生成,实现精确的人体运动控制,减少2D引导伪影并提升3D结构建模能力。

Comments Project page: https://jingyunliang.github.io/MeshToken/

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.01048 2026-06-02 cs.CV 79%

Decoupled Residual Denoising Diffusion Models for Unified and Data Efficient Image-to-Image Translation

解耦残差去噪扩散模型用于统一且数据高效的图像到图像翻译

Ziyue Lin, Jiahe Hou, Hongyu Xia, Xinrui Xie, Feifei Wang, Yuyin Zhou, Wei Wang, Jiawei Liu, Liangqiong Qu

机构 * The University of Hong Kong(香港大学) Shenyang Institute of Automation, Chinese Academy of Sciences(中国科学院沈阳自动化研究所) The Chinese University of Hong Kong(香港中文大学) University of California, Santa Cruz(加州大学圣克鲁兹分校)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 提出解耦残差去噪扩散模型(DRDD),通过将扩散过程解耦为随机噪声扩散和确定性残差扩散两个独立阶段,实现统一且数据高效的图像到图像翻译。

Comments CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.00957 2026-06-02 cs.CV 79%

Boundary-Protection W8A8 HiFloat8 Quantization for Large-Scale Text-to-Video Diffusion Transformers

面向大规模文生视频扩散Transformer的边界保护W8A8 HiFloat8量化

Yiming Zhao

机构 * Yiming Zhao(赵毅铭)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 针对Wan2.1-T2V-14B模型,提出一种边界保护策略的W8A8 HiF8后训练量化方法,通过保留首尾边界块为BF16而量化中间块,在VBench五个维度上匹配或略优于BF16基线。

Comments 6 pages, 5 figures. Accepted to ICME 2026 Grand Challenge

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.00803 2026-06-02 astro-ph.CO cs.CV cs.LG 79%

Generative Diffusion Priors for 3D Mapping of the Dark Universe

用于暗宇宙三维映射的生成扩散先验

Brandon Zhao, Diana Scognamiglio, Olivier Doré, Katherine L. Bouman

机构 * Department of Computing and Mathematical Sciences, California Institute of Technology(加州理工学院计算与数学科学系) Jet Propulsion Laboratory, California Institute of Technology(加州理工学院喷气推进实验室) Department of Physics, Duke University(杜克大学物理系) Cahill Center for Astronomy and Astrophysics, California Institute of Technology(加州理工学院卡希尔天文与天体物理中心)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 利用扩散模型学习宇宙模拟中的先验分布,结合物理正向模型解决弱引力透镜三维暗物质反问题,显著提升重建精度并生成统计一致的后验样本。

Comments Accepted to CVPR 2026 (Highlight)

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.00798 2026-06-02 cs.CV cs.AI cs.LG 79%

DASH: Dual-Branch Score Distillation for Guidance-Calibrated Compact Diffusion Models

DASH: 用于引导校准紧凑扩散模型的双分支分数蒸馏

Abdullah Al Shafi, Kazi Saeed Alam, Sk Imran Hossain, Engelbert Mephu Nguifo

机构 * Khulna University of Engineering & Technology(Khulna 工程与技术大学) University Clermont Auvergne(克莱蒙特-奥弗涅大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 针对类条件扩散模型参数压缩中无监督无条件分数分支导致引导失效的问题,提出双分支蒸馏框架DASH,通过独立监督两个分支并引入锚点正则化和课程迁移,在5.9倍压缩下保持与教师模型相近的FID和引导保真度。

Comments 14 pages, 7 figures, 4 tables; appendix with additional ablations and qualitative results

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.00393 2026-06-02 eess.IV cs.CV 79%

AutoIQ: An Ensemble Framework for Automatic Assessment of Geometric Distortion in Prostate Diffusion-Weighted Imaging

AutoIQ:前列腺扩散加权成像中几何畸变自动评估的集成框架

Haoran Sun, Lixia Wang, Yin-Chen Hsu, Hsu-Lei Lee, Chang Gao, Fei Han, Robert Grimm, Vibhas Deshpande, Ziyang Long, Hsin-Jung Yang, Rola Saouaf, Alessandro D'Agnolo, Timothy Daskivich, Hyung Kim, Debiao Li, Yibin Xie

机构 * Biomedical Imaging Research Institute, Cedars-Sinai Medical Center(生物医学成像研究 institute, Cedars-Sinai 医疗中心) Department of Bioengineering, University of California(生物工程系,加州大学) Siemens Medical Solutions USA Inc.(西门子医疗解决方案美国公司) Siemens Healthineers AG(西门子健康影像股份有限公司) Department of Imaging, Cedars-Sinai Medical Center(成像部,Cedars-Sinai 医疗中心) Department of Nuclear Medicine, Cedars-Sinai Medical Center(核医学部,Cedars-Sinai 医疗中心) Department of Urology, Cedars-Sinai Medical Center(泌尿科,Cedars-Sinai 医疗中心)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 提出AutoIQ集成机器学习框架,结合分割和配准方法量化DWI几何畸变,用于自动分类畸变严重程度,在独立测试集上达到0.95准确率。

Comments Original research; 11 pages, 7 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.00299 2026-06-02 cs.CV cs.AI 79%

Real2SAM2Real: Generative 3D Caches as Complementary Context for Video Diffusion

Real2SAM2Real: 生成式3D缓存作为视频扩散的互补上下文

Jiayi Wu, Haoming Cai, Cornelia Fermuller, Christopher Metzler, Yiannis Aloimonos

机构 * University of Maryland(马里兰大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 提出Real2SAM2Real框架,通过3D提升模型提取可编辑的3D缓存作为几何支架,结合软空间对齐注入和微调策略,实现视频扩散模型对相机轨迹和多实体运动的精确解耦控制。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.00153 2026-06-02 cs.CV cs.AI 79%

DiffCrossGait: Trajectory-Level Alignment for 2D-3D Cross-Modal Gait Recognition via Latent Diffusion

DiffCrossGait:基于潜在扩散的2D-3D跨模态步态识别轨迹级对齐

Zhiyang Lu, Ming Cheng

机构 * University of Science and Technology of China(中国科学技术大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 针对2D-3D跨模态步态识别中的域差异问题,提出DiffCrossGait,通过潜在扩散空间中的轨迹级对齐实现连续模态对齐,并引入三阶段对齐策略确保身份锚定、动态一致性和跨模态结构可恢复性,在SUSTech1K和FreeGait基准上达到最优性能。

Comments Accepted by ICML2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.00109 2026-06-02 cs.CV cs.AI cs.LG 79%

VDSB-GWSyn: Diffusion Schrödinger Bridge for Controllable and Anatomically Feasible Guidewire Synthesis in Coronary Angiography

VDSB-GWSyn: 用于冠状动脉造影中可控且解剖学可行的导丝合成的扩散薛定谔桥

Haoyuan Tang, Zhuo Zhang, Jialin Li, Shuai Xiao, Jiachen Yang

机构 * Tianjin University(天津大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 提出基于扩散薛定谔桥的VDSB-GWSyn框架,通过形状先验和血管分割约束生成可控、高保真导丝样本,显著提升下游导丝端点定位精度。

Comments Early accept to MICCAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.31162 2026-06-02 cs.CV cs.LG 79%

Guidance for Low-Level Perceptual Editing in Unconditional Diffusion Models

无条件扩散模型中低级感知编辑的引导

Shreyansh Modi, Akshat Tomar, Aarush Aggarwal

机构 * Indian Institute of Technology Roorkee(印度理工学院罗尔基)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 针对无条件扩散模型在美学和感知增强中难以进行全局低级变换的问题,提出一种无需训练的推理时机制,通过提取退化概念向量并结合瓶颈修补与无分类器引导,实现图像编辑与质量提升。

Comments 11 pages, 12 figures, Generative Models for Computer Vision Workshop CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.16415 2026-06-02 cs.CV cs.LG 79%

Diffusion Models, Denoiser Architecture and Creativity

扩散模型、去噪器架构与创造力

Itamar Levine, Yair Weiss

机构 * The Hebrew University of Jerusalem(海法大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文通过理论和实验表明,扩散模型的创造力源于去噪器架构与目标分布之间的相互作用,并指出去噪器架构的归纳偏差必须与真实目标分布高度一致才能成功。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11158 2026-06-02 eess.IV cs.CV 79%

Diffusion Models for Hyperspectral Image Analysis: A Comprehensive Review

扩散模型在高光谱图像分析中的应用:综述

Xing Hu, Xiangcheng Liu, Qianqian Duan, Lian Zhang, Huiliang Shang, Linhua Jiang, Haima Yang, Dawei Zhang

机构 * School of Optical-Electrical and Computer Engineering, University of Shanghai for Science and Technology(上海理工大学光学电子与计算机工程学院) School of Electronics and Electrical Engineering, Shanghai University of Engineering Science(上海工程技术大学电子与电气工程学院) Medical Artificial Intelligence Lab, The First Hospital of Hebei Medical University, Hebei Medical University(河北医科大学第一医院医学人工智能实验室) Hangzhou Institute of Technology, xidian University(杭州职业技术学院)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文系统综述了扩散模型(包括去噪扩散概率模型和基于随机微分方程的生成框架)在高光谱图像处理中的最新进展,分类现有方法,强调其处理高维数据的优势,并与传统方法比较性能,特别关注变化检测和灾后异常识别等关键应用,同时讨论计算成本和训练稳定性等局限,并展望未来研究方向。

Comments Published in Neural Networks

Journal ref Neural Networks (2026) 109109

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02214 2026-06-02 cs.CV 79%

Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video Generation

因果强迫:自回归扩散蒸馏的正确方法,用于高质量实时交互式视频生成

Hongzhou Zhu, Min Zhao, Guande He, Hang Su, Chongxuan Li, Jun Zhu

机构 * Hongzhou Zhu(朱洪洲) Min Zhao(赵敏) Guande He(何冠德) Hang Su(苏hang) Chongxuan Li(李崇轩) Jun Zhu(朱军)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 针对双向扩散模型蒸馏为自回归模型时的架构差距问题,提出因果强迫方法,通过自回归教师进行ODE初始化并应用DMD过程,显著提升视频生成质量。

Comments Project page and the code: \href{https://thu-ml.github.io/CausalForcing.github.io/}{https://thu-ml.github.io/CausalForcing.github.io/}; https://github.com/thu-ml/Causal-Forcing. ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.09503 2026-06-02 cs.CV 79%

PermuQuant: Lowering Per-Group Quantization Error by Reordering Channels for Diffusion Models

PermuQuant:通过重新排列通道降低扩散模型每组量化误差

Yongsen Cheng, Kai Liu, Kaiwen Tao, Junxian Li, Zhixin Wang, Zhikai Chen, Renjing Pei, Yulun Zhang

机构 * Shanghai Jiao Tong University(上海交通大学) Huawei Noah’s Ark Lab(华为诺亚实验室)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 提出PermuQuant框架,通过基于联合二阶矩的通道重排序和校准接受规则,降低低比特扩散模型每组量化误差,实现显著加速和内存压缩。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.09063 2026-06-02 cs.CV cs.AI 79%

Frequency-Enhanced Diffusion Models: Curriculum-Guided Semantic Alignment for Zero-Shot Skeleton Action Recognition

频率增强扩散模型:基于课程引导语义对齐的零样本骨架动作识别

Yuxi Zhou, Zhengbo Zhang, Jingyu Pan, Zhiyu Lin, Zhigang Tu

机构 * State Key Laboratory of Information Engineering in Surveying, Mapping and Remote Sensing(测绘遥感信息工程国家重点实验室) Wuhan University(武汉大学) Information Systems Technology and Design Pillar(信息系统技术与设计学院) Singapore University of Technology and Design(新加坡科技与设计大学) School of Geodesy and Geomatics(测绘学院) School of Mathematics and Statistics(数学与统计学院) Wuhan University Shenzhen Research Institute(武汉大学深圳研究院)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 提出频率感知扩散模型FDSM,通过语义引导频谱残差模块、时间步自适应频谱损失和课程语义抽象,解决扩散模型频谱偏差导致的高频动态过度平滑问题,实现零样本骨架动作识别,在多个数据集上达到最优性能。

Comments Accepted by The Visual Computer

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.27645 2026-06-02 cs.CV 79%

OpenDPR: Open-Vocabulary Change Detection via Vision-Centric Diffusion-Guided Prototype Retrieval for Remote Sensing Imagery

OpenDPR:面向遥感影像的基于视觉中心扩散引导原型检索的开放词汇变化检测

Qi Guo, Jue Wang, Yinhe Liu, Yanfei Zhong

机构 * Wuhan University(武汉大学) Beijing Institute of Technology(北京理工大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 提出OpenDPR框架,通过扩散模型构建原型并检索视觉相似性,解决开放词汇变化检测中类别识别瓶颈,并设计S2C模块增强变化定位能力。

Comments Accepted by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.19848 2026-06-02 cs.CV 79%

DerMAE: Improving skin lesion classification through conditioned latent diffusion and MAE distillation

DerMAE: 通过条件潜在扩散和MAE蒸馏改进皮肤病变分类

Francisco Filho, Kelvin Cunha, Fábio Papais, Emanoel dos Santos, Rodrigo Mota, Thales Bezerra, Erico Medeiros, Paulo Borba, Tsang Ing Ren

机构 * Universidade Federal do Pernambuco(佛罗里达州帕尔马大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 针对皮肤病变分类中恶性样本不足导致的类别不平衡问题,提出使用类别条件扩散模型生成合成图像,结合自监督MAE预训练学习鲁棒特征,并通过知识蒸馏将大模型知识迁移至轻量级ViT学生模型,在提升分类性能的同时实现高效设备端推理。

Comments 4 pages, 2 figures, 1 table, Published in: 2026 IEEE 23rd International Symposium on Biomedical Imaging (ISBI)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20072 2026-06-02 cs.CV cs.LG cs.RO 79%

Discrete Diffusion VLA: Bringing Discrete Diffusion to Action Decoding in Vision-Language-Action Policies

离散扩散VLA:将离散扩散引入视觉-语言-动作策略中的动作解码

Zhixuan Liang, Yizhuo Li, Tianshuo Yang, Chengyue Wu, Sitong Mao, Liuao Pei, Tian Nian, Shunbo Zhou, Xiaokang Yang, Jiangmiao Pang, Yao Mu, Ping Luo

机构 * University of Science and Technology of China(中国科学技术大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 提出离散扩散VLA,通过将动作块离散化并在统一Transformer骨干内使用离散扩散模式进行渐进细化,实现自适应解码顺序和错误纠正,在多个基准上取得高性能并保留预训练的视觉-语言先验。

Comments Accepted by ICML 2026. 17 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06136 2026-06-02 cs.CV cs.AI 79%

GSV3D: Gaussian Splatting-based Geometric Distillation with Stable Video Diffusion for Single-Image 3D Object Generation

GSV3D: 基于高斯溅射的几何蒸馏与稳定视频扩散用于单图像3D物体生成

Ye Tao, Jiawei Zhang, Yahao Shi, Dongqing Zou, Bin Zhou

机构 * State Key Laboratory of Virtual Reality Technology and Systems, Beihang University(虚拟现实技术与系统国家重点实验室,北京航空航天大学) SenseTime Research(商汤科技研究院) PBVR

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 提出一种结合2D扩散模型隐式3D推理能力与高斯溅射几何蒸馏的方法,通过高斯溅射解码器将SV3D潜变量输出转换为显式3D表示,实现多视图一致性和高质量3D生成。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.31596 2026-06-01 cs.CV cs.LG 79%

KLIP: localized distribution shift detection via KL-divergence with diffusion priors in Inverse Problems

KLIP:通过逆问题中扩散先验的KL散度进行局部分布偏移检测

Alireza Kheirandish, Jihoon Hong, Sara Fridovich-Keil

机构 * Georgia Institute of Technology(佐治亚理工学院)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 提出基于KL散度的OOD检测指标,无需校准数据或偏移分布知识,可检测并定位图像中的局部分布偏移。

Comments CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.31590 2026-06-01 cs.CV cs.AI 79%

TunerDiT: Training-free Progressive Steering of Diffusion Transformer for Multi-Event Video Generation

TunerDiT: 无需训练的多事件视频生成扩散变压器渐进式引导

Ruotong Liao, Guowen Huang, Qing Cheng, Guangyao Zhai, Lei Zhang, Xun Xiao, Thomas Seidl, Daniel Cremers, Volker Tresp

机构 * Ludwig Maximilian University of Munich(慕尼黑路德维希-马克西米利安大学) Technical University of Munich(慕尼黑技术大学) MCML University of Hamburg(汉堡大学) Huawei European Research Institute(华为欧洲研究院)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 针对长视频多事件生成难题,提出无需额外训练的TunerDiT方法,通过事件分区掩码和跨事件提示融合实现渐进式引导,在8项指标上达到最优。

Comments 17 pages, 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.31115 2026-06-01 cs.CV 79%

Polyphony: Diffusion-based Dual-Hand Action Segmentation with Alternating Vision Transformer and Semantic Conditioning

Polyphony: 基于扩散的双手动作分割,采用交替视觉Transformer和语义条件

Hao Zheng, Hu Wang, Tiantian Zheng, Prajjwal Bhattarai, Tuka Alhanai

机构 * New York University Abu Dhabi(纽约大学阿布扎赫尔分校) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 提出Polyphony三阶段方法,通过交替训练双手视觉Transformer、语义特征条件化和扩散分割,解决双手动作分割中的手间依赖、视觉不对称和语义模糊问题,在多个数据集上达到最优性能。

Comments CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.31057 2026-06-01 cs.CV cs.LG 79%

LVSA: Training-Free Sparse Attention for Long Video Diffusion

LVSA:长视频扩散的无训练稀疏注意力

Gael Glorian, Ioannis Lamprou, Zhen Zhang, Yujie Yuan, Hongsheng Liu

机构 * Distributed Parallel Technology Laboratory, Paris Research Center, Huawei Technologies France(华为法国巴黎研究中心分布式并行技术实验室) AI Framework and Data Technology Lab, Huawei Technologies Co., Ltd.(华为技术有限公司人工智能框架与数据技术实验室)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 提出一种无需训练、模型无关的块稀疏注意力方法LVSA,通过结构化窗口模式与旋转全局锚点结合,在降低长视频扩散推理计算成本的同时消除固定网格偏差,支持超训练时域的视频生成。

Comments 10 pages, 5 figures, 4 tables. Code: https://github.com/JiusiServe/LongVideoSparseAttention

详情

展开后加载摘要…

URL PDF HTML 收藏