arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

共收录 86872 信号源:cs.CV, cs.GR, cs.MM

1. 扩散模型 70277 篇

2608.21710 2026-08-25 cs.CV 新提交 79%

StereoDiffuer: Diffusion-based Progressive Geometry Modeling with Saliency Attention Perception for Stereo Matching

StereoDiffuer:基于扩散的显著性注意力感知渐进几何建模的立体匹配方法

Bohan Li

机构 * Shanghai Jiaotong University(上海交通大学) Ningbo Institute of Digital Twin, Eastern Institute of Technology(宁波东方理工大学数字孪生研究院)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 针对现有立体匹配方法难以保留细粒度几何细节的问题,提出基于扩散的StereoDiffuer框架,通过显著性注意力感知模块提取几何线索,结合迭代去噪扩散过程优化视差,在Scene Flow与KITTI基准上表现具竞争力。

Comments 17 pages, 9 figures, and 11 tables. Accepted by Signal Processing: Image Communication

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.26203 2026-08-25 cs.CV 版本更新 79%

WildShadowRemover: In-the-Wild Video Shadow Removal via Detail-Preserving Video Diffusion Models

WildShadowRemover:基于细节保留视频扩散模型的野外视频阴影去除方法

Jiamin Xu, Cong Wang, Zheng Dong, Chi Wang, Renshu Gu, Weiwei Xu, Gang Xu

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 针对野外视频阴影去除难题,提出WildShadowRemover框架,通过LoRA微调适配视频扩散模型,结合细节注入、频率分解调制及Depth Anything 3深度先验,构建对应数据集,在阴影去除质量与时间一致性上优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.17825 2026-08-25 cs.CV 版本更新 79%

Steering Video Diffusion Transformers with Massive Activations

通过大规模激活引导视频扩散变换器

Xianhang Cheng, Yujian Zheng, Zhenyu Xie, Tingting Liao, Hao Li

机构 * MBZUAI

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 研究视频扩散变换器中大规模激活的作用,提出STAS方法通过引导首帧和边界token的激活值提升视频生成质量与时间一致性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.24013 2026-08-25 cs.CV 版本更新 79%

Ordinal Diffusion Models for Color Fundus Images

序数扩散模型用于彩色视网膜图像

Gustav Schmidt, Philipp Berens, Sarah Müller

机构 * Hertie Institute for AI in Brain Health, University of T\"ubingen, Germany

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出序数潜在扩散模型,用于生成具有连续严重程度结构的彩色视网膜图像,通过标量疾病表示实现相邻阶段的平滑过渡,提升了生成质量与临床一致性。

Comments MICCAI 2026 accepted manuscript (post-rebuttal)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05575 2026-08-25 cs.CV 版本更新 79%

DiffSwap++: 3D Latent-Controlled Diffusion for Identity-Preserving Face Swapping

DiffSwap++:用于身份保留人脸交换的3D潜在控制扩散模型

Weston Bondurant, Arkaprava Sinha, Hieu Le, Srijan Das, Stephanie Schuckers

机构 * UNC Charlotte(北卡罗来纳大学夏洛特分校)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本研究提出DiffSwap++,一种融入3D人脸潜在特征的扩散人脸交换方法,可在保留源身份的同时维持目标姿态表情,在多数据集实验及评估中性能优于现有方法。

Comments IJCB 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.20756 2026-08-25 cs.CV 版本更新 79%

StereoDiff: Stereo-Diffusion Synergy for Video Depth Estimation

StereoDiff:用于视频深度估计的立体-扩散协同框架

Haodong Li, Chen Wang, Jiahui Lei, Kostas Daniilidis, Lingjie Liu

机构 * Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) University of Pennsylvania(宾夕法尼亚大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 StereoDiff将立体匹配与视频深度扩散协同,针对视频静态和动态区域分别采用对应技术,在真实世界动态视频深度基准上实现了当前最优的零样本估计性能。

Comments We no longer wish this manuscript to stand as an active preprint and will not be maintaining or updating it further; we would prefer it not be cited as current work. It has not been published or accepted at any journal or conference, so the request is not motivated by publication elsewhere. We are requesting a standard withdrawal, not a full removal

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.05389 2026-08-25 cs.CV cs.AI 版本更新 79%

Image-Conditional Diffusion Transformer for Underwater Image Enhancement

用于水下图像增强的图像条件扩散Transformer

Xingyang Nie, Caoliang Zhang, Xiaoyu Zhai, Fengzhong Qu, Biao Wang, Huilin Ge

机构 * Ocean College, Jiangsu University of Science and Technology(江苏科技大学海洋学院) Nanjing Research Institute of Electronic Equipment, China Aerospace Science and Industry Corporation(中国航天科工集团有限公司南京电子设备研究所) Ocean College, Zhejiang University(浙江大学海洋学院)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 该研究提出基于图像条件扩散Transformer(ICDT)的水下图像增强方法,以Transformer替换DDPM的U-Net骨干,在水下ImageNet数据集上验证,其最大模型ICDT-XL/2实现了当前最优增强性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.21229 2026-08-24 cs.CV 新提交 79%

Anchoring Instruction Outside Mask: Exact Reference Caching for Efficient In-Context Diffusion Transformers

在掩码外锚定指令:用于高效上下文扩散Transformer的精确引用缓存

Yangshuai Liu, Zheming Li, Jiaao Li, Kang He, Ziliang Lai, Zhitai Liu, Chengru Song

机构 * Harbin Institute of Technology(哈尔滨工业大学) KlingAI Research(KlingAI研究院)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 该研究针对上下文扩散Transformer中引用增多导致计算量过大的问题,提出掩码外锚定指令的精确引用缓存方法,通过静态文本锚点结合速度蒸馏实现高效图像编辑,在保持生成质量的同时大幅提升了去噪速度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.20759 2026-08-24 cs.CV 新提交 79%

DiGS-Avatar: Single-Image Animatable 3D Human Reconstruction via UV-Space Diffusion

DiGS-Avatar:基于UV空间扩散的单图像可动画三维人体重建

Jiakun Li, Li Fang, Hao Zhu, Fei Hu, Long Ye, Yuan Zhang, Jinyao Yan

机构 * Key Laboratory of Media Audio and Video (Communication University of China)(中国传媒大学媒体音频与视频重点实验室) School of Intelligence Science and Technology, Nanjing University(南京大学智能科学与技术学院)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出DiGS-Avatar,将单图像可动画三维人体重建转化为UV空间扩散的潜变量补全任务,通过师生框架优化后解码为三维高斯基元,实现了高效高质量的重建与零样本泛化。

Comments ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.20308 2026-08-21 cs.CV 新提交 79%

DreamHand: Repurposing Video Diffusion Models for Occlusion-Robust Egocentric 3D Hand Motion Recovery

DreamHand:复用视频扩散模型实现遮挡鲁棒的第一人称视角3D手部运动恢复

Yufei Liu, Xixi Wang, Hao Li, Ganlong Zhao, Kaitong Cai, Chengkai Jin, Chunxiao Liu, Jianbo Liu, Siyuan Huang, Xingang Pan, Hongsheng Li

机构 * ACE Robotics(ACE机器人公司) Nanyang Technological University(南洋理工大学) The Chinese University of Hong Kong(香港中文大学) Shanghai Jiao Tong University(上海交通大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 DreamHand是复用视频扩散模型的离线片段级框架,通过确定性干净潜在编码器与双向时空解码器恢复带度量位置的连续双手轨迹,在五个第一人称视角基准测试中实现最佳性能,为机器人操作数据提供可扩展路径。

Comments Project Page: https://ggxxii.github.io/dreamhand/

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.19871 2026-08-21 cs.CV 新提交 79%

DIFFCZSL: Compositional Zero-Shot Learning Regularized by Diffusion Representations

DIFFCZSL:基于扩散表示正则化的组合零样本学习

Hangyu Tian, Zhenqi He, Yanghao Wang, Long Chen

机构 * The Hong Kong University of Science and Technology(香港科技大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 DIFFCZSL将预训练扩散模型的生成先验注入基于CLIP的组合零样本学习流程,通过对比对齐提升性能,在两类设置下均优于强基线,凸显了扩散表示与视觉-语言模型的互补优势。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.19644 2026-08-21 cs.CV 新提交 79%

When Guidance Goes Off-Scale: Recalibrating Diffusion Transformers under Analog Compute-in-Memory Nonidealities

当引导超出规模:模拟存算非理想性下的扩散变换器重新校准

Wenshuai Yao, Wenyong Zhou

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 该研究针对模拟存算非理想性下扩散变换器的CFG残差问题,提出采样器侧引导尺度重新校准方法,可大幅消除CIM导致的FID差距,提升生成质量。

Comments 9 pages, 8 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.19556 2026-08-21 cs.CV cs.AI 新提交 79%

Stream4D: 4D-Consistency for Streaming Autoregressive Diffusion Video Models

Stream4D:面向流式自回归扩散视频模型的4D一致性

Yuanhao Ban, Jiaqi Feng, Hengguang Zhou, Xiaohuan Pei, Justin Cui, Cho-Jui Hsieh

机构 * UCLA(加州大学洛杉矶分校) Tsinghua University(清华大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 Stream4D用显式建模场景动力学的前馈4D重建奖励替代静态评判器,结合运动先验与感知锚,提升流式自回归扩散视频模型的4D重建质量、运动保留效果及人类对齐偏好。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.22700 2026-08-21 cs.CV 版本更新 79%

4DLoG: Generative Modeling of Neurodegenerative Brain Anatomy with 4D Longitudinal Diffusion Model

利用4D纵向扩散模型生成神经退行性脑解剖结构

Nivetha Jayakumar, Swakshar Deb, Bahram Jafrasteh, Qingyu Zhao, Miaomiao Zhang

机构 * Department of Electrical and Computer Engineering(电气与计算机工程系) Department of Radiology(放射学系) Department of Electrical and Computer Engineering, Department of Computer Science(电气与计算机工程系、计算机科学系)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出一种4D纵向扩散模型,用于生成神经退行性疾病的脑解剖结构,通过临床变量条件生成准确的脑部变化轨迹,验证了其在疾病分类和分割中的有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.18400 2026-08-21 eess.IV cs.CV 版本更新 79%

Exploiting Completeness Perception with Diffusion Transformer for Unified 3D MRI Synthesis

利用扩散变换器的完整性感知实现统一的3D MRI合成

Junkai Liu, Nay Aung, Theodoros N. Arvanitis, Joao A. C. Lima, Steffen E. Petersen, Le Zhang

机构 * School of Engineering, University of Birmingham, UK(伯明翰大学工程学院) William Harvey Research Institute, Queen Mary University London, UK(女王玛丽大学伦敦威廉·哈里维研究所) Barts Heart Centre, St Bartholomew’s Hospital, Barts Health NHS Trust, UK(巴特勒心脏中心,圣巴塞洛缪医院,巴特勒健康 NHS信托) Division of Cardiology, Johns Hopkins University School of Medicine, US(约翰霍普金斯大学医学院心脏病科)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出CoPeDiT模型,通过完整性感知机制提升3D MRI合成的语义一致性与鲁棒性。

Comments Accepted to TMLR

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.15648 2026-08-21 cs.LG cs.CE cs.CV 版本更新 79%

Guided Diffusion by Optimized Loss Functions on Relaxed Parameters for Inverse Material Design

通过放松参数优化损失函数引导扩散用于逆材料设计

Jens U. Kreber, Christian Weißenfels, Joerg Stueckler

机构 * Intelligent Perception in Technical Systems Group, University of Augsburg, Germany(技术系统智能感知组,乌尔姆大学,德国) Faculty of Mathematics, Natural Science and Engineering, University of Augsburg, Germany(数学、自然科学与工程学院,乌尔姆大学,德国) Centre for Advanced Analytics and Predictive Sciences, University of Augsburg, Germany(高级分析与预测科学中心,乌尔姆大学,德国)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 通过优化损失函数引导扩散模型,用于逆材料设计,实现高效且多样化的设计生成。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05441 2026-08-21 cs.CV cs.LG 版本更新 79%

Decoupling High and Low Frequencies for Faithful Image Generation with Fine Details

解耦高低频以生成具有精细细节的忠实图像

Tejaswini Medi, Hsien-Yi Wang, Arianna Rampini, Margret Keuper

机构 * University of Mannheim(曼海姆大学) Autodesk AI Lab(Autodesk人工智能实验室) MPI for Informatics(信息研究所)

专题命中 扩散模型 :image generation(title);diffusion(abstract);分类 cs.CV

AI总结 该研究针对潜在生成模型难以恢复图像高频细节的问题,提出DeBaT解耦频带分词器,将其集成到潜在扩散模型后,可生成更锐利、更真实的图像样本。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.19000 2026-08-20 cs.CV 新提交 79%

Mise-en-Scène: Implicit Layout Emergence in Diffusion Transformers for Human-AI Design Co-Creation

场景布置(Mise-en-Scène):用于人机协同设计共创的扩散Transformer中隐式布局的涌现

Zipeng Xu, Ryan Murdock, Umberto Michieli

机构 * Canva Research(Canva研究院)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出Mise-en-Scène框架,通过微调扩散Transformer实现隐式布局涌现,结合匹配放置步骤保证素材保真度,在PrismLayersPlus基准上生成的设计感知质量显著优于现有方法。

Comments Best Paper Award at ECCV Human-AI Co-Creation Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.11748 2026-08-20 cs.CV 版本更新 79%

Dual Modality Prompted Diffusion Priors for Zero Shot Hyperspectral Pansharpening

用于零样本高光谱 pansharpening 的双模态提示扩散先验

Pengwei Xie, Fei Zhu, Jiajun Li, Xiangyuan Liu, Xiangyuan Liu, Kangqing Shen, Gemine Vivone

机构 * School of Artificial Intelligence, Beijing Normal University(北京师范大学人工智能学院) Peking University(北京大学) National Center for Applied Mathematics Shenzhen (NCAMS), Southern University of Science and Technology(南方科技大学深圳应用数学国家中心) Department of Automation, Tsinghua University(清华大学自动化系) National Research Council of Italy, Institute of Integrated Methodologies for Earth Observation (CNR-IMIOT)(意大利国家研究委员会综合地球观测方法研究所)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 该研究针对零样本高光谱 pansharpening 问题,提出双模态提示扩散模型 DIDM,通过双模态提示注入与全色引导正则化实现空间细节与光谱保真的平衡,在多数据集实验中取得最优性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.08289 2026-08-20 cs.CV cs.AI cs.LG 79%

Large Intestine 3D Shape Refinement Using Point Diffusion Models for Digital Phantom Generation

Kaouther Mouheb, Mobina Ghojogh Nejad, Lavsen Dahal, Ehsan Samei, Kyle J. Lafata, W. Paul Segars, Joseph Y. Lo

机构 * Center for Virtual Imaging Trials, Department of Radiology, Duke University School of Medicine, Durham, NC, USA Biomedical Imaging Group Rotterdam, Department of Radiology \& Nuclear Medicine, Erasmus MC, Rotterdam, the Netherlands Electrical Computer Engineering, Pratt School of Engineering, Duke University, Durham, NC, USA

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

Journal ref Shape in Medical Imaging (ShapeMI 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.21307 2026-08-20 cs.CV 版本更新 79%

The Linear Geometry of Interpretable Tokens: Jailbreaking Attacks and Defenses for Unlearned Diffusion Models

可解释令牌的线性几何:未学习扩散模型的越狱攻击与防御

Siyi Chen, Yimeng Zhang, Sijia Liu, Qing Qu

机构 * Department of Electrical Engineering & Computer Science, University of Michigan(电气工程与计算机科学系,密歇根大学) Department of Computer Science, Michigan State University(计算机科学系,密歇根州立大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本研究揭示未学习扩散模型中擦除的有害概念以线性子空间存在,提出SubAttack越狱攻击与SubDefense防御,实验表明其提升了对扩散未学习漏洞的理解与缓解效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.17973 2026-08-19 cs.CV 新提交 79%

LinCa: Accelerating Diffusion Models via Learnable Decomposed Feature Caching

LinCa:基于可学习分解特征缓存的扩散模型加速方法

Jinshan Liu, Haoran Qin, Xiaobing Tu, Jiacheng Liu, Jiahui Hu, Zhengan Yan, Yukun Xie, Kerui Shen, Jinkui Ren, Yuqi Lin, Xiantao Zhang, Linfeng Zhang

机构 * Shanghai Jiao Tong University(上海交通大学) Shandong University(山东大学) Terminal Intelligent Computing Division, Alibaba Cloud(阿里云终端智能计算事业部) South China University of Technology(华南理工大学) Jilin University(吉林大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出LinCa框架,通过可学习可逆网络分解缓存特征并差异化预测,仅需少量额外参数即可在5-7倍加速下使扩散模型保持近无损质量,性能优于现有方法。

Comments Accepted to ECCV 2026. 28 pages including appendix. Code: https://github.com/QHR69/LinCa

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.26747 2026-08-19 cs.CV cs.LG 版本更新 79%

From Diffusion to Flow: Efficient Motion Generation in MotionGPT3

从扩散到流:在MotionGPT3中实现高效的运动生成

Jaymin Bhan, JiHong Jeon, SangYeop Jeong

机构 * Department of Applied Artificial Intelligence(应用人工智能系)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文比较了扩散与校正流在MotionGPT3中的表现,发现流在训练速度、性能和效率上更具优势。

Comments ReALM-GEN Workshop ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.11534 2026-08-19 cs.CV 版本更新 79%

Risk-Controllable Multi-View Diffusion for Driving Scenario Generation

可控风险多视角扩散用于驾驶场景生成

Hongyi Lin, Wenxiu Shi, Heye Huang, Dingyi Zhuang, Song Zhang, Yang Liu, Xiaobo Qu, Jinhua Zhao

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 RiskMV-DPO通过整合风险水平与物理风险建模,实现可控风险的多视角驾驶场景生成,提升3D检测性能并减少FID,推动安全导向的具身智能发展。

Comments 10 pages, 4 figures; accepted at the CVPR 2026 Workshop on Video Generative Models: Benchmarks and Evaluation (VGBE). Updated to the complete camera-ready version

Journal ref Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.00239 2026-08-19 eess.IV cs.CV cs.LG 79%

IgCONDA-PET: Weakly-Supervised PET Anomaly Detection using Implicitly-Guided Attention-Conditional Counterfactual Diffusion Modeling -- a Multi-Center, Multi-Cancer, and Multi-Tracer Study

Shadab Ahamed, Arman Rahmim

机构 * organization= Department of Physics \& Astronomy, University of British Columbia , city= Vancouver , state= BC , country= Canada organization= Department of Integrative Oncology, BC Cancer Research Institute , city= Vancouver , state= BC , country= Canada organization= Department of Radiology, University of British Columbia , city= Vancouver , state= BC , country= Canada

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

Comments 48 pages, 13 figures, 4 tables

Journal ref Computerized Medical Imaging and Graphics, 124, 102615 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.16786 2026-08-18 cs.CV 新提交 79%

Revisiting Classifier-Free Guidance Methods in Latent Diffusion Models

重新审视潜在扩散模型中无分类器引导方法

Artem Sergievskii, Artyom Turevich, Sergey Kastryulin

机构 * HSE University(高等经济大学) Yandex(Yandex公司)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文重新评估了八种源于无分类器引导(CFG)的无训练技术,发现无一种方法始终优于CFG,APG仅获名义最佳分数,注意力扰动方法在不同模型上表现有差异,CFG仍是低成本竞争基线。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.15343 2026-08-18 cs.CV 新提交 79%

Feed-Forward Hierarchical Gaussian Diffusion for Extreme CT Reconstruction

用于极端CT重建的前馈分层高斯扩散

Yuezhe Yang, Li Cheng

机构 * University of Alberta(阿尔伯塔大学) University of Sydney(悉尼大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 针对极端CT重建的不适定问题,提出HiGDiff前馈分层高斯扩散框架,通过结构与细节两阶段扩散实现最先进性能,在LDCT-PD数据集上PSNR提升5.81dB、SSIM提升0.113。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.10162 2026-08-18 cs.CV 版本更新 79%

MAD-HOI: Masked Autoregressive Diffusion for Generating Articulated Hand Object Interactions from Text

MAD-HOI:用于从文本生成关节手-物体交互的掩码自回归扩散模型

Ananya Bal, Kartik Sharma, Ethan Lai, Samyak Tiwari, Liza Dahiya, Chaitanya Chawla, Laszlo A. Jeni

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 MAD-HOI是一种结合掩码自回归与扩散的HOI生成模型,可生成多样且符合物理规律的手-物体交互序列,支持可变长度生成、运动补全填充等功能,在ARCTIC和GRAB数据集上优于开源基线。

Comments 17 pages, 9 figures, 8 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.18736 2026-08-18 cs.CV 版本更新 79%

Spectral Progressive Diffusion for Efficient Image and Video Generation

频域渐进扩散用于高效图像和视频生成

Howard Xiao, Brian Chao, Lior Yariv, Gordon Wetzstein

机构 * Stanford University(斯坦福大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出了一种频域渐进扩散框架,通过在预训练扩散模型的去噪轨迹中逐步提高分辨率,实现高效的图像和视频生成,同时改进了效率和质量。

Comments Project website at https://howardxiao.ca/speed

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06223 2026-08-18 cs.CV 79%

RedDiffuser: Auditing Multimodal Safety Failures in Vision-Language Models via Reinforced Diffusion

RedDiffuser:通过强化扩散审计多模态安全故障

Ruofan Wang, Xingjun Ma

机构 * Fudan University(复旦大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 研究多模态系统在有害上下文暴露下的安全审计,提出RedDiffuser框架通过扩散模型生成视觉输入,揭示隐藏的安全漏洞,实验显示VLMs在部分有毒文本与视觉上下文结合时存在广泛安全问题。

详情

展开后加载摘要…

URL PDF HTML 收藏