arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

共收录 70082 信号源:cs.CV, cs.GR, cs.MM

1. 扩散模型 70082 篇

2604.22700 2026-08-21 cs.CV 版本更新 79%

4DLoG: Generative Modeling of Neurodegenerative Brain Anatomy with 4D Longitudinal Diffusion Model

利用4D纵向扩散模型生成神经退行性脑解剖结构

Nivetha Jayakumar, Swakshar Deb, Bahram Jafrasteh, Qingyu Zhao, Miaomiao Zhang

机构 * Department of Electrical and Computer Engineering(电气与计算机工程系) Department of Radiology(放射学系) Department of Electrical and Computer Engineering, Department of Computer Science(电气与计算机工程系、计算机科学系)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出一种4D纵向扩散模型,用于生成神经退行性疾病的脑解剖结构,通过临床变量条件生成准确的脑部变化轨迹,验证了其在疾病分类和分割中的有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.18400 2026-08-21 eess.IV cs.CV 版本更新 79%

Exploiting Completeness Perception with Diffusion Transformer for Unified 3D MRI Synthesis

利用扩散变换器的完整性感知实现统一的3D MRI合成

Junkai Liu, Nay Aung, Theodoros N. Arvanitis, Joao A. C. Lima, Steffen E. Petersen, Le Zhang

机构 * School of Engineering, University of Birmingham, UK(伯明翰大学工程学院) William Harvey Research Institute, Queen Mary University London, UK(女王玛丽大学伦敦威廉·哈里维研究所) Barts Heart Centre, St Bartholomew’s Hospital, Barts Health NHS Trust, UK(巴特勒心脏中心,圣巴塞洛缪医院,巴特勒健康 NHS信托) Division of Cardiology, Johns Hopkins University School of Medicine, US(约翰霍普金斯大学医学院心脏病科)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出CoPeDiT模型,通过完整性感知机制提升3D MRI合成的语义一致性与鲁棒性。

Comments Accepted to TMLR

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.15648 2026-08-21 cs.LG cs.CE cs.CV 版本更新 79%

Guided Diffusion by Optimized Loss Functions on Relaxed Parameters for Inverse Material Design

通过放松参数优化损失函数引导扩散用于逆材料设计

Jens U. Kreber, Christian Weißenfels, Joerg Stueckler

机构 * Intelligent Perception in Technical Systems Group, University of Augsburg, Germany(技术系统智能感知组,乌尔姆大学,德国) Faculty of Mathematics, Natural Science and Engineering, University of Augsburg, Germany(数学、自然科学与工程学院,乌尔姆大学,德国) Centre for Advanced Analytics and Predictive Sciences, University of Augsburg, Germany(高级分析与预测科学中心,乌尔姆大学,德国)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 通过优化损失函数引导扩散模型,用于逆材料设计,实现高效且多样化的设计生成。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05441 2026-08-21 cs.CV cs.LG 版本更新 79%

Decoupling High and Low Frequencies for Faithful Image Generation with Fine Details

解耦高低频以生成具有精细细节的忠实图像

Tejaswini Medi, Hsien-Yi Wang, Arianna Rampini, Margret Keuper

机构 * University of Mannheim(曼海姆大学) Autodesk AI Lab(Autodesk人工智能实验室) MPI for Informatics(信息研究所)

专题命中 扩散模型 :image generation(title);diffusion(abstract);分类 cs.CV

AI总结 该研究针对潜在生成模型难以恢复图像高频细节的问题,提出DeBaT解耦频带分词器,将其集成到潜在扩散模型后,可生成更锐利、更真实的图像样本。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.19000 2026-08-20 cs.CV 新提交 79%

Mise-en-Scène: Implicit Layout Emergence in Diffusion Transformers for Human-AI Design Co-Creation

场景布置(Mise-en-Scène):用于人机协同设计共创的扩散Transformer中隐式布局的涌现

Zipeng Xu, Ryan Murdock, Umberto Michieli

机构 * Canva Research(Canva研究院)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出Mise-en-Scène框架,通过微调扩散Transformer实现隐式布局涌现,结合匹配放置步骤保证素材保真度,在PrismLayersPlus基准上生成的设计感知质量显著优于现有方法。

Comments Best Paper Award at ECCV Human-AI Co-Creation Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.11748 2026-08-20 cs.CV 版本更新 79%

Dual Modality Prompted Diffusion Priors for Zero Shot Hyperspectral Pansharpening

用于零样本高光谱 pansharpening 的双模态提示扩散先验

Pengwei Xie, Fei Zhu, Jiajun Li, Xiangyuan Liu, Xiangyuan Liu, Kangqing Shen, Gemine Vivone

机构 * School of Artificial Intelligence, Beijing Normal University(北京师范大学人工智能学院) Peking University(北京大学) National Center for Applied Mathematics Shenzhen (NCAMS), Southern University of Science and Technology(南方科技大学深圳应用数学国家中心) Department of Automation, Tsinghua University(清华大学自动化系) National Research Council of Italy, Institute of Integrated Methodologies for Earth Observation (CNR-IMIOT)(意大利国家研究委员会综合地球观测方法研究所)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 该研究针对零样本高光谱 pansharpening 问题,提出双模态提示扩散模型 DIDM,通过双模态提示注入与全色引导正则化实现空间细节与光谱保真的平衡,在多数据集实验中取得最优性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.08289 2026-08-20 cs.CV cs.AI cs.LG 79%

Large Intestine 3D Shape Refinement Using Point Diffusion Models for Digital Phantom Generation

Kaouther Mouheb, Mobina Ghojogh Nejad, Lavsen Dahal, Ehsan Samei, Kyle J. Lafata, W. Paul Segars, Joseph Y. Lo

机构 * Center for Virtual Imaging Trials, Department of Radiology, Duke University School of Medicine, Durham, NC, USA Biomedical Imaging Group Rotterdam, Department of Radiology \& Nuclear Medicine, Erasmus MC, Rotterdam, the Netherlands Electrical Computer Engineering, Pratt School of Engineering, Duke University, Durham, NC, USA

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

Journal ref Shape in Medical Imaging (ShapeMI 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.21307 2026-08-20 cs.CV 版本更新 79%

The Linear Geometry of Interpretable Tokens: Jailbreaking Attacks and Defenses for Unlearned Diffusion Models

可解释令牌的线性几何:未学习扩散模型的越狱攻击与防御

Siyi Chen, Yimeng Zhang, Sijia Liu, Qing Qu

机构 * Department of Electrical Engineering & Computer Science, University of Michigan(电气工程与计算机科学系,密歇根大学) Department of Computer Science, Michigan State University(计算机科学系,密歇根州立大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本研究揭示未学习扩散模型中擦除的有害概念以线性子空间存在,提出SubAttack越狱攻击与SubDefense防御,实验表明其提升了对扩散未学习漏洞的理解与缓解效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.17973 2026-08-19 cs.CV 新提交 79%

LinCa: Accelerating Diffusion Models via Learnable Decomposed Feature Caching

LinCa:基于可学习分解特征缓存的扩散模型加速方法

Jinshan Liu, Haoran Qin, Xiaobing Tu, Jiacheng Liu, Jiahui Hu, Zhengan Yan, Yukun Xie, Kerui Shen, Jinkui Ren, Yuqi Lin, Xiantao Zhang, Linfeng Zhang

机构 * Shanghai Jiao Tong University(上海交通大学) Shandong University(山东大学) Terminal Intelligent Computing Division, Alibaba Cloud(阿里云终端智能计算事业部) South China University of Technology(华南理工大学) Jilin University(吉林大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出LinCa框架,通过可学习可逆网络分解缓存特征并差异化预测,仅需少量额外参数即可在5-7倍加速下使扩散模型保持近无损质量,性能优于现有方法。

Comments Accepted to ECCV 2026. 28 pages including appendix. Code: https://github.com/QHR69/LinCa

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.26747 2026-08-19 cs.CV cs.LG 版本更新 79%

From Diffusion to Flow: Efficient Motion Generation in MotionGPT3

从扩散到流:在MotionGPT3中实现高效的运动生成

Jaymin Bhan, JiHong Jeon, SangYeop Jeong

机构 * Department of Applied Artificial Intelligence(应用人工智能系)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文比较了扩散与校正流在MotionGPT3中的表现,发现流在训练速度、性能和效率上更具优势。

Comments ReALM-GEN Workshop ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.11534 2026-08-19 cs.CV 版本更新 79%

Risk-Controllable Multi-View Diffusion for Driving Scenario Generation

可控风险多视角扩散用于驾驶场景生成

Hongyi Lin, Wenxiu Shi, Heye Huang, Dingyi Zhuang, Song Zhang, Yang Liu, Xiaobo Qu, Jinhua Zhao

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 RiskMV-DPO通过整合风险水平与物理风险建模,实现可控风险的多视角驾驶场景生成,提升3D检测性能并减少FID,推动安全导向的具身智能发展。

Comments 10 pages, 4 figures; accepted at the CVPR 2026 Workshop on Video Generative Models: Benchmarks and Evaluation (VGBE). Updated to the complete camera-ready version

Journal ref Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.00239 2026-08-19 eess.IV cs.CV cs.LG 79%

IgCONDA-PET: Weakly-Supervised PET Anomaly Detection using Implicitly-Guided Attention-Conditional Counterfactual Diffusion Modeling -- a Multi-Center, Multi-Cancer, and Multi-Tracer Study

Shadab Ahamed, Arman Rahmim

机构 * organization= Department of Physics \& Astronomy, University of British Columbia , city= Vancouver , state= BC , country= Canada organization= Department of Integrative Oncology, BC Cancer Research Institute , city= Vancouver , state= BC , country= Canada organization= Department of Radiology, University of British Columbia , city= Vancouver , state= BC , country= Canada

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

Comments 48 pages, 13 figures, 4 tables

Journal ref Computerized Medical Imaging and Graphics, 124, 102615 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.16786 2026-08-18 cs.CV 新提交 79%

Revisiting Classifier-Free Guidance Methods in Latent Diffusion Models

重新审视潜在扩散模型中无分类器引导方法

Artem Sergievskii, Artyom Turevich, Sergey Kastryulin

机构 * HSE University(高等经济大学) Yandex(Yandex公司)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文重新评估了八种源于无分类器引导(CFG)的无训练技术,发现无一种方法始终优于CFG,APG仅获名义最佳分数,注意力扰动方法在不同模型上表现有差异,CFG仍是低成本竞争基线。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.15343 2026-08-18 cs.CV 新提交 79%

Feed-Forward Hierarchical Gaussian Diffusion for Extreme CT Reconstruction

用于极端CT重建的前馈分层高斯扩散

Yuezhe Yang, Li Cheng

机构 * University of Alberta(阿尔伯塔大学) University of Sydney(悉尼大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 针对极端CT重建的不适定问题,提出HiGDiff前馈分层高斯扩散框架,通过结构与细节两阶段扩散实现最先进性能,在LDCT-PD数据集上PSNR提升5.81dB、SSIM提升0.113。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.10162 2026-08-18 cs.CV 版本更新 79%

MAD-HOI: Masked Autoregressive Diffusion for Generating Articulated Hand Object Interactions from Text

MAD-HOI:用于从文本生成关节手-物体交互的掩码自回归扩散模型

Ananya Bal, Kartik Sharma, Ethan Lai, Samyak Tiwari, Liza Dahiya, Chaitanya Chawla, Laszlo A. Jeni

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 MAD-HOI是一种结合掩码自回归与扩散的HOI生成模型,可生成多样且符合物理规律的手-物体交互序列,支持可变长度生成、运动补全填充等功能,在ARCTIC和GRAB数据集上优于开源基线。

Comments 17 pages, 9 figures, 8 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.18736 2026-08-18 cs.CV 版本更新 79%

Spectral Progressive Diffusion for Efficient Image and Video Generation

频域渐进扩散用于高效图像和视频生成

Howard Xiao, Brian Chao, Lior Yariv, Gordon Wetzstein

机构 * Stanford University(斯坦福大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出了一种频域渐进扩散框架,通过在预训练扩散模型的去噪轨迹中逐步提高分辨率,实现高效的图像和视频生成,同时改进了效率和质量。

Comments Project website at https://howardxiao.ca/speed

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06223 2026-08-18 cs.CV 79%

RedDiffuser: Auditing Multimodal Safety Failures in Vision-Language Models via Reinforced Diffusion

RedDiffuser:通过强化扩散审计多模态安全故障

Ruofan Wang, Xingjun Ma

机构 * Fudan University(复旦大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 研究多模态系统在有害上下文暴露下的安全审计,提出RedDiffuser框架通过扩散模型生成视觉输入,揭示隐藏的安全漏洞,实验显示VLMs在部分有毒文本与视觉上下文结合时存在广泛安全问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.17555 2026-08-18 cs.CV cs.AI 版本更新 79%

FrescoDiffusion: 4K Image-to-Video with Prior-Regularized Tiled Diffusion

FrescoDiffusion: 4K图像到视频的先验正则化拼接扩散

Hugo Caselles-Dupré, Mathis Koroglu, Guillaume Jeanneret, Arnaud Dapogny, Matthieu Cord

机构 * Obvious Research, Paris, France Institute of Intelligent Systems(Obvious Research,巴黎,法国智能系统研究所) Robotics - Sorbonne University, Paris, France(机器人学 - 索邦大学,巴黎,法国)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 FrescoDiffusion通过先验正则化拼接扩散方法,实现从单张复杂图像生成4K图像到视频,提升全局一致性与细节保真度。

Comments 5 authors. Hugo Caselles-Dupré, Mathis Koroglu, and Guillaume Jeanneret contributed equally. 15 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10076 2026-08-18 cs.CV 版本更新 79%

Mitigating Error Accumulation in Co-Speech Motion Generation via Global Rotation Diffusion and Multi-Level Constraints

通过全局旋转扩散和多级约束缓解语音同步运动生成中的误差累积

Xiangyue Zhang, Jianfang Li, Jianqiang Ren, Jiaxu Zhang

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出GlobalDiff,通过全局旋转扩散和多级约束缓解语音同步运动生成中的误差累积,提升生成运动的平滑度和准确性。

Comments AAAI 2026. Project page: https://xiangyuezhang.com/GlobalDiff/; code and pretrained models: https://github.com/Xiangyue-Zhang/GlobalDiff

Journal ref Proceedings of the AAAI Conference on Artificial Intelligence, 40(15), 12834-12842, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.18911 2026-08-18 cs.LG cs.AI cs.CV 版本更新 79%

Rethinking Token-wise Feature Caching: Accelerating Diffusion Transformers with Dual Feature Caching

重新思考逐令牌特征缓存:通过双特征缓存加速扩散Transformer

Chang Zou, Shikang Zheng, Evelyn Zhang, Runlin Guo, Haohang Xu, Zhengyi Shi, Conghui He, Xuming Hu, Linfeng Zhang

机构 * Shanghai Jiao Tong University(上海交通大学) South China University of Technology(南方科技大学) Beihang University(北航) Huawei Technologies Ltd(华为技术有限公司) Shanghai AI Lab(上海人工智能实验室) Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文针对逐令牌特征缓存的核心问题提出质疑,提出双特征缓存方法DuCa,在DiT等生成模型上实现了优于现有逐令牌特征缓存的性能。

Journal ref IEEE Transactions on Image Processing, vol. 35, pp. 6211-6220, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.14430 2026-08-17 cs.LG cs.CV stat.ML 新提交 79%

Designing Reinforcement Learning for Diffusion Models: A Unified Path-Space View

为扩散模型设计强化学习:统一的路径空间视角

Yixian Xu, Yuanrui Zhang, Shengjie Luo, Liwei Wang, Di He

机构 * State Key Laboratory of General Artificial Intelligence(通用人工智能国家重点实验室) School of Intelligence Science and Technology, Peking University(北京大学智能科学与技术学院) ByteDance(字节跳动)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文从路径空间视角统一了扩散模型RL算法的原理,推导得到降方差值梯度形式,提出多样本KDE估计器和尺度受限权重族,在SD3.5-M等模型上验证了方法有效性并优于基线。

Comments 29 pages, 9 figures, 4 tables; work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.13669 2026-08-17 cs.CV 新提交 79%

Multiphase-Diff: Diffusion-Based Generative Modeling for High-Contrast Multiphase Physical Systems with Sharp Interfaces

Multiphase-Diff:面向具有尖锐界面的高对比度多相物理系统的基于扩散的生成建模

Yining Huang, Zhenyu Liang

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 针对高对比度多相物理系统的扩散生成建模难题,提出Multiphase-Diff方法,通过三项改进实现更优的物理与分布保真度,在多相基准上优于7个基线模型,适用于该场景的科学样本生成。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.09231 2026-08-17 cs.CV 版本更新 79%

BAG: Budget-Aware Gating for Diffusion Caching

BAG:面向扩散缓存的预算感知门控机制

Tong Zhao, Mingkun Lei, Yucheng Han, Chi Zhang

机构 * Westlake AGI Lab(西湖AGI实验室)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 提出BAG(预算感知门控机制),通过离线到在线调度蒸馏训练轻量门控网络,在FLUX.1-dev和Wan2.1上实现优于现有最优缓存方法的性能。

Comments 23 pages, 13 figures, and 14 tables. Code Link: see AGI-Lab/BAG" target="_blank" rel="noopener">https://github.com/Westlake-AGI-Lab/BAG

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.03265 2026-08-17 cs.CV 79%

Optimizing for the Shortest Path in Denoising Diffusion Model

Ping Chen, Xingpeng Zhang, Zhaoxiang Liu, Huan Hu, Xiang Liu, Kai Wang, Min Wang, Yanlin Qian, Shiguo Lian

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

Comments Accepet by CVPR 2025 (10 pages, 6 figures)

Journal ref Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025, pp. 18021-18030

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.13255 2026-08-14 cs.CV cs.AI 新提交 79%

GeoCache: Training-Free Acceleration of Multi-View Texture Diffusion via Geometric Delta Transport

GeoCache:通过几何增量传输实现多视图纹理扩散的无训练加速

Haotang Li, Zhenyu Qi, Shaohan Henry Wang, Kebin Peng, Yutong Zhao, Zi Wang, Bo Liu, Huanrui Yang, Sen He

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本研究提出无训练插件GeoCache,通过传输锚点视图的几何对齐更新实现多视图纹理扩散加速,在三个基准数据集上实现优于时间缓存和步骤缩减的速度-保真度权衡,2倍以上加速时表现突出。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.01677 2026-08-14 cs.CV cs.LG 版本更新 79%

Generative Brownian Bridge Diffusion In Motion Space For Enhanced Myocardial Strain Analysis

用于增强心肌应变分析的运动空间生成布朗桥扩散模型

Rishov Paul, Frederick H. Epstein, Miaomiao Zhang

机构 * University of Virginia(弗吉尼亚大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 该研究提出运动空间生成布朗桥扩散模型,以标准CMR图像为条件学习运动映射,在多中心CMR数据集上验证其可提升应变分析准确性,为开发临床可用的低成本心脏评估AI工具提供新范式。

Comments 12 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17492 2026-08-14 cs.CV 版本更新 79%

EvDiff: Event-Based Video Reconstruction using One-Step Diffusion Models

EvDiff:高质视频与事件相机

Weilun Li, Lei Sun, Ruixi Gao, Qi Jiang, Yuqin Ma, Kaiwei Wang, Ming-Hsuan Yang, Luc Van Gool, Danda Pani Paudel

机构 * Zhejiang University(浙江大学) INSAIT UC Merced(加州大学默塞德分校) Google DeepMind(谷歌DeepMind)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 EvDiff通过基于事件的扩散模型和替代训练框架,从单色事件流生成高质量彩色视频,提升视频生成的保真度和现实感。

Comments Replacement note: This manuscript has been transferred from the CVPR format to the ECCV 2026 format, with the corresponding title and template updated accordingly. The technical content remains largely unchanged from the previous version. (Current version: 21 pages, 6 figures, and 3 tables.) Accepted by ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.12308 2026-08-13 cs.CV cs.AI 新提交 79%

DreamFly: Causal Memory and Receding-Horizon Diffusion Planning for Aerial Vision-Language Navigation

DreamFly:面向空中视觉语言导航的因果记忆与后退时域扩散规划

Yan Deng, Fei Xu

机构 * School of Electronic Information Engineering, Xi’an Technological University(西安工业大学电子信息工程学院) School of Computer Science and Engineering, Xi’an Technological University(西安工业大学计算机科学与工程学院)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 DreamFly是基于Dream-VLA的扩散式空中VLN框架,通过因果记忆、后退时域扩散规划和LiteStop解耦终止,在OpenFly基准上显著优于对比方法。

Comments 24 pages, 6 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.12175 2026-08-13 cs.CV 新提交 79%

TGRHuman: Text-Guided Realistic 3D Human Generation via Diffusion Renderer

TGRHuman:基于扩散渲染器的文本引导逼真3D人体生成

Muxin Zhang, Chaohui Yu, Yuanwang Yang, Min Wei, Zhuo Su, Kun Li

机构 * Tianjin University(天津大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本研究提出TGRHuman方法,解耦3D人体的几何与纹理生成,采用显式多视图优化结合扩散渲染器,实现高效高质量文本引导的逼真3D人体生成,性能优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.12155 2026-08-13 cs.CV 新提交 79%

Understanding Why Foundation Models Work for Diffusion-Generated Image Detection

探究基础模型为何适用于扩散生成图像检测

Davide Cozzolino, Giovanni Poggi, Luisa Verdoliva

机构 * University Federico II of Naples(那不勒斯费德里科二世大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本研究探究基础模型用于扩散生成图像检测的原因,通过DDIM反演、频率交换及潜在空间分析,发现其利用中低频分布差异实现检测,为该类方法的可解释性提供了新方向。

详情

展开后加载摘要…

URL PDF HTML 收藏