arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

2025-12-04 至 2025-12-04 共收录 45 信号源:cs.CV, cs.GR, cs.MM

1. 扩散模型 45 篇

2512.03247 2025-12-04 cs.CV 85%

PixPerfect: Seamless Latent Diffusion Local Editing with Discriminative Pixel-Space Refinement

PixPerfect: 基于判别像素空间的无缝潜在扩散局部编辑

Haitian Zheng, Yuan Yao, Yongsheng Yu, Yuqian Zhou, Jiebo Luo, Zhe Lin

机构 * Adobe Research(Adobe研究院) University of Rochester(罗切斯特大学)

专题命中 扩散模型 :diffusion(title,abstract);image editing(abstract);inpainting(abstract);分类 cs.CV

AI总结 PixPerfect通过判别像素空间和伪影模拟流程,实现跨不同LDM架构和任务的无缝高保真局部编辑。

Comments Published in the Thirty-ninth Annual Conference on Neural Information Processing Systems (NeurIPS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.11056 2025-12-04 cs.CV 83%

Flow to the Mode: Mode-Seeking Diffusion Autoencoders for State-of-the-Art Image Tokenization

流到模式:用于最新图像标记化的模式寻求扩散自编码器

Kyle Sargent, Kyle Hsu, Justin Johnson, Li Fei-Fei, Jiajun Wu

机构 * Stanford University(斯坦福大学) University of Michigan(密歇根大学)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV

AI总结 FlowMo是一种基于Transformer的扩散自编码器,通过模式匹配和模式寻求阶段实现图像标记化的新SOTA,无需卷积、对抗损失等。

Comments ICCV 2025, 19 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.06424 2025-12-04 cs.CV 83%

Margin-aware Preference Optimization for Aligning Diffusion Models without Reference

基于边界的偏好优化:无需参考的扩散模型对齐

Jiwoo Hong, Sayak Paul, Noah Lee, Kashif Rasul, James Thorne, Jongheon Jeong

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);分类 cs.CV

AI总结 本文提出MaPO,一种无需参考的扩散模型对齐方法,通过优化偏好输出与非偏好输出之间的边界,提升T2I任务的适应性能,减少训练时间并优于现有方法。

Comments Accepted to AAAI 2026 Main Technical Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03056 2025-12-04 cs.LG cs.AI 82%

Delta Sampling: Data-Free Knowledge Transfer Across Diffusion Models

Delta Sampling: 数据无源的知识迁移跨扩散模型

Zhidong Gao, Zimeng Pan, Yuhang Yao, Chenyue Xie, Wei Wei

机构 * Shanxi University(山西大学) Google Cloud(谷歌云) Carnegie Mellon University(卡内基梅隆大学) University of Science and Technology of China(中国科学技术大学)

专题命中 扩散模型 :diffusion(title,abstract);image synthesis(abstract)

AI总结 Delta Sampling通过推理时利用模型预测差异实现跨不同架构基模型的知识迁移,无需原始训练数据。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03566 2025-12-04 cs.CV cs.MM 81%

GAOT: Generating Articulated Objects Through Text-Guided Diffusion Models

通过文本引导的扩散模型生成拟人化物体:GAOT

Hao Sun, Lei Fan, Donglin Di, Shaohui Liu

机构 * Harbin Institute of Technology(哈尔滨工业大学) University of New South Wales(新南威尔士大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV、cs.MM

AI总结 GAOT通过文本引导的扩散模型和超图学习,实现从文本提示到拟人化物体的生成,取得优于现有方法的性能。

Comments Accepted by ACM MM Asia2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03673 2025-12-04 cs.CV 79%

ConvRot: Rotation-Based Plug-and-Play 4-bit Quantization for Diffusion Transformers

ConvRot:基于旋转的插拔式4位量化用于扩散变换器

Feice Huang, Zuliang Han, Xing Zhou, Yihuang Chen, Lifei Zhu, Haoqian Wang

机构 * SIGS, Tsinghua University(清华大学信号与信息系) Central Media Technology Institute, Huawei(华为中央媒体技术研究所)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 ConvRot提出基于旋转的4位量化方法,通过正则Hadamard变换抑制异常值,实现扩散变换器的插拔式W4A4推理,提升速度和内存效率同时保持图像质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19499 2025-12-04 cs.CV cs.AI 79%

Sat2Flow: A Structure-Aware Diffusion Framework for Human Flow Generation from Satellite Imagery

Sat2Flow: 一种结构感知的扩散框架,用于从卫星图像生成人类流量

Xiangxu Wang, Tianhong Zhao, Wei Tu, Bowen Zhang, Guanzhou Chen, Jinzhou Cao

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 Sat2Flow通过结构感知的扩散框架,利用卫星图像生成结构一致的OD流量,解决了现有方法对辅助数据依赖和空间拓扑变化敏感的问题。

Comments 9 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.14430 2025-12-04 cs.CV cs.AI cs.PF 79%

PipeFusion: Patch-level Pipeline Parallelism for Diffusion Transformers Inference

PipeFusion: 基于扩散变换器推理的片级流水线并行

Jiarui Fang, Jinzhe Pan, Aoyu Li, Xibo Sun, Jiannan Wang

机构 * ByteDance(字节跳动) HUST(华中科技大学) Tencent(腾讯) The University of Hong Kong(香港大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 PipeFusion通过片级流水线并行方法,提高扩散变换器推理的效率和性能,适用于大模型如Flux.1。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03451 2025-12-04 cs.CV cs.AI cs.LG 79%

GalaxyDiT: Efficient Video Generation with Guidance Alignment and Adaptive Proxy in Diffusion Transformers

GalaxyDiT:通过引导对齐和自适应代理实现高效的视频生成

Zhiye Song, Steve Dai, Ben Keller, Brucek Khailany

机构 * Massachusetts Institute of Technology(麻省理工学院) NVIDIA(英伟达)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 GalaxyDiT通过引导对齐和自适应代理选择,提升视频生成效率,实现高达2.37倍的速度提升并保持高质量输出

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03450 2025-12-04 cs.CV cs.LG 79%

KeyPointDiffuser: Unsupervised 3D Keypoint Learning via Latent Diffusion Models

KeyPointDiffuser: 通过潜在扩散模型实现无监督的3D关键点学习

Rhys Newbury, Juyan Zhang, Tin Tran, Hanna Kurniawati, Dana Kulić

机构 * Monash University(蒙纳士大学) Australian National University(澳大利亚国立大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 KeyPointDiffuser通过潜在扩散模型实现无监督3D关键点学习,提升关键点一致性6个百分点。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03430 2025-12-04 cs.CV 79%

Label-Efficient Hyperspectral Image Classification via Spectral FiLM Modulation of Low-Level Pretrained Diffusion Features

通过低级预训练扩散特征的光谱FiLM调制实现标签高效的超光谱图像分类

Yuzhen Hu, Biplab Banerjee, Saurabh Prasad

机构 * University of Houston, Texas, USA(德克萨斯大学) Indian Institute of Technology Bombay, Mumbai, India(印度班加罗尔理工学院)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出了一种基于预训练扩散模型的标签高效超光谱图像分类方法,通过光谱FiLM调制融合空间和光谱信息,提升分类性能。

Comments Accepted to the ICML 2025 TerraBytes Workshop (June 9, 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03317 2025-12-04 cs.CV cs.AI cs.LG cs.RO 79%

NavMapFusion: Diffusion-based Fusion of Navigation Maps for Online Vectorized HD Map Construction

NavMapFusion: 基于扩散的导航地图融合用于在线向量化的高精度地图构建

Thomas Monninger, Zihan Zhang, Steffen Staab, Sihao Ding

机构 * Mercedes-Benz Research & Development North America, USA(梅赛德斯-奔驰北美研究与开发) University of Stuttgart, Germany(斯图加特大学) University of California, San Diego, USA(加州大学圣地亚哥分校) University of Southampton, United Kingdom(南安普顿大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 NavMapFusion通过结合低保真先验与高保真传感器数据,利用扩散模型实现在线高精度地图构建,提升环境表示的准确性和实时性。

Comments Accepted to 2026 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03813 2025-12-04 math.DS 78%

Hopf bifurcations in a reaction-diffusion model with a general advection term and delay effect

具有通用对流项和延迟效应的反应扩散模型中的Hopf分岔

Jingxiao Song, Chengwei Ren, Shaofen Zou

专题命中 扩散模型 :diffusion(title,abstract)

AI总结 本文研究了具有通用对流项和延迟效应的反应扩散模型,通过分析Hopf分岔的方向和稳定性,验证了模型在种群动态中的有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03759 2025-12-04 cs.CL cs.AI cs.LG 78%

Principled RL for Diffusion LLMs Emerges from a Sequence-Level Perspective

从序列层面视角出发的扩散大语言模型原理化强化学习

Jingyang Ou, Jiaqi Han, Minkai Xu, Shaoxuan Xu, Jianwen Xie, Stefano Ermon, Yi Wu, Chongxuan Li

机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学人工智能学院) Beijing Key Laboratory of Research on Large Models and Intelligent Governance(北京大模型与智能治理研究重点实验室) Engineering Research Center of Next-Generation Intelligent Search and Recommendation, MOE(下一代智能搜索与推荐工程技术研究中心) Stanford University(斯坦福大学) Lambda, Inc(Lambda公司) Tsinghua University(清华大学)

专题命中 扩散模型 :diffusion(title,abstract)

AI总结 本文提出ESPO,一种基于序列级优化的原理化强化学习框架,用于提升扩散大语言模型的生成能力,通过ELBO作为似然代理,显著优于token级方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.08873 2025-12-04 cond-mat.soft physics.chem-ph 78%

Coherent X-rays reveal anomalous molecular diffusion and cage effects in crowded protein solutions

相干X射线揭示拥挤蛋白质溶液中异常分子扩散和笼效应

Anita Girelli, Maddalena Bin, Mariia Filianina, Michelle Dargasz, Nimmi Das Anthuparambil, Johannes Möller, Alexey Zozulya, Iason Andronis, Sonja Timmermann, Sharon Berkowicz, Sebastian Retzbach, Mario Reiser, Agha Mohammad Raza, Marvin Kowalski, Mohammad Sayed Akhundzadeh, Jenny Schrage, Chang Hee Woo, Maximilian D. Senft, Lara Franziska Reichart, Aliaksandr Leonau, Prince Prabhu Rajaiah, William Chèvremont, Tilo Seydel, Jörg Hallmann, Angel Rodriguez-Fernandez, Jan-Etienne Pudell, Felix Brausse, Ulrike Boesenberg, James Wrigley, Mohamed Youssef, Wei Lu, Wonhyuk Jo, Roman Shayduk, Trey Guest, Anders Madsen, Felix Lehmkühler, Michael Paulus, Fajun Zhang, Frank Schreiber, Christian Gutt, Fivos Perakis

专题命中 扩散模型 :diffusion(title,abstract)

AI总结 利用MHz-XPCS研究铁蛋白在拥挤环境中的异常扩散及笼效应,通过δγ理论建模揭示蛋白质在细胞质中的复杂运动机制。

Comments This version of the article has been accepted for publication after peer review, but is not the Version of Record and does not reflect post-acceptance improvements or any corrections. The Version of Record is available online at: https://doi.org/10.1038/s41467-025-66972-6

Journal ref Nature Communications volume 16, Article number: 10814 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.05794 2025-12-04 eess.IV 78%

Robust Physics-based Deep MRI Reconstruction Via Diffusion Purification

通过扩散净化实现鲁棒的基于物理的深度MRI重建

Ismail Alkhouri, Shijun Liang, Rongrong Wang, Qing Qu, Saiprasad Ravishankar

专题命中 扩散模型 :diffusion(title,abstract)

AI总结 本文提出通过扩散模型净化提升MRI重建的鲁棒性,避免对抗训练的minimax优化,仅需微调即可提高模型稳定性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2102.13480 2025-12-04 math.AP 78%

Kinks and solitons in linear and nonlinear-diffusion Keller-Segel type models with logarithmic sensitivity

线性和非线性扩散凯勒-塞格尔型模型中的结节和孤子

Juan Campos, Claudia García, Carlos Pulido, Juan Soler

专题命中 扩散模型 :diffusion(title,abstract)

AI总结 本文研究了凯勒-塞格尔模型中线性和非线性扩散情况下旅行波模式的存在性及差异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03564 2025-12-04 cs.LG cs.CR 78%

Towards Irreversible Machine Unlearning for Diffusion Models

向扩散模型的不可逆机器遗忘学习迈进

Xun Yuan, Zilong Zhao, Jiayu Li, Aryan Pasikhani, Prosanta Gope, Biplab Sikdar

机构 * Department of Electrical and Computer Engineering, College of Design and Engineering, National University of Singapore(新加坡国立大学电子与计算机工程系,设计与工程学院) Betterdata, Singapore(新加坡Betterdata公司) Department of Computer Science, University of Sheffield(谢菲尔德大学计算机科学系)

专题命中 扩散模型 :diffusion(title,abstract)

AI总结 本文提出DiMRA攻击可逆转基于微调的扩散模型机器遗忘学习方法,并提出DiMUM方法通过记忆替代数据提升模型鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03234 2025-12-04 stat.ML cs.LG 78%

Iterative Tilting for Diffusion Fine-Tuning

迭代倾斜用于扩散微调

Jean Pachebat, Giovanni Conforti, Alain Durmus, Yazid Janati

机构 * CMAP, École Polytechnique(巴黎高等理工学院计算机学院) Università degli Studi di Padova(帕多瓦大学) Institute of Foundation Models(基础模型研究所)

专题命中 扩散模型 :diffusion(title,abstract)

AI总结 本文提出迭代倾斜方法,通过分解奖励倾斜并利用泰勒展开实现扩散模型的高效微调,适用于具有线性奖励的二维高斯混合物场景。

Comments 14 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03185 2025-12-04 math.AP math-ph math.MP 78%

Nonlinear diffusion limit of non-local interactions on a sphere

球面上非局部相互作用的非线性扩散极限

Mark A. Peletier, Anna Shalova

专题命中 扩散模型 :diffusion(title,abstract)

AI总结 本文研究了球面上非局部相互作用的非线性扩散极限,通过变分与谐分析方法证明了解收敛于具有渗流型扩散项的方程,并探讨了其在变压器模型中的应用。

Comments 44 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03127 2025-12-04 cs.LG cs.AI physics.chem-ph 78%

Atomic Diffusion Models for Small Molecule Structure Elucidation from NMR Spectra

原子扩散模型用于从NMR谱解析小分子结构

Ziyu Xiong, Yichi Zhang, Foyez Alauddin, Chu Xin Cheng, Joon Soo An, Mohammad R. Seyedsayamdost, Ellen D. Zhong

机构 * Princeton University(普林斯顿大学) California Institute of Technology(加州理工学院)

专题命中 扩散模型 :diffusion(title,abstract)

AI总结 ChefNMR通过原子扩散模型从NMR光谱直接预测小分子结构,实现高精度自动解析。

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01214 2025-12-04 physics.geo-ph 78%

Diffusion Models Bridge Deep Learning and Physics in ENSO Forecasting

扩散模型将深度学习与物理在ENSO预测中结合

Weifeng Xu, Xiang Zhu, Xiaoyong Li, Qiang Yao, Xiaoli Ren, Kefeng Deng, Song Wu, Chengcheng Shao, Xiaolong Xu, Juan Zhao, Chengwu Zhao, Jianping Cao, Jingnan Wang, Wuxin Wang, Qixiu Li, Xiaori Gao, Xinrong Wu, Huizan Wang, Xiaoqun Cao, Weiming Zhang, Junqiang Song, Kaijun Ren

专题命中 扩散模型 :diffusion(title,abstract)

AI总结 本文提出基于条件扩散模型的ENSO预测方法,通过高阶马尔可夫链构建概率映射,量化不确定性并提升预测能力,揭示反向扩散过程与充电-放电机制的联系,为ENSO预测提供新的理论框架。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22501 2025-12-04 cs.SI cs.IT math.IT math.OC 78%

A Novel Discrete-time Model of Information Diffusion on Social Networks Considering Users Behavior

一种考虑用户行为的社交网络信息扩散离散时间模型

Tran Van Khanh, Do Xuan Cho, Hoang Phi Dung

专题命中 扩散模型 :diffusion(title,abstract)

AI总结 本文提出SDIR模型,通过引入延迟状态来更精确地描述社交网络中用户的信息扩散行为,并设计算法优化信息扩散的传播效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18603 2025-12-04 cs.LG 78%

Demystify Protein Generation with Hierarchical Conditional Diffusion Models

通过分层条件扩散模型解开蛋白质生成之谜

Zinan Ling, Yi Shi, Brett McKinney, Da Yan, Yang Zhou, Bo Hui

机构 * University of Tulsa(塔尔萨大学) Johns Hopkins University(约翰霍普金斯大学) Indiana University Bloomington(印第安纳大学布卢明顿分校) Auburn University(阿伯茨维尔大学)

专题命中 扩散模型 :diffusion(title,abstract)

AI总结 本文提出了一种分层条件扩散模型,整合序列和结构信息,以高效生成功能导向的蛋白质,并引入新的评估指标评估生成质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25127 2025-12-04 cs.CV cs.AI cs.LG 77%

Score Distillation of Flow Matching Models

流匹配模型的分数蒸馏

Mingyuan Zhou, Yi Gu, Huangjie Zheng, Liangchen Song, Guande He, Yizhe Zhang, Wenze Hu, Yinfei Yang

机构 * Apple(苹果公司)

专题命中 扩散模型 :image generation(abstract);text-to-image(abstract);diffusion(abstract);分类 cs.CV

AI总结 本文提出了一种适用于文本到图像流匹配模型的分数蒸馏方法,通过统一高斯扩散和流匹配模型,实现了高效的生成加速。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03683 2025-12-04 cs.CV 70%

GaussianBlender: Instant Stylization of 3D Gaussians with Disentangled Latent Spaces

GaussianBlender: 通过解耦潜在空间实现3D高斯体的即时风格化

Melis Ocal, Xiaoyan Xing, Yue Li, Ngo Anh Vien, Sezer Karaoglu, Theo Gevers

机构 * University of Amsterdam(阿姆斯特丹大学) Bosch Center for AI(博世人工智能中心)

专题命中 扩散模型 :text-to-image(abstract);diffusion(abstract);分类 cs.CV

AI总结 GaussianBlender通过解耦潜在空间实现文本驱动的3D风格化,提供即时高保真风格化并超越传统方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.18084 2025-12-04 cs.CV cs.RO 70%

DynamicCity: Large-Scale 4D Occupancy Generation from Dynamic Scenes

DynamicCity: 从动态场景生成大规模4D占用图

Hengwei Bian, Lingdong Kong, Haozhe Xie, Liang Pan, Yu Qiao, Ziwei Liu

机构 * WorldBench Team(WorldBench团队)

专题命中 扩散模型 :diffusion(abstract);inpainting(abstract);分类 cs.CV

AI总结 DynamicCity通过VAE和DiT模型生成高质量动态4D占用图,提升拟合质量与训练效率。

Comments ICLR 2025 Spotlight; 35 pages, 18 figures, 15 tables; Project Page at https://dynamic-city.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02933 2025-12-04 cs.CV 70%

LoVoRA: Text-guided and Mask-free Video Object Removal and Addition with Learnable Object-aware Localization

LoVoRA: 基于文本引导和无掩码的视频对象移除与添加方法

Zhihan Xiao, Lin Liu, Yixin Gao, Xiaopeng Zhang, Haoxuan Che, Songping Mai, Qi Tian

机构 * Tsinghua University(清华大学) Huawei Inc.(华为公司) University of Science and Technology of China(中国科学技术大学)

专题命中 扩散模型 :diffusion(abstract);inpainting(abstract);分类 cs.CV

AI总结 LoVoRA通过可学习的对象感知定位机制实现无掩码的视频对象移除与添加,结合图像到视频转换和光流传播,实现端到端视频编辑。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03453 2025-12-04 cs.CV 57%

GeoVideo: Introducing Geometric Regularization into Video Generation Model

GeoVideo: 在视频生成模型中引入几何正则化

Yunpeng Bai, Shaoheng Fang, Chaohui Yu, Fan Wang, Qixing Huang

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校) DAMO Academy, Alibaba Group(阿里云达摩院) Hupan Lab(鸿篇实验室)

专题命中 扩散模型 :diffusion(abstract);分类 cs.CV

AI总结 GeoVideo通过引入几何正则化损失,提升视频生成的时空一致性和几何结构合理性。

Comments Project Page: https://geovideo.github.io/GeoVideo/

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03404 2025-12-04 cs.CV 57%

MOS: Mitigating Optical-SAR Modality Gap for Cross-Modal Ship Re-Identification

MOS:缓解光学-合成孔径雷达模态差距以实现跨模态舰船重识别

Yujian Zhao, Hankun Liu, Guanglin Niu

机构 * School of Artificial Intelligence, Beihang University(北京航空航天大学人工智能学院) School of Computer Science and Engineering, Beihang University(北京航空航天大学计算机科学与工程学院)

专题命中 扩散模型 :diffusion(abstract);分类 cs.CV

AI总结 MOS通过模态一致表示学习和跨模态数据生成与融合,有效缓解光学与SAR图像间的模态差距,提升舰船跨模态重识别的准确率。

详情

展开后加载摘要…

URL PDF HTML 收藏