arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

共收录 70082 信号源:cs.CV, cs.GR, cs.MM

1. 扩散模型 70082 篇

2301.13622 2023-04-06 cs.LG cs.CV stat.ML 80%

Learning Data Representations with Joint Diffusion Models

Kamil Deja, Tomasz Trzcinski, Jakub M. Tomczak

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

Comments Code: https://github.com/KamilDeja/joint_diffusion

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.17598 2023-03-31 cs.CV 80%

Consistent View Synthesis with Pose-Guided Diffusion Models

Hung-Yu Tseng, Qinbo Li, Changil Kim, Suhib Alsisan, Jia-Bin Huang, Johannes Kopf

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

Comments CVPR 2023. Project page: https://poseguided-diffusion.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.09295 2023-03-17 cs.CV 80%

DIRE for Diffusion-Generated Image Detection

Zhendong Wang, Jianmin Bao, Wengang Zhou, Weilun Wang, Hezhen Hu, Hong Chen, Houqiang Li

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

Comments A general diffusion-generated image detector

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.10668 2023-02-24 cs.CV cs.AI cs.LG 80%

$PC^2$: Projection-Conditioned Point Cloud Diffusion for Single-Image 3D Reconstruction

Luke Melas-Kyriazi, Christian Rupprecht, Andrea Vedaldi

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

Comments Project page: https://lukemelas.github.io/projection-conditioned-point-cloud-diffusion

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.01714 2023-01-18 cs.CV cs.AI cs.LG 80%

Compositional Visual Generation with Composable Diffusion Models

Nan Liu, Shuang Li, Yilun Du, Antonio Torralba, Joshua B. Tenenbaum

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

Comments ECCV 2022. First three authors contributed equally. Project website: https://energy-based-model.github.io/Compositional-Visual-Generation-with-Composable-Diffusion-Models/

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.09012 2023-01-10 cs.LG cs.CV 80%

Diffusion models as plug-and-play priors

Alexandros Graikos, Nikolay Malkin, Nebojsa Jojic, Dimitris Samaras

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

Comments NeurIPS 2022; code: https://github.com/AlexGraikos/diffusion_priors

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.06135 2022-12-13 cs.CV 80%

Rodin: A Generative Model for Sculpting 3D Digital Avatars Using Diffusion

Tengfei Wang, Bo Zhang, Ting Zhang, Shuyang Gu, Jianmin Bao, Tadas Baltrusaitis, Jingjing Shen, Dong Chen, Fang Wen, Qifeng Chen, Baining Guo

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

Comments Project Webpage: https://3d-avatar-diffusion.microsoft.com/

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.02349 2022-10-06 eess.IV cs.CV cs.LG q-bio.NC 80%

Fitting a Directional Microstructure Model to Diffusion-Relaxation MRI Data with Self-Supervised Machine Learning

Jason P. Lim, Stefano B. Blumberg, Neil Narayan, Sean C. Epstein, Daniel C. Alexander, Marco Palombo, Paddy J. Slator

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

Comments Oral Presentation in: Computational Diffusion MRI Workshop (CDMRI) at Medical Image Computing and Computer Assisted Intervention (MICCAI) 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.03530 2022-03-22 cs.CV cs.AI 80%

A Conditional Point Diffusion-Refinement Paradigm for 3D Point Cloud Completion

Zhaoyang Lyu, Zhifeng Kong, Xudong Xu, Liang Pan, Dahua Lin

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

Comments Accepted to ICLR 2022. Code is released at https://github.com/ZhaoyangLyu/Point_Diffusion_Refinement

详情

展开后加载摘要…

URL PDF HTML 收藏
1805.12071 2018-10-03 cs.CV stat.AP 80%

Automatic, fast and robust characterization of noise distributions for diffusion MRI

Samuel St-Jean, Alberto De Luca, Max A. Viergever, Alexander Leemans

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

Comments v2: added publisher DOI statement, fixed text typo in appendix A2

Journal ref St-Jean S. et al. (2018) Automatic, Fast and Robust Characterization of Noise Distributions for Diffusion MRI. In: Medical Image Computing and Computer Assisted Intervention - MICCAI 2018. LNCS, vol 11070. Springer, Cham

详情

展开后加载摘要…

URL PDF HTML 收藏
1503.05768 2015-03-26 cs.CV 80%

On learning optimized reaction diffusion processes for effective image restoration

Yunjin Chen, Wei Yu, Thomas Pock

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

Comments 9 pages, 3 figures, 3 tables. CVPR2015 oral presentation together with the supplemental material of 13 pages, 8 pages (Notes on diffusion networks)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.03938 2026-08-05 cs.RO cs.SY eess.SY 新提交 80%

Bimanual Manipulation Within an 8 GB Budget: Zero-Copy Sensing and Quantized ACT on an Entry-Level Jetson

8 GB预算内的双臂操作:入门级Jetson上的零拷贝感知与量化ACT

Ekansh Singh, Eva Samuel, Alessandra Reneau, Ryan Schmeelk, Yashvi Gandhi

专题命中 扩散模型 :diffusion(summary_cn,abstract)

AI总结 该研究在8 GB入门级Jetson上实现了双臂操作,采用零拷贝感知、量化ACT,发现ACT在相同演示下优于Diffusion Policy,INT8量化可大幅降低延迟且保留任务成功率。

Comments 9 pages, 8 tables. Work conducted at the Georgia Tech Research Institute (GTRI), Aerospace, Transportation and Advanced Systems Laboratory (ATAS)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.00730 2026-08-04 cs.RO 新提交 80%

Push-Wiper: Toward General-Purpose Robotic Cleaning across Varied Stains and Surfaces with Segmented Pushing Trajectories

Push-Wiper:基于分段推动轨迹的通用机器人清洁方案,用于处理不同污渍与表面

Renhao Lu, Mingxin Wang, Chenyang Cao, Yang Yang, Guoping Pan, Kangkang Dong, Yi Cheng, Houde Liu

机构 * Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院) Z-Lab, Zerith Robotics(Zerith机器人公司Z-Lab实验室) University of Toronto(多伦多大学)

专题命中 扩散模型 :diffusion(summary_cn,abstract)

AI总结 Push-Wiper 框架将粘性污渍清洁转化为聚合问题,通过分段推动轨迹结合 Diffusion Policy 与 ASPI 控制器,清洁得分比基线高130%,可零样本泛化至多种污渍与表面。

Comments 8 pages, 8 figures. Accepted to the 2026 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.17310 2026-08-04 cs.CR 版本更新 80%

Cryptanalysis of LDPC-Based Pseudorandom Error-Correcting Codes

基于LDPC的伪随机纠错码的密码分析

Tianrui Wang, Anyu Wang, Tianshuo Cong, Delong Ran, Jinyuan Liu, Xiaoyun Wang

专题命中 扩散模型 :diffusion(summary_cn,abstract)

AI总结 首次对LDPC-PRC进行密码分析,提出三种攻击破坏其不可检测性和安全性,在DeepSeek和Stable Diffusion等模型上验证有效,并给出防御建议。

Comments Accepted by USENIX Security 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.09235 2026-07-01 cs.LG cs.AI stat.ML 版本更新 80%

On Variance Reduction in Learning Mean Flows

在学习均值流中的方差缩减

Juanwu Lu, Ziran Wang

机构 * Purdue University(普渡大学)

专题命中 扩散模型 :diffusion(summary_cn,abstract)

AI总结 本文通过理论分析揭示了均值流训练中损失非递减和梯度方差无界的问题根源,提出最优系数的闭式解,并在基准和Diffusion Transformer上验证了其提升样本质量和FID趋势的效果。

Comments 27 pages, 8 figures, 8 tables. Added supplementary experiment: independent validation of the small-bias regime, to break the circularity in bias estimation

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.30642 2026-06-30 cs.SD cs.AI 80%

LeVo 2: Stable and Melodious Song Generation via Hierarchical Representation Modeling and Progressive Post-Training

LeVo 2:通过层次表示建模和渐进式后训练实现稳定悦耳的歌曲生成

Shun Lei, Huaicheng Zhang, Dapeng Wu, Yaoxun Xu, Lishi Zuo, Wei Tan, Hangting Chen, Guangzheng Li, Jianwei Yu, Zhiyong Wu, Dong Yu

机构 * Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院) Tencent(腾讯) Wuhan University(武汉大学) Hong Kong Polytechnic University(香港理工大学)

专题命中 扩散模型 :diffusion(summary_cn,abstract)

AI总结 提出LeVo 2混合LLM-Diffusion框架,通过层次化建模(先预测混合令牌进行语义规划,再并行预测人声和伴奏令牌)解决全曲生成中协调性与细节保真度的权衡,并引入美学引导训练策略,在主观和客观指标上超越开源基线,接近商业系统。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.08044 2026-05-11 cs.CL cs.AI cs.LG 80%

Fast Byte Latent Transformer

快速字节潜在变换器

Julie Kallini, Artidoro Pagnoni, Tomasz Limisiewicz, Gargi Ghosh, Luke Zettlemoyer, Christopher Potts, Xiaochuang Han, Srinivasan Iyer

机构 * FAIR at Meta(Meta的FAIR) Stanford University(斯坦福大学) University of Washington(华盛顿大学)

专题命中 扩散模型 :diffusion(summary_cn,abstract)

AI总结 本文提出BLT Diffusion和BLT Self-speculation等方法,通过并行生成和验证步骤提升字节级语言模型的生成速度和质量,降低内存带宽消耗。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.13203 2026-04-16 cs.HC cs.AI 80%

Inclusive Kitchen Design for Older Adults: Generative AI Visualizations to Support Mild Cognitive Impairment

包容性厨房设计用于老年人:生成式AI可视化以支持轻度认知障碍

Ibrahim Bilau, Nicole Li, Terrence Malayvong, Eunhwa Yang

机构 * Georgia Institute of Technology(佐治亚理工学院)

专题命中 扩散模型 :diffusion(summary_cn,abstract)

AI总结 本研究利用生成式AI创建MCI友好的厨房设计,通过训练Stable Diffusion模型提升可视化效果,帮助老年人更易独立生活。

Comments 19 pages, 7 figures, 5 tables, IAFOR Agen2026 Conference Proceedings

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10471 2026-04-15 cond-mat.mtrl-sci cs.AI 80%

Siamese Foundation Models for Crystal Structure Prediction

孪生基础模型用于晶体结构预测

Liming Wu, Wenbing Huang, Rui Jiao, Jianxing Huang, Liwei Liu, Yipeng Zhou, Hao Sun, Yang Liu, Fuchun Sun, Yuxiang Ren, Jirong Wen

机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学人工智能学院) Beijing Key Laboratory of Research on Large Models and Intelligent Governance(北京大型模型与智能治理研究重点实验室) Department of Computer Science and Technology, Tsinghua University(清华大学计算机科学与技术系) Institute for AI Industry Research, Tsinghua University(清华大学人工智能产业研究院) Advanced Computing and Storage Lab, Huawei Technologies(华为技术有限公司先进计算与存储实验室) School of Intelligence Science and Technology, Nanjing University(南京大学智能科学与技术学院) Engineering Research Center of Next-Generation Intelligent Search and Recommendation, MOE(教育部下一代智能搜索与推荐工程研究中心)

专题命中 扩散模型 :diffusion(summary_cn,abstract)

AI总结 本文提出Diffusion-based Crystal Omni框架,结合孪生生成模型和能量预测模型,提升晶体结构预测性能,并在实际超导材料中验证其有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16233 2026-01-15 astro-ph.GA cs.LG 80%

Can AI Dream of Unseen Galaxies? Conditional Diffusion Model for Galaxy Morphology Augmentation

AI能否梦见未见的星系?面向星系形态增强的条件扩散模型

Chenrui Ma, Zechang Sun, Tao Jing, Zheng Cai, Yuan-Sen Ting, Song Huang, Mingyu Li

机构 * Tsinghua Shenzhen International Graduate School, Tsinghua University, Shenzhen China(清华大学深圳国际研究生院,清华大学,深圳中国) Department of Strategic and Advanced Interdisciplinary Research, Pengcheng Laboratory, Shenzhen China(战略与先进跨学科研究部,鹏城实验室,深圳中国) Department of Astronomy, Tsinghua University, Beijing China(天文学系,清华大学,北京中国) School of Mathematics and Physics, Qinghai University, Xining China(数学物理学院,青海大学,西宁中国) Department of Astronomy, The Ohio State University, Columbus, OH 43210, USA(天文学系,俄亥俄州立大学,哥伦布,OH 43210,美国) Center for Cosmology and AstroParticle Physics (CCAPP), The Ohio State University, Columbus, OH 43210, USA(宇宙学与天体粒子物理中心(CCAPP),俄亥俄州立大学,哥伦布,OH 43210,美国)

专题命中 扩散模型 :diffusion(title,abstract);image generation(comments)

AI总结 本研究提出条件扩散模型GalaxySD,通过合成逼真的星系图像增强ML训练数据,提升标准形态分类性能,并有效提高罕见星系检测的实例数量。

Comments 29 pages, 17 figures, accepted version for ApJS. Comments welcome. See another independent work for further reference, Category-based Galaxy Image Generation via Diffusion Models (Fan, Tang et al.)

Journal ref The Astrophysical Journal Supplement Series, 282, 25 (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
1109.2262 2014-03-27 cond-mat.mtrl-sci 80%

Diffusion in LanCoIn3n+2 phases studied by perturbed angular correlation

Randal Newhouse, Gary S. Collins

专题命中 扩散模型 :diffusion(title,abstract)

Comments Accepted for publication in Defect and Diffusion Forum, DIMAT 2011 conference, 6 pages, 5 figures, 1 table

Journal ref Defect and Diffusion Forum 323-325 (2012) 453-458

详情

展开后加载摘要…

URL PDF HTML 收藏
1109.2261 2014-03-27 cond-mat.mtrl-sci 80%

Diffusion in binary and pseudo-binary L12 indides, stannides, gallides and aluminides of rare-earth elements as studied using perturbed angular correlation of 111In/Cd

Randal Newhouse, Justine Minish, Gary S. Collins

专题命中 扩散模型 :diffusion(title,abstract)

Comments Accepted for publication in Defect and Diffusion Forum, DIMAT 2011 conference, 6 pages, 4 figures

Journal ref Defect and Diffusion Forum 323-325 (2012) 447-452

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.21229 2026-08-24 cs.CV 新提交 79%

Anchoring Instruction Outside Mask: Exact Reference Caching for Efficient In-Context Diffusion Transformers

在掩码外锚定指令:用于高效上下文扩散Transformer的精确引用缓存

Yangshuai Liu, Zheming Li, Jiaao Li, Kang He, Ziliang Lai, Zhitai Liu, Chengru Song

机构 * Harbin Institute of Technology(哈尔滨工业大学) KlingAI Research(KlingAI研究院)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 该研究针对上下文扩散Transformer中引用增多导致计算量过大的问题,提出掩码外锚定指令的精确引用缓存方法,通过静态文本锚点结合速度蒸馏实现高效图像编辑,在保持生成质量的同时大幅提升了去噪速度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.20818 2026-08-24 cs.LG cs.AI cs.CV 新提交 79%

Scaling Muon for Diffusion Transformers

针对扩散Transformer的Muon优化器的缩放研究

Chenghao Li, Xiao Han, Xinxin Huang, Wei Liu, Boyang Li, Bing Xiao, Heran Zhang, Juanma Perez Rua, Ke Xu, Kangning Liu, Linjun Kuang, Na Li, Tan Wang, Tian Xie, Wei Peng, Yang Pei, Yifan Xu, Yuanhao Zhai, Yuwei Lin, Zhe Wang, Zihao He, Daniel Li, Junbiao Tang, Ziyang Jiang, Dake Chen

机构 * University of Southern California(南加州大学) Meta

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 该研究针对大型扩散Transformer的Muon优化器,提出周期性行级Muon,在保留其生成质量优势的同时,大幅降低训练的计算、通信开销与时间。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.20759 2026-08-24 cs.CV 新提交 79%

DiGS-Avatar: Single-Image Animatable 3D Human Reconstruction via UV-Space Diffusion

DiGS-Avatar:基于UV空间扩散的单图像可动画三维人体重建

Jiakun Li, Li Fang, Hao Zhu, Fei Hu, Long Ye, Yuan Zhang, Jinyao Yan

机构 * Key Laboratory of Media Audio and Video (Communication University of China)(中国传媒大学媒体音频与视频重点实验室) School of Intelligence Science and Technology, Nanjing University(南京大学智能科学与技术学院)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出DiGS-Avatar,将单图像可动画三维人体重建转化为UV空间扩散的潜变量补全任务,通过师生框架优化后解码为三维高斯基元,实现了高效高质量的重建与零样本泛化。

Comments ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.19567 2026-08-24 cs.CV 版本更新 79%

Block3D: Efficient Text-to-3D Generation via Block-Wise Diffusion

Block3D:基于分块扩散的高效文本到3D生成

Bowen Cui, Weijie Wang, Zeyu Zhang, Yefei He, Mingda Lin, Haoyu Zhao, Yuanyu He, Donny Y. Chen, Feng Chen, Bohan Zhuang

机构 * ZipLab, Zhejiang University(浙江大学ZipLab) University of California, Berkeley(加州大学伯克利分校) Wuhan University(武汉大学) Monash University(莫纳什大学) University of Adelaide(阿德莱德大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 针对文本到3D生成成本高的问题,提出Block3D分块扩散框架,通过分块生成、联合去噪及置信度引导的块内修正,在保持几何保真度的同时实现5.15倍的速度提升。

Comments Code is not ready for publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.20308 2026-08-21 cs.CV 新提交 79%

DreamHand: Repurposing Video Diffusion Models for Occlusion-Robust Egocentric 3D Hand Motion Recovery

DreamHand:复用视频扩散模型实现遮挡鲁棒的第一人称视角3D手部运动恢复

Yufei Liu, Xixi Wang, Hao Li, Ganlong Zhao, Kaitong Cai, Chengkai Jin, Chunxiao Liu, Jianbo Liu, Siyuan Huang, Xingang Pan, Hongsheng Li

机构 * ACE Robotics(ACE机器人公司) Nanyang Technological University(南洋理工大学) The Chinese University of Hong Kong(香港中文大学) Shanghai Jiao Tong University(上海交通大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 DreamHand是复用视频扩散模型的离线片段级框架,通过确定性干净潜在编码器与双向时空解码器恢复带度量位置的连续双手轨迹,在五个第一人称视角基准测试中实现最佳性能,为机器人操作数据提供可扩展路径。

Comments Project Page: https://ggxxii.github.io/dreamhand/

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.19871 2026-08-21 cs.CV 新提交 79%

DIFFCZSL: Compositional Zero-Shot Learning Regularized by Diffusion Representations

DIFFCZSL:基于扩散表示正则化的组合零样本学习

Hangyu Tian, Zhenqi He, Yanghao Wang, Long Chen

机构 * The Hong Kong University of Science and Technology(香港科技大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 DIFFCZSL将预训练扩散模型的生成先验注入基于CLIP的组合零样本学习流程,通过对比对齐提升性能,在两类设置下均优于强基线,凸显了扩散表示与视觉-语言模型的互补优势。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.19644 2026-08-21 cs.CV 新提交 79%

When Guidance Goes Off-Scale: Recalibrating Diffusion Transformers under Analog Compute-in-Memory Nonidealities

当引导超出规模:模拟存算非理想性下的扩散变换器重新校准

Wenshuai Yao, Wenyong Zhou

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 该研究针对模拟存算非理想性下扩散变换器的CFG残差问题,提出采样器侧引导尺度重新校准方法,可大幅消除CIM导致的FID差距,提升生成质量。

Comments 9 pages, 8 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.19556 2026-08-21 cs.CV cs.AI 新提交 79%

Stream4D: 4D-Consistency for Streaming Autoregressive Diffusion Video Models

Stream4D:面向流式自回归扩散视频模型的4D一致性

Yuanhao Ban, Jiaqi Feng, Hengguang Zhou, Xiaohuan Pei, Justin Cui, Cho-Jui Hsieh

机构 * UCLA(加州大学洛杉矶分校) Tsinghua University(清华大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 Stream4D用显式建模场景动力学的前馈4D重建奖励替代静态评判器,结合运动先验与感知锚,提升流式自回归扩散视频模型的4D重建质量、运动保留效果及人类对齐偏好。

详情

展开后加载摘要…

URL PDF HTML 收藏