arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

共收录 86872 信号源:cs.CV, cs.GR, cs.MM

1. 扩散模型 70277 篇

2211.10437 2023-09-07 cs.CV 80%

A Structure-Guided Diffusion Model for Large-Hole Image Completion

Daichi Horita, Jiaolong Yang, Dong Chen, Yuki Koyama, Kiyoharu Aizawa, Nicu Sebe

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

Comments BMVC2023. Code: https://github.com/UdonDa/Structure_Guided_Diffusion_Model

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.16897 2023-07-11 cs.CV cs.LG cs.SD eess.AS 80%

Physics-Driven Diffusion Models for Impact Sound Synthesis from Videos

Kun Su, Kaizhi Qian, Eli Shlizerman, Antonio Torralba, Chuang Gan

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

Comments CVPR 2023. Project page: https://sukun1045.github.io/video-physics-sound-diffusion/

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.14674 2023-05-25 cs.CV 80%

T1: Scaling Diffusion Probabilistic Fields to High-Resolution on Unified Visual Modalities

Kangfu Mei, Mo Zhou, Vishal M. Patel

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

Comments for project page, see https://t1-diffusion-model.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.08031 2023-05-16 cs.CV cs.AI 80%

On enhancing the robustness of Vision Transformers: Defensive Diffusion

Raza Imam, Muhammad Huzaifa, Mohammed El-Amine Azz

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

Comments Our code is publicly available at https://github.com/Muhammad-Huzaifaa/Defensive_Diffusion

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.10530 2023-04-21 cs.CV 80%

Collaborative Diffusion for Multi-Modal Face Generation and Editing

Ziqi Huang, Kelvin C. K. Chan, Yuming Jiang, Ziwei Liu

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

Comments CVPR 2023. Project page: https://ziqihuangg.github.io/projects/collaborative-diffusion.html Code: https://github.com/ziqihuangg/Collaborative-Diffusion

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.13622 2023-04-06 cs.LG cs.CV stat.ML 80%

Learning Data Representations with Joint Diffusion Models

Kamil Deja, Tomasz Trzcinski, Jakub M. Tomczak

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

Comments Code: https://github.com/KamilDeja/joint_diffusion

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.17598 2023-03-31 cs.CV 80%

Consistent View Synthesis with Pose-Guided Diffusion Models

Hung-Yu Tseng, Qinbo Li, Changil Kim, Suhib Alsisan, Jia-Bin Huang, Johannes Kopf

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

Comments CVPR 2023. Project page: https://poseguided-diffusion.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.09295 2023-03-17 cs.CV 80%

DIRE for Diffusion-Generated Image Detection

Zhendong Wang, Jianmin Bao, Wengang Zhou, Weilun Wang, Hezhen Hu, Hong Chen, Houqiang Li

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

Comments A general diffusion-generated image detector

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.10668 2023-02-24 cs.CV cs.AI cs.LG 80%

$PC^2$: Projection-Conditioned Point Cloud Diffusion for Single-Image 3D Reconstruction

Luke Melas-Kyriazi, Christian Rupprecht, Andrea Vedaldi

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

Comments Project page: https://lukemelas.github.io/projection-conditioned-point-cloud-diffusion

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.01714 2023-01-18 cs.CV cs.AI cs.LG 80%

Compositional Visual Generation with Composable Diffusion Models

Nan Liu, Shuang Li, Yilun Du, Antonio Torralba, Joshua B. Tenenbaum

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

Comments ECCV 2022. First three authors contributed equally. Project website: https://energy-based-model.github.io/Compositional-Visual-Generation-with-Composable-Diffusion-Models/

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.09012 2023-01-10 cs.LG cs.CV 80%

Diffusion models as plug-and-play priors

Alexandros Graikos, Nikolay Malkin, Nebojsa Jojic, Dimitris Samaras

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

Comments NeurIPS 2022; code: https://github.com/AlexGraikos/diffusion_priors

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.06135 2022-12-13 cs.CV 80%

Rodin: A Generative Model for Sculpting 3D Digital Avatars Using Diffusion

Tengfei Wang, Bo Zhang, Ting Zhang, Shuyang Gu, Jianmin Bao, Tadas Baltrusaitis, Jingjing Shen, Dong Chen, Fang Wen, Qifeng Chen, Baining Guo

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

Comments Project Webpage: https://3d-avatar-diffusion.microsoft.com/

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.02349 2022-10-06 eess.IV cs.CV cs.LG q-bio.NC 80%

Fitting a Directional Microstructure Model to Diffusion-Relaxation MRI Data with Self-Supervised Machine Learning

Jason P. Lim, Stefano B. Blumberg, Neil Narayan, Sean C. Epstein, Daniel C. Alexander, Marco Palombo, Paddy J. Slator

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

Comments Oral Presentation in: Computational Diffusion MRI Workshop (CDMRI) at Medical Image Computing and Computer Assisted Intervention (MICCAI) 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.03530 2022-03-22 cs.CV cs.AI 80%

A Conditional Point Diffusion-Refinement Paradigm for 3D Point Cloud Completion

Zhaoyang Lyu, Zhifeng Kong, Xudong Xu, Liang Pan, Dahua Lin

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

Comments Accepted to ICLR 2022. Code is released at https://github.com/ZhaoyangLyu/Point_Diffusion_Refinement

详情

展开后加载摘要…

URL PDF HTML 收藏
1805.12071 2018-10-03 cs.CV stat.AP 80%

Automatic, fast and robust characterization of noise distributions for diffusion MRI

Samuel St-Jean, Alberto De Luca, Max A. Viergever, Alexander Leemans

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

Comments v2: added publisher DOI statement, fixed text typo in appendix A2

Journal ref St-Jean S. et al. (2018) Automatic, Fast and Robust Characterization of Noise Distributions for Diffusion MRI. In: Medical Image Computing and Computer Assisted Intervention - MICCAI 2018. LNCS, vol 11070. Springer, Cham

详情

展开后加载摘要…

URL PDF HTML 收藏
1503.05768 2015-03-26 cs.CV 80%

On learning optimized reaction diffusion processes for effective image restoration

Yunjin Chen, Wei Yu, Thomas Pock

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

Comments 9 pages, 3 figures, 3 tables. CVPR2015 oral presentation together with the supplemental material of 13 pages, 8 pages (Notes on diffusion networks)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.22111 2026-08-26 cs.SD 版本更新 80%

FlowSep 2: Self-Supervised Flow Matching for Language-Queried Audio Source Separation

FlowSep 2:用于语言查询音频源分离的自监督流匹配方法

Yi Yuan, Xubo Liu, Haohe Liu, Xiyuan Kang, Mark D. Plumbley, Wenwu Wang

机构 * School of Computer Science and Electronic Engineering, University of Surrey(萨里大学计算机科学与电子工程学院) Meta Superintelligence Labs(元宇宙超级智能实验室) Department of Informatics, King’s College London(伦敦国王学院信息学系)

专题命中 扩散模型 :diffusion(summary_cn,abstract)

AI总结 本研究提出FlowSep2,一种结合Self-Flow与Diffusion Transformer的文本条件流匹配生成模型,用于语言查询音频源分离,在多个基准上达到SOTA性能,可有效分离重叠声源。

Comments Submission to IEEE/ACM Transactions on Audio, Speech, and Language Processing

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.03938 2026-08-05 cs.RO cs.SY eess.SY 新提交 80%

Bimanual Manipulation Within an 8 GB Budget: Zero-Copy Sensing and Quantized ACT on an Entry-Level Jetson

8 GB预算内的双臂操作:入门级Jetson上的零拷贝感知与量化ACT

Ekansh Singh, Eva Samuel, Alessandra Reneau, Ryan Schmeelk, Yashvi Gandhi

专题命中 扩散模型 :diffusion(summary_cn,abstract)

AI总结 该研究在8 GB入门级Jetson上实现了双臂操作,采用零拷贝感知、量化ACT,发现ACT在相同演示下优于Diffusion Policy,INT8量化可大幅降低延迟且保留任务成功率。

Comments 9 pages, 8 tables. Work conducted at the Georgia Tech Research Institute (GTRI), Aerospace, Transportation and Advanced Systems Laboratory (ATAS)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.00730 2026-08-04 cs.RO 新提交 80%

Push-Wiper: Toward General-Purpose Robotic Cleaning across Varied Stains and Surfaces with Segmented Pushing Trajectories

Push-Wiper:基于分段推动轨迹的通用机器人清洁方案,用于处理不同污渍与表面

Renhao Lu, Mingxin Wang, Chenyang Cao, Yang Yang, Guoping Pan, Kangkang Dong, Yi Cheng, Houde Liu

机构 * Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院) Z-Lab, Zerith Robotics(Zerith机器人公司Z-Lab实验室) University of Toronto(多伦多大学)

专题命中 扩散模型 :diffusion(summary_cn,abstract)

AI总结 Push-Wiper 框架将粘性污渍清洁转化为聚合问题,通过分段推动轨迹结合 Diffusion Policy 与 ASPI 控制器,清洁得分比基线高130%,可零样本泛化至多种污渍与表面。

Comments 8 pages, 8 figures. Accepted to the 2026 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.17310 2026-08-04 cs.CR 版本更新 80%

Cryptanalysis of LDPC-Based Pseudorandom Error-Correcting Codes

基于LDPC的伪随机纠错码的密码分析

Tianrui Wang, Anyu Wang, Tianshuo Cong, Delong Ran, Jinyuan Liu, Xiaoyun Wang

专题命中 扩散模型 :diffusion(summary_cn,abstract)

AI总结 首次对LDPC-PRC进行密码分析,提出三种攻击破坏其不可检测性和安全性,在DeepSeek和Stable Diffusion等模型上验证有效,并给出防御建议。

Comments Accepted by USENIX Security 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.09235 2026-07-01 cs.LG cs.AI stat.ML 版本更新 80%

On Variance Reduction in Learning Mean Flows

在学习均值流中的方差缩减

Juanwu Lu, Ziran Wang

机构 * Purdue University(普渡大学)

专题命中 扩散模型 :diffusion(summary_cn,abstract)

AI总结 本文通过理论分析揭示了均值流训练中损失非递减和梯度方差无界的问题根源,提出最优系数的闭式解,并在基准和Diffusion Transformer上验证了其提升样本质量和FID趋势的效果。

Comments 27 pages, 8 figures, 8 tables. Added supplementary experiment: independent validation of the small-bias regime, to break the circularity in bias estimation

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.30642 2026-06-30 cs.SD cs.AI 80%

LeVo 2: Stable and Melodious Song Generation via Hierarchical Representation Modeling and Progressive Post-Training

LeVo 2:通过层次表示建模和渐进式后训练实现稳定悦耳的歌曲生成

Shun Lei, Huaicheng Zhang, Dapeng Wu, Yaoxun Xu, Lishi Zuo, Wei Tan, Hangting Chen, Guangzheng Li, Jianwei Yu, Zhiyong Wu, Dong Yu

机构 * Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院) Tencent(腾讯) Wuhan University(武汉大学) Hong Kong Polytechnic University(香港理工大学)

专题命中 扩散模型 :diffusion(summary_cn,abstract)

AI总结 提出LeVo 2混合LLM-Diffusion框架,通过层次化建模(先预测混合令牌进行语义规划,再并行预测人声和伴奏令牌)解决全曲生成中协调性与细节保真度的权衡,并引入美学引导训练策略,在主观和客观指标上超越开源基线,接近商业系统。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.08044 2026-05-11 cs.CL cs.AI cs.LG 80%

Fast Byte Latent Transformer

快速字节潜在变换器

Julie Kallini, Artidoro Pagnoni, Tomasz Limisiewicz, Gargi Ghosh, Luke Zettlemoyer, Christopher Potts, Xiaochuang Han, Srinivasan Iyer

机构 * FAIR at Meta(Meta的FAIR) Stanford University(斯坦福大学) University of Washington(华盛顿大学)

专题命中 扩散模型 :diffusion(summary_cn,abstract)

AI总结 本文提出BLT Diffusion和BLT Self-speculation等方法,通过并行生成和验证步骤提升字节级语言模型的生成速度和质量,降低内存带宽消耗。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.13203 2026-04-16 cs.HC cs.AI 80%

Inclusive Kitchen Design for Older Adults: Generative AI Visualizations to Support Mild Cognitive Impairment

包容性厨房设计用于老年人:生成式AI可视化以支持轻度认知障碍

Ibrahim Bilau, Nicole Li, Terrence Malayvong, Eunhwa Yang

机构 * Georgia Institute of Technology(佐治亚理工学院)

专题命中 扩散模型 :diffusion(summary_cn,abstract)

AI总结 本研究利用生成式AI创建MCI友好的厨房设计,通过训练Stable Diffusion模型提升可视化效果,帮助老年人更易独立生活。

Comments 19 pages, 7 figures, 5 tables, IAFOR Agen2026 Conference Proceedings

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10471 2026-04-15 cond-mat.mtrl-sci cs.AI 80%

Siamese Foundation Models for Crystal Structure Prediction

孪生基础模型用于晶体结构预测

Liming Wu, Wenbing Huang, Rui Jiao, Jianxing Huang, Liwei Liu, Yipeng Zhou, Hao Sun, Yang Liu, Fuchun Sun, Yuxiang Ren, Jirong Wen

机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学人工智能学院) Beijing Key Laboratory of Research on Large Models and Intelligent Governance(北京大型模型与智能治理研究重点实验室) Department of Computer Science and Technology, Tsinghua University(清华大学计算机科学与技术系) Institute for AI Industry Research, Tsinghua University(清华大学人工智能产业研究院) Advanced Computing and Storage Lab, Huawei Technologies(华为技术有限公司先进计算与存储实验室) School of Intelligence Science and Technology, Nanjing University(南京大学智能科学与技术学院) Engineering Research Center of Next-Generation Intelligent Search and Recommendation, MOE(教育部下一代智能搜索与推荐工程研究中心)

专题命中 扩散模型 :diffusion(summary_cn,abstract)

AI总结 本文提出Diffusion-based Crystal Omni框架,结合孪生生成模型和能量预测模型,提升晶体结构预测性能,并在实际超导材料中验证其有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16233 2026-01-15 astro-ph.GA cs.LG 80%

Can AI Dream of Unseen Galaxies? Conditional Diffusion Model for Galaxy Morphology Augmentation

AI能否梦见未见的星系?面向星系形态增强的条件扩散模型

Chenrui Ma, Zechang Sun, Tao Jing, Zheng Cai, Yuan-Sen Ting, Song Huang, Mingyu Li

机构 * Tsinghua Shenzhen International Graduate School, Tsinghua University, Shenzhen China(清华大学深圳国际研究生院,清华大学,深圳中国) Department of Strategic and Advanced Interdisciplinary Research, Pengcheng Laboratory, Shenzhen China(战略与先进跨学科研究部,鹏城实验室,深圳中国) Department of Astronomy, Tsinghua University, Beijing China(天文学系,清华大学,北京中国) School of Mathematics and Physics, Qinghai University, Xining China(数学物理学院,青海大学,西宁中国) Department of Astronomy, The Ohio State University, Columbus, OH 43210, USA(天文学系,俄亥俄州立大学,哥伦布,OH 43210,美国) Center for Cosmology and AstroParticle Physics (CCAPP), The Ohio State University, Columbus, OH 43210, USA(宇宙学与天体粒子物理中心(CCAPP),俄亥俄州立大学,哥伦布,OH 43210,美国)

专题命中 扩散模型 :diffusion(title,abstract);image generation(comments)

AI总结 本研究提出条件扩散模型GalaxySD,通过合成逼真的星系图像增强ML训练数据,提升标准形态分类性能,并有效提高罕见星系检测的实例数量。

Comments 29 pages, 17 figures, accepted version for ApJS. Comments welcome. See another independent work for further reference, Category-based Galaxy Image Generation via Diffusion Models (Fan, Tang et al.)

Journal ref The Astrophysical Journal Supplement Series, 282, 25 (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
1109.2262 2014-03-27 cond-mat.mtrl-sci 80%

Diffusion in LanCoIn3n+2 phases studied by perturbed angular correlation

Randal Newhouse, Gary S. Collins

专题命中 扩散模型 :diffusion(title,abstract)

Comments Accepted for publication in Defect and Diffusion Forum, DIMAT 2011 conference, 6 pages, 5 figures, 1 table

Journal ref Defect and Diffusion Forum 323-325 (2012) 453-458

详情

展开后加载摘要…

URL PDF HTML 收藏
1109.2261 2014-03-27 cond-mat.mtrl-sci 80%

Diffusion in binary and pseudo-binary L12 indides, stannides, gallides and aluminides of rare-earth elements as studied using perturbed angular correlation of 111In/Cd

Randal Newhouse, Justine Minish, Gary S. Collins

专题命中 扩散模型 :diffusion(title,abstract)

Comments Accepted for publication in Defect and Diffusion Forum, DIMAT 2011 conference, 6 pages, 4 figures

Journal ref Defect and Diffusion Forum 323-325 (2012) 447-452

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.26794 2026-08-28 cs.CV 新提交 79%

Ring Forcing: Towards Precise Long-Term Memory for Autoregressive Video Diffusion

环强制法:面向自回归视频扩散模型的精准长期记忆

Bowen Xue, Brandon Y. Feng, Chenguo Lin, Yuchen Lin, Yujia Zeng, Lvmin Zhang, Maneesh Agrawala, Honglei Yan, Panwang Pan

机构 * Stanford University(斯坦福大学) Massachusetts Institute of Technology(麻省理工学院) Peking University(北京大学) University of California, Berkeley(加州大学伯克利分校) ByteDance(字节跳动)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 针对自回归视频扩散模型的长期记忆瓶颈,提出Ring Forcing框架,通过环形训练策略、压缩与时间步组合策略及稀疏RoPE机制,实现分钟级视频生成的优异连贯性与物体恒常性,性能优于现有方法。

Comments Project page: https://ringforcing.com

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.26580 2026-08-28 cs.CV cs.CL 新提交 79%

Visual Information-Guided Parallel Decoding for Diffusion Multimodal Large Language Models

面向扩散多模态大语言模型的视觉信息引导并行解码

Insu Lee, Wooje Park, Wonseok Shin, Jinwoo Son, Byonghyo Shim

机构 * Seoul National University(首尔大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 针对扩散多模态大语言模型解码时未充分利用输入图像信息的问题,提出视觉信息引导采样器 VIG-Sampler,在 7 个基准及 3 个开源模型上验证其性能优于 Info-Gain 采样器。

详情

展开后加载摘要…

URL PDF HTML 收藏