arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

2026-01-08 至 2026-01-08 共收录 37 信号源:cs.CV, cs.GR, cs.MM

1. 扩散模型 37 篇

2503.18719 2026-01-08 cs.CV 83%

Boosting Resolution Generalization of Diffusion Transformers with Randomized Positional Encodings

通过随机化位置编码提升扩散变换器的分辨率泛化能力

Liang Hou, Cong Liu, Mingwu Zheng, Xin Tao, Pengfei Wan, Di Zhang, Kun Gai

机构 * Kling Team, Kuaishou Technology(快手科技 Kling Team)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV

AI总结 本文提出RPE-2D,通过随机化位置编码提升扩散变换器的分辨率泛化能力,实现高分辨率和低分辨率图像生成的无缝过渡。

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07960 2026-01-08 cs.CV 83%

VisualCloze: A Universal Image Generation Framework via Visual In-Context Learning

VisualCloze: 一种通过视觉上下文学习实现通用图像生成的框架

Zhong-Yu Li, Ruoyi Du, Juncheng Yan, Le Zhuo, Qilong Wu, Zhen Li, Peng Gao, Zhanyu Ma, Ming-Ming Cheng

机构 * VCIP, CS, Nankai University(VCIP、计算机科学、南开大学) Beijing University of Posts and Telecommunications(北京邮电大学) Tsinghua University(清华大学) Shanghai AI Laboratory(上海人工智能实验室) The Chinese University of Hong Kong(香港中文大学)

专题命中 扩散模型 :image generation(title,abstract);diffusion(abstract);分类 cs.CV

AI总结 VisualCloze通过视觉上下文学习实现通用图像生成,支持多种任务泛化与反向生成,利用图结构数据集提升任务密度和知识迁移。

Comments Accepted at ICCV 2025. Project page: https://visualcloze.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.04056 2026-01-08 cs.CL 82%

Bridging the Discrete-Continuous Gap: Unified Multimodal Generation via Coupled Manifold Discrete Absorbing Diffusion

弥合离散-连续鸿沟:通过耦合流形离散吸收扩散实现统一多模态生成

Yuanfeng Xu, Yuhao Chen, Liang Lin, Guangrun Wang

机构 * Sun Yat-sen University(中山大学) Guangdong Key Lab of Big Data Analysis & Processing(广东省大数据分析与处理重点实验室) X-Era AI Lab(X-Era AI实验室)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract)

AI总结 CoM-DAD通过耦合流形离散吸收扩散框架,实现离散与连续数据的统一多模态生成,解决多模态对齐与训练稳定性问题。

Comments 10 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03959 2026-01-08 cs.CV 79%

FUSION: Full-Body Unified Motion Prior for Body and Hands via Diffusion

FUSION: 全身统一运动先验用于身体和手部的扩散模型

Enes Duran, Nikos Athanasiou, Muhammed Kocabas, Michael J. Black, Omid Taheri

机构 * Max Planck Institute for Intelligent Systems(马克斯·普朗克智能系统研究所) University of Tübingen(图宾根大学) Meshcapade GmbH(Meshcapade公司)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 FUSION提出首个基于扩散的全身运动先验模型,联合建模身体和手部运动,实现更自然的运动生成和多样化应用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.18455 2026-01-08 cs.CV 79%

Plasticine: A Traceable Diffusion Model for Medical Image Translation

Plasticine: 一种用于医学图像翻译的可追溯扩散模型

Tianyang Zhang, Xinxing Cheng, Jun Cheng, Shaoming Zheng, He Zhao, Huazhu Fu, Alejandro F Frangi, Jiang Liu, Jinming Duan

机构 * Department of Computer Science, University of Birmingham(计算机科学系,伯明翰大学) Institute for Infocomm Research, A*STAR(信息与通信研究所,A*STAR) Imperial College London(伦敦帝国学院) Department of Eye and Vision Science, University of Liverpool(眼科与视觉科学系,利物浦大学) Institute of High Performance Computing, A*STAR(高性能计算研究所,A*STAR) Department of Computer Science, and the Division of Informatics, Imaging and Data Sciences, School of Health Sciences, The University of Manchester(计算机科学系,信息、成像与数据科学分会,健康科学学院,曼彻斯特大学) Southern University of Science and Technology(南方科技大学) University of Birmingham(伯明翰大学) University of Manchester(曼彻斯特大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 Plasticine是一种以可追溯性为核心目标的端到端医学图像翻译框架,通过结合强度翻译和空间变换,在去噪扩散框架中生成具有可解释性强度过渡和空间一致变形的合成图像。

Comments Accepted by IEEE Transactions on Artificial Intelligence

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09996 2026-01-08 cs.CV q-bio.NC q-bio.QM 79%

Leveraging Swin Transformer for enhanced diagnosis of Alzheimer's disease using multi-shell diffusion MRI

利用Swin Transformer提升多壳扩散MRI在阿尔茨海默病诊断中的应用

Quentin Dessain, Nicolas Delinte, Bernard Hanseeuw, Laurence Dricot, Benoît Macq

机构 * ICTEAM Institute, UCLouvain(ICTEAM研究院,UCLouvain大学) Institute of Neuroscience (IoNS), UCLouvain(神经科学研究所(IoNS),UCLouvain大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本研究利用Swin Transformer模型,通过多壳扩散MRI数据提升阿尔茨海默病诊断和淀粉样蛋白检测的准确率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.18414 2026-01-08 cs.CV 79%

U-REPA: Aligning Diffusion U-Nets to ViTs

U-REPA: 对齐扩散U-Net与ViT

Yuchuan Tian, Hanting Chen, Mengyu Zheng, Yuchen Liang, Chao Xu, Yunhe Wang

机构 * State Key Lab of General AI, School of Intelligence Science and Technology, Peking University(人工智能通用基础理论国家重点实验室,智能科学与技术学院,北京大学) Huawei Noah’s Ark Lab(华为诺亚实验室) The University of Sydney(悉尼大学) School of Mathematical Sciences, Peking University(数学科学学院,北京大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 U-REPA通过改进的表示对齐方法,提升扩散模型在U-Net架构中的生成质量和收敛速度。

Comments 22 pages, 8 figures

Journal ref NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03305 2026-01-08 cs.CV cs.AI cs.CY 79%

Mass Concept Erasure in Diffusion Models with Concept Hierarchy

扩散模型中的概念擦除与概念层次

Jiahang Tu, Ye Li, Yiming Wu, Hanbin Zhao, Chao Zhang, Hui Qian

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出了一种基于概念层次的扩散模型擦除方法,通过群体擦除和超类型保留低秩适应技术,提高擦除效率并减少生成质量退化。

Comments This paper has been accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12631 2026-01-08 cs.CV cs.AI 79%

Multivariate Diffusion Transformer with Decoupled Attention for High-Fidelity Mask-Text Collaborative Facial Generation

多变量扩散变换器与解耦注意力用于高保真遮罩-文本协同面部生成

Yushe Cao, Dianxi Shi, Xing Fu, Xuechao Zou, Haikuo Peng, Xueqi Li, Chun Yu, Junliang Xing

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 MDiTFace通过解耦注意力机制和多变量变换块提升遮罩-文本协同面部生成的保真度与一致性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00412 2026-01-08 cs.CV 79%

Sortblock: Similarity-Aware Feature Reuse for Diffusion Model

Sortblock: 基于相似性的特征重用以提升扩散模型

Hanqi Chen, Xu Zhang, Xiaoliu Guan, Lielin Jiang, Guanzhong Wang, Zeyu Chen, Yi Liu

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 Sortblock通过动态缓存块级特征和轻量级预测机制,实现扩散模型的高效推理加速,提升生成质量的同时显著减少推理时间。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.09632 2026-01-08 cs.RO cs.CV 79%

Adaptive Anomaly Recovery for Telemanipulation: A Diffusion Model Approach to Vision-Based Tracking

适应性异常恢复用于远程操控:一种基于扩散模型的视觉跟踪方法

Haoyang Wang, Haoran Guo, Lingfeng Tao, Zhengxiong Li

机构 * Oklahoma State University(俄克拉荷马州立大学) University of Colorado Denver(科罗拉多大学丹佛分校) Kennesaw State University(凯斯威克州立大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出基于扩散模型的DET框架,通过帧差检测技术识别和分割视频流中的异常,利用扩散模型重建异常片段以提升远程操控在复杂视觉条件下的鲁棒性。

Comments This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.04054 2026-01-08 cs.LG 78%

LinkD: AutoRegressive Diffusion Model for Mechanical Linkage Synthesis

LinkD:用于机械连杆合成的自回归扩散模型

Yayati Jadhav, Amir Barati Farimani

机构 * Department of Mechanical Engineering, Carnegie Mellon University, Pittsburgh, PA, USA(机械工程系,卡内基梅隆大学,匹兹堡,PA,USA)

专题命中 扩散模型 :diffusion(title,abstract)

AI总结 LinkD通过自回归扩散模型实现机械连杆的高效逆向设计,结合因果Transformer和DDPM,支持动态生成和自适应修正,适用于大规模复杂机械系统。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03903 2026-01-08 cs.IR 78%

Unleashing the Potential of Neighbors: Diffusion-based Latent Neighbor Generation for Session-based Recommendation

释放邻居的潜力:基于扩散的潜在邻居生成用于基于会话的推荐

Yuhan Yang, Jie Zou, Guojia An, Jiwei Wei, Yang Yang, Heng Tao Shen

专题命中 扩散模型 :diffusion(title,abstract)

AI总结 DiffSBR通过基于扩散的潜在邻居生成方法,提升基于会话的推荐性能。

Comments This paper has been accepted by KDD 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03701 2026-01-08 cs.LG cs.AI 78%

Inference Attacks Against Graph Generative Diffusion Models

针对图生成扩散模型的推断攻击

Xiuling Wang, Xin Huang, Guibo Luo, Jianliang Xu

机构 * Hong Kong Baptist University(香港 Baptist 大学) Peking University(北京大学)

专题命中 扩散模型 :diffusion(title,abstract)

AI总结 本文针对图生成扩散模型提出三种推断攻击,并设计防御机制以减少信息泄露风险。

Comments This work has been accepted by USENIX Security 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03397 2026-01-08 cs.CE cs.LG 78%

PIVONet: A Physically-Informed Variational Neuro ODE Model for Efficient Advection-Diffusion Fluid Simulation

PIVONet:一种结合物理信息的变分神经ODE模型,用于高效对流-扩散流体模拟

Hei Shing Cheung, Qicheng Long, Zhiyue Lin

机构 * Division of Engineering Science University of Toronto(工程科学系 马他大学)

专题命中 扩散模型 :diffusion(title,abstract)

AI总结 PIVONet结合神经ODE与连续归一化流,通过变分方法实现高效且能捕捉湍流的流体模拟。

Comments 13 pages, 14 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03259 2026-01-08 cs.IR 78%

LLMDiRec: LLM-Enhanced Intent Diffusion for Sequential Recommendation

LLMDiRec: 基于大语言模型的意图扩散增强序列推荐

Bo-Chian Chen, Manel Slokom

专题命中 扩散模型 :diffusion(title,abstract)

AI总结 LLMDiRec通过整合大语言模型与意图感知扩散模型,提升序列推荐中复杂用户意图捕捉和长尾项推荐效果。

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.20604 2026-01-08 cs.CL 78%

MoE-DiffuSeq: Enhancing Long-Document Diffusion Models with Sparse Attention and Mixture of Experts

MoE-DiffuSeq:通过稀疏注意力和专家混合架构增强长文档扩散模型

Alexandros Christoforos, Chadbourne Davis

专题命中 扩散模型 :diffusion(title,abstract)

AI总结 MoE-DiffuSeq通过稀疏注意力和专家混合架构提升长文档扩散模型的效率与生成质量。

Comments Under submission

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04924 2026-01-08 eess.SP cs.SY eess.SY 78%

Steady-State Spread Bounds for Graph Diffusion via Laplacian Regularisation in Networked Systems

图扩散中通过拉普拉斯正则化实现的稳态扩散界限

Ardavan Rahimian

专题命中 扩散模型 :diffusion(title,abstract)

AI总结 本文通过拉普拉斯正则化方法,研究了图扩散过程中稳态扩散的界限,并提出了一种基于正则化强度的设计规则,以控制扩散范围。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07770 2026-01-08 eess.SP cs.LG 78%

Channel Estimation for RIS-Assisted mmWave Systems via Diffusion Models

通过扩散模型实现RIS辅助毫米波系统的信道估计

Yang Wang, Yin Xu, Cixiao Zhang, Zhiyong Chen, Mingzeng Dai, Haiming Wang, Bingchao Liu, Dazhi He, Meixia Tao

机构 * Cooperative Medianet Innovation Center, Shanghai Jiao Tong University(上海交通大学合作中研网创新中心) Lenovo Group, Lenovo Research(联想集团,联想研究)

专题命中 扩散模型 :diffusion(title,abstract)

AI总结 本文提出基于扩散模型的RIS辅助毫米波系统信道估计方法,通过改进的DDIMs采样算法和轻量级BRCNet网络,实现更高效的信道估计性能。

Comments 5 pages, 3 figures

Journal ref IEEE Communications Letters, vol.30, pp.597-601, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.09503 2026-01-08 math.OC 78%

Data-driven identification of reaction-diffusion dynamics from finitely many non-local noisy measurements by exponential fitting

基于指数拟合的数据驱动识别反应扩散动力学:从有限非局部噪声测量中获取主导特征值和初始条件模式

Rami Katz, Giulia Giordano, Dmitry Batenkov

专题命中 扩散模型 :diffusion(title,abstract)

AI总结 本文提出了一种基于指数拟合的数据驱动方法,用于从有限非局部噪声测量中准确估计反应扩散方程的主导特征值和初始条件模式。

Journal ref IFAC Journal of Systems and Control, 35, 100350, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.18109 2026-01-08 cs.CV 74%

Difficulty Controlled Diffusion Model for Synthesizing Effective Training Data

困难控制扩散模型用于合成有效训练数据

Zerun Wang, Jiafeng Mao, Xueting Wang, Toshihiko Yamasaki

机构 * CyberAgent

专题命中 扩散模型 :diffusion(title);分类 cs.CV

AI总结 本文提出了一种困难控制扩散模型,通过引入学习难度作为额外条件信号,高效生成具有显著性能提升的困难样本,从而在多个数据集上实现了更低的生成成本和更高的性能。

Comments AAAI 2026 accepted

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03591 2026-01-08 cond-mat.stat-mech cond-mat.soft physics.bio-ph 71%

Interplay of activity and non-reciprocity in tracer dynamics: From non-equilibrium fluctuation-dissipation to giant diffusion

活动性与非互惠性相互作用在示踪动力学中的作用:从非平衡波动-耗散到巨扩散

Subhajit Paul, Debasish Chaudhuri

专题命中 扩散模型 :diffusion(title)

AI总结 研究揭示了非互惠相互作用如何通过非平衡波动-耗散关系影响示踪体扩散,发现非互惠性增强导致巨扩散,对主动软物质和生物系统有重要影响。

Comments 14 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.01045 2026-01-08 cs.CV cs.GR 62%

WonderHuman: Hallucinating Unseen Parts in Dynamic 3D Human Reconstruction

WonderHuman: 在动态3D人体重建中生成未见部分

Zilong Wang, Zhiyang Dou, Yuan Liu, Cheng Lin, Xiao Dong, Yunhui Guo, Chenxu Zhang, Xin Li, Wenping Wang, Xiaohu Guo

机构 * Department of Computer Science, The University of Texas at Dallas(德克萨斯大学达拉斯分校计算机科学系) Computer Graphics Group, The University of Hong Kong(香港大学计算机图形组) School of Engineering, The Hong Kong University of Science and Technology(香港科学与技术大学工程学院) Guangdong Provincial/Zhuhai Key Laboratory of IRADS, Beijing Normal-Hong Kong Baptist University(广东/珠海IRADS重点实验室,北京师范大学-香港 Baptist大学) Department of Computer Science & Engineering, Texas A&M University(德克萨斯A&M大学计算机科学与工程系)

专题命中 扩散模型 :diffusion(abstract);分类 cs.CV、cs.GR

AI总结 WonderHuman通过双空间优化和分数蒸馏采样技术,从单目视频实现高保真动态人体重建,尤其擅长生成未见人体部分。

Journal ref IEEE Transactions on Visualization and Computer Graphics, vol. 31, no. 12, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03993 2026-01-08 cs.CV 57%

PosterVerse: A Full-Workflow Framework for Commercial-Grade Poster Generation with HTML-Based Scalable Typography

PosterVerse: 一个用于商业级海报生成的全流程框架,基于HTML的可扩展字体

Junle Liu, Peirong Zhang, Yuyi Zhang, Pengyu Yan, Hui Zhou, Xinyue Zhou, Fengjun Guo, Lianwen Jin

专题命中 扩散模型 :diffusion(abstract);分类 cs.CV

AI总结 PosterVerse通过全流程框架和HTML驱动技术,实现商业级海报的高密度文本渲染与灵活设计。

Journal ref AAAI 2026 Oral

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03665 2026-01-08 cs.CV 57%

PhysVideoGenerator: Towards Physically Aware Video Generation via Latent Physics Guidance

PhysVideoGenerator: 通过潜在物理引导实现物理感知的视频生成

Siddarth Nilol Kundur Satish, Devesh Jaiswal, Hongyu Chen, Abhishek Bakshi

机构 * Center for Data Science New York University(数据科学中心 新 York 大学) Department of Computer Science New York University(计算机科学系 新 York 大学)

专题命中 扩散模型 :diffusion(abstract);分类 cs.CV

AI总结 PhysVideoGenerator通过在视频生成过程中引入可学习的物理先验,实现物理感知的视频生成,展示了联合训练范式的可行性。

Comments 9 pages, 2 figures, project page: https://github.com/CVFall2025-Project/PhysVideoGenerator

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13720 2026-01-08 cs.CV 57%

Back to Basics: Let Denoising Generative Models Denoise

回归基础:让去噪生成模型去噪

Tianhong Li, Kaiming He

机构 * MIT(麻省理工学院)

专题命中 扩散模型 :diffusion(abstract);分类 cs.CV

AI总结 本文提出通过直接预测干净数据的JiT模型,在高维空间中实现有效的去噪生成模型,展示在ImageNet上取得竞争性结果。

Comments Tech report. Code at https://github.com/LTH14/JiT

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07828 2026-01-08 cs.CV cs.AI 57%

Benchmarking Content-Based Puzzle Solvers on Corrupted Jigsaw Puzzles

在受腐蚀的图片拼图上基准测试基于内容的谜题求解器

Richard Dirauf, Florian Wolz, Dario Zanca, Björn Eskofier

机构 * Machine Learning and Data Analytics Lab, Department Artificial Intelligence in Biomedical Engineering, Friedrich-Alexander-Universität Erlangen-Nürnberg (FAU)(机器学习与数据分析实验室,人工智能生物医学工程系,弗里德里希-亚历山大-亚琛-纽伦堡大学) Siemens Healthineers AG(西门子医疗影像股份有限公司)

专题命中 扩散模型 :diffusion(abstract);分类 cs.CV

AI总结 本文评估了三种类型的图片拼图腐蚀对基于内容的谜题求解器的影响,发现深度学习模型通过微调增强数据能显著提高鲁棒性,位置扩散模型表现尤为突出。

Journal ref ICIAP 2025, Lecture Notes in Computer Science 16167 (2026) pp 286-298

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.04013 2026-01-08 physics.ao-ph physics.flu-dyn 50%

Effects of Horizontal Discretization on Triangular and Hexagonal Grids on Linear Baroclinic and Symmetric Instabilities

水平离散化对三角形和六边形网格上线性巴克林ic和对称不稳定性的影响

Steffen Maaß, Sergey Danilov

专题命中 扩散模型 :diffusion(abstract)

AI总结 研究三角形和六边形网格上水平离散化对巴克林ic和对称不稳定性的影响,揭示了不同离散化方案在不稳定性增长率上的差异及校准参数的重要性。

Comments 36 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.14134 2026-01-08 cond-mat.soft 50%

Enhancing Swelling Kinetics of pNIPAM Lyogels: The Role of Crosslinking, Copolymerization, and Solvent

增强pNIPAM凝胶的吸水动力学:交联、共聚和溶剂的作用

Kathrin Marina Eckert, Jelisa Bonsen, Anja Hajnal, Johannes Gmeiner, Jonah Hasse, Muhammad Adrian, Julian Karsten, Patrick A. Kißling, Alexander Penn, Bodo Fiedler, Gerrit A. Luinstra, Irina Smirnova

专题命中 扩散模型 :diffusion(abstract)

AI总结 本研究探讨了通过交联、共聚和溶剂调控提高pNIPAM凝胶吸水动力学的方法,发现交联度影响机械稳定性与响应时间,溶剂性质显著影响膨胀行为。

Comments 26 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
1504.08139 2026-01-08 math.AP 50%

On massless electron limit for a multispecies kinetic system with external magnetic field

多物种动能系统在外部磁场下的无质量电子极限

Maxime Herda

专题命中 扩散模型 :diffusion(abstract)

AI总结 本文研究了在外部磁场下多物种动能系统中电子质量趋于零时的极限行为,推导出各向异性漂移扩散方程作为宏观电子密度的描述。

详情

展开后加载摘要…

URL PDF HTML 收藏