Conditional Diffusion Models for Global Precipitation Map Inpainting
专题命中 扩散模型 :diffusion(title,abstract);inpainting(title,abstract)
视觉与机器人
图像生成、文生图、图像编辑、扩散模型和可控生成。
专题命中 扩散模型 :diffusion(title,abstract);inpainting(title,abstract)
专题命中 扩散模型 :diffusion(title,abstract);inpainting(title,abstract)
Journal ref In Advances in Continuous and Discrete Models, Vol. 2025, Article No. 74, 2025
专题命中 扩散模型 :text-to-image(title,abstract);diffusion(title,abstract)
专题命中 扩散模型 :diffusion(title,abstract);inpainting(title,abstract)
Comments Accepted to ICLRW 2025 (Oral)
专题命中 扩散模型 :diffusion(title,abstract);inpainting(title,abstract)
Comments Submitted for publication to the Journal of Audio Engineering Society on January 30th, 2023
Journal ref Journal of the Audio Engineering Society 72, no. 3 (2024): 100-113
专题命中 扩散模型 :image generation(title,abstract);diffusion(title,abstract)
专题命中 扩散模型 :diffusion(title,abstract);text-to-image(title);image generation(abstract)
Comments 5 pages, 7 figures
专题命中 扩散模型 :diffusion(title,abstract);inpainting(title,abstract)
专题命中 扩散模型 :text-to-image(title,abstract);diffusion(title,abstract)
Comments 7 pages, 2 figures
专题命中 扩散模型 :diffusion(title,abstract);inpainting(title,abstract)
Comments 6 pages text + 2 pages references, 10 figures
专题命中 扩散模型 :diffusion(title,abstract);inpainting(title,abstract)
专题命中 扩散模型 :diffusion(title,abstract);inpainting(title,abstract)
专题命中 扩散模型 :diffusion(title,abstract);inpainting(title,abstract)
Journal ref In L. Calatroni, M. Donatelli, S. Morigi, M. Prato, M. Santacesaria (Eds.): Scale Space and Variational Methods in Computer Vision. Lecture Notes in Computer Science, Vol. 14009. Springer, Cham, 588-600, 2023
专题命中 扩散模型 :image generation(title,abstract);diffusion(title,abstract)
专题命中 扩散模型 :diffusion(title,abstract);inpainting(title,abstract)
专题命中 扩散模型 :diffusion(title,abstract);inpainting(title,abstract)
Comments 30 pages, 5 figures, 1 Table
专题命中 扩散模型 :diffusion(title,abstract);inpainting(title,abstract)
专题命中 扩散模型 :diffusion(title,abstract);inpainting(title,abstract)
专题命中 扩散模型 :diffusion(title,abstract);inpainting(title,abstract)
优化你的采样:基于贝叶斯优化的调优扩散采样
机构 * Cornell University(康奈尔大学)
专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);text-to-image(abstract);inpainting(abstract)
AI总结 本研究提出OYS方法,将扩散模型采样的时间步长选择作为黑盒优化问题,用贝叶斯优化直接优化目标指标,可提升多类图像生成任务性能,大幅降低推理成本且无需额外训练。
FoR-SALE:基于参考坐标系引导的LLM驱动扩散编辑中的空间调整
机构 * Department of Computer Science and Engineering(计算机科学与工程系) ; Michigan State University(密歇根州立大学)
专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);text-to-image(abstract);分类 cs.CV
AI总结 FoR-SALE是SLD的扩展,通过参考坐标系引导的潜在空间操作调整图像朝向与深度,在三个空间理解基准上仅单轮校正就将SOTA T2I模型性能提升最高7.0个百分点,解决了非相机视角空间表达的生成偏差问题。
Comments Published in COLM 2026. 10 main pages, 4 tables, 4 figures
剖析艺术领域的扩散模型:交互式模型调整与基于实践的可解释性
机构 * School of Interactive Arts and Technology, Surrey, Canada(交互艺术与技术学院,英国苏城)
专题命中 扩散模型 :diffusion(title,summary_cn);分类 cs.MM
AI总结 研究探讨创意实践中可解释人工智能,提出以实验和干预为中心的方法,通过集成模型调整和交互界面到工作流程,经对Stable Diffusion 1.5分析,助艺术家理解模型组件对生成图像的影响,让大型模型成为创意材料。
Comments Under review for "Explainable AI for the Arts" (N. Bryan-Kinns, Ed.), Springer
先进图像生成:负提示优化与潜在分类器引导
专题命中 扩散模型 :diffusion(summary_cn,abstract);image generation(title);分类 cs.CV
AI总结 该研究提出新系统,通过微调序列到序列语言模型集成负提示优化与潜在空间分类器引导,自动生成优化负提示,用混合分类器评估引导扩散步骤,减少伪影并提高语义保真度,提升Stable Diffusion图像生成质量。
稀疏-LaViDa:稀疏多模态离散扩散语言模型
机构 * Adobe(Adobe公司) ; UCLA(加州大学洛杉矶分校)
专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);text-to-image(abstract);image editing(abstract)
AI总结 针对掩码离散扩散模型推理速度慢的问题,提出Sparse-LaViDa框架,通过动态截断冗余掩码令牌加速采样,引入专用寄存器令牌和注意力掩码保证质量与一致性,基于LaViDa - O构建,在多任务中实现加速且保持质量。
Comments 18 pages (12 pages for the main paper and 6 pages for the appendix), 9 figures
用于文化遗产纺织品的人工智能:针对新型乌洛花纹合成的微调潜在扩散模型
机构 * Information Systems Study Program, Faculty of Informatics and Electrical Engineering, Institut Teknologi Del(信息系统研究项目,信息与电气工程学院,技术学院)
专题命中 扩散模型 :diffusion(title,summary_cn);分类 cs.CV
AI总结 该研究针对传统乌洛编织局限,提出在高分辨率乌洛图案数据集上微调Protogen v3.4和Stable Diffusion v1.4两个潜在扩散模型来生成新颖设计,通过多种方式评估模型性能,发现特定引导尺度能平衡保真度与多样性,支持非物质文化遗产创新更新。
Comments 21 pages, 8 figures, 3 tables. The manuscript is currently under review at the 2026 4th International Conference on Data, Information and Computing Science (https://www.cdics.org/)
扩散模型中时间步嵌入的冗余性研究
机构 * Independent Researcher, Lima, Peru(独立研究者,秘鲁利马)
专题命中 扩散模型 :diffusion(title,summary_cn);分类 cs.CV
AI总结 本文通过理论和实验证明,在U-Net和Diffusion Transformer架构中,扩散模型无需显式时间步嵌入也能达到全局最优,甚至在某些指标上超越有条件模型。
Comments 17 pages
TexTailor:用于多模态扩散变压器的推理时文本引导剪裁
机构 * Fudan University(复旦大学) ; Shanghai Innovation Institute(上海创新研究院) ; Shanghai Academy of AI for Science(上海人工智能科学研究院)
专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);text-to-image(abstract);image editing(abstract)
AI总结 研究基于变压器的扩散模型中不同模块与文本条件的交互。通过系统管道分析各模块功能,提出推理时逐块文本引导剪裁方法,提升文本图像对齐,有多种下游应用,实验表明优于基线。
Comments To appear in ECCV 2026
UltraImageGen: 高效超高清图像生成与层级局部注意力
专题命中 扩散模型 :image generation(title,abstract);text-to-image(abstract);diffusion(abstract);image synthesis(abstract)
AI总结 提出UltraImageGen框架,通过层级局部注意力和低分辨率全局引导,将预训练扩散模型扩展到8K以上分辨率,实现近线性复杂度、10倍加速和更低内存占用。
Comments 31 pages
Adv-TGD:面向人脸识别冒充攻击的对抗性文本引导扩散
机构 * University of South Florida, Bellini College of Artificial Intelligence, Cybersecurity and Computing(南佛罗里达大学贝利尼人工智能、网络安全与计算学院)
专题命中 扩散模型 :diffusion(title,summary_cn);分类 cs.CV
AI总结 提出Adv-TGD框架,利用Stable Diffusion和LoRA微调生成逼真对抗人脸,在保持视觉质量的同时实现高成功率身份冒充攻击,平均ASR达85.90%。
破碎的记忆:通过退化生成检测和缓解扩散模型中的记忆化
机构 * Fudan University(复旦大学) ; East China University of Science and Technology(东华大学)
专题命中 扩散模型 :diffusion(title,summary_cn);分类 cs.CV
AI总结 本文首次发现扩散模型中的记忆化会导致内部数值不稳定性并表现为视觉“破碎”伪影,基于此提出了一种基于潜变量更新范数的经验稳定区域来量化稳定行为,并设计了一个即时的逐步骤检测与自适应缓解框架,在不改变提示或引导的情况下抑制记忆化,在Stable Diffusion 1.4上实现了AUC>0.999的检测性能和0.0%的记忆化率。
Comments KDD 2026, extended version