Kiss3DGen: Repurposing Image Diffusion Models for 3D Asset Generation
专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV、cs.GR、cs.MM
Comments The first three authors contributed equally to this work
视觉与机器人
图像生成、文生图、图像编辑、扩散模型和可控生成。
专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV、cs.GR、cs.MM
Comments The first three authors contributed equally to this work
专题命中 扩散模型 :diffusion(title,abstract);image synthesis(abstract);分类 cs.CV、cs.GR、cs.MM
Comments Accepted in The International Conference on Pattern Recognition (ICPR) 2024
专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);image editing(abstract)
Comments Published at 2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)
专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);text-to-image(abstract)
专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);personalized generation(abstract)
Comments ECCV 2024 Camera-Ready Version
专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);image synthesis(abstract)
专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);text-to-image(abstract)
专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);text-to-image(abstract)
Comments ICML 2024
专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);text-to-image(abstract)
Comments Published in Transactions on Machine Learning Research (07/2023)
专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);image synthesis(abstract)
Comments 41 pages, 12 figures
专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);inpainting(abstract)
Comments Demo and implementation at https://auffusion.github.io
专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);text-to-image(abstract)
专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);inpainting(abstract)
Comments This survey paper is accepted by IJCAI 2023
专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);image synthesis(abstract)
Comments 7 pages, 8 figures, ICML 2023 Workshop on Challenges in Deployable Generative AI
专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);image synthesis(abstract)
专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);image synthesis(abstract)
专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);分类 cs.CV、cs.GR
Comments 11 pages, 9 figures, Project page: https://unity-research.github.io/Geometry-Image-Diffusion.github.io/
专题命中 扩散模型 :diffusion(title,abstract);inpainting(abstract);分类 cs.CV、cs.GR
Comments CVPR 2024. Code and additional visualizations available: https://single-mesh-diffusion.github.io/
专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV、cs.GR
Comments 44 pages, 28 figures. A briefer version was presented at NeurIPS23 Workshop on Diffusion Models [arXiv:2311.10892]
专题命中 扩散模型 :diffusion(title,abstract);image editing(abstract);分类 cs.CV、cs.GR
Comments CVPR 2022. Code is available at: https://omriavrahami.com/blended-diffusion-page/
几何至关重要:用于学习语义对应的3D基础先验
机构 * University of Freiburg(弗赖堡大学) ; Max Planck Institute for Informatics(马克斯·普朗克信息研究所) ; CISPA Helmholtz Center for Information Security(CISPA 河岸信息安全中心)
专题命中 扩散模型 :diffusion(summary_cn,abstract);text-to-image(abstract);分类 cs.CV
AI总结 提出一种3D感知的后训练框架,利用3D基础模型(SAM3D)估计物体几何和姿态,生成几何感知特征图,结合DINO和Stable Diffusion特征,通过测地距离过滤候选对应,训练轻量适配器改进语义对应。
Comments 9 pages (main paper), 21 pages (total), 4 figures
通过测试时训练线性化视觉Transformer
机构 * Tsinghua University(清华大学)
专题命中 扩散模型 :diffusion(summary_cn,abstract);text-to-image(abstract);分类 cs.CV
AI总结 提出利用测试时训练(TTT)架构与Softmax注意力的结构对齐性,结合键实例归一化和局部性增强模块,实现从预训练Transformer到线性注意力模型的有效权重迁移,在Stable Diffusion 3.5上仅需1小时微调即可达到相近的图像生成质量并加速推理。
Comments ICML 2026
MMCORE:多模态连接与表征对齐的潜在嵌入
机构 * ByteDance Seed(字节跳动种子)
专题命中 扩散模型 :image generation(abstract);text-to-image(abstract);diffusion(abstract);image editing(abstract)
AI总结 MMCORE通过预训练视觉-语言模型生成语义视觉嵌入,结合扩散模型实现多模态图像生成与编辑,提升生成质量并降低计算开销。
MaMe & MaRe:基于矩阵的令牌合并与恢复用于高效的视觉感知与合成
专题命中 扩散模型 :diffusion(summary_cn,abstract);image synthesis(abstract);分类 cs.CV
AI总结 本文提出MaMe和MaRe,通过矩阵运算实现免训练的令牌合并与恢复,提升视觉模型效率,实验显示在ViT-B上吞吐量提升2%且精度下降1.0%,在图像合成中减少Stable Diffusion v2.1生成延迟31%。
Comments 20 pages. Extended version of CVPR 2026 Findings paper. Neurocomputing (Elsevier) under review
专题命中 扩散模型 :image generation(title);text-to-image(title);分类 cs.CV
专题命中 扩散模型 :image generation(title);diffusion(title);分类 cs.CV
少量通道绘制整体画面:揭示扩散变换器中的大规模激活
机构 * University of Modena and Reggio Emilia(摩德纳和雷吉奥艾米利亚大学) ; University of Pisa(比萨大学)
专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);分类 cs.CV、cs.MM
AI总结 研究扩散变换器中少量关键通道对图像生成的决定性作用,揭示其在功能、空间组织和可迁移性上的核心贡献。
Comments Project page: https://aimagelab.github.io/MAs-DiT/
PixGS: 像素空间扩散用于直接三维高斯泼溅生成
机构 * Qualcomm AI Research(高通人工智能研究院)
专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);分类 cs.CV、cs.GR
AI总结 提出单阶段管道PixGS,利用像素空间扩散从文本或图像直接生成高质量3D高斯泼溅,避免潜在压缩伪影,并引入表面法线、深度和高频结构监督,在单张A100 GPU上1秒内生成,性能超越现有方法。
Comments Accepted at ECCV 2026
Edit3DGS:通过2D指令引导扩散与3D高斯泼溅的动态头部编辑统一框架
机构 * University of Science, VNU-HCM, Ho Chi Minh, Vietnam(越南胡志明市国家大学) ; Vietnam National University, Ho Chi Minh, Vietnam(越南国家大学)
专题命中 扩散模型 :diffusion(title,abstract);inpainting(abstract);分类 cs.CV、cs.GR
AI总结 提出Edit3DGS统一框架,结合2D指令引导扩散与3D高斯泼溅,实现动态3D头部的可控编辑,支持表情变换、属性修改等操作,并保持身份与运动动态的一致性。
Comments SOICT 2025
上下文空间中的即时排斥以实现扩散变换器的丰富多样性
机构 * Tel Aviv University(特拉维夫大学) ; Snap Research Israel(Snap以色列研究)
专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);分类 cs.CV、cs.GR
AI总结 针对文本到图像扩散模型多样性不足的问题,提出在扩散变换器的上下文空间中通过多模态注意力通道施加即时排斥,在不牺牲视觉保真度和语义一致性的前提下显著提升生成多样性,且计算开销小,适用于现代Turbo和蒸馏模型。
Comments SIGGRAPH 2026. Project page: https://contextual-repulsion.github.io/