Mask-ControlNet: Higher-Quality Image Generation with An Additional Mask Prompt
专题命中 可控生成 :image generation(title,abstract);text-to-image(abstract);diffusion(abstract);分类 cs.CV
视觉与机器人
图像生成、文生图、图像编辑、扩散模型和可控生成。
专题命中 可控生成 :image generation(title,abstract);text-to-image(abstract);diffusion(abstract);分类 cs.CV
专题命中 可控生成 :image generation(title,abstract);text-to-image(abstract);diffusion(abstract);分类 cs.CV
Comments CVPR 2024
专题命中 可控生成 :image synthesis(title,abstract);text-to-image(abstract);diffusion(abstract);分类 cs.CV
Comments Demo: https://huggingface.co/spaces/Pusheen/LoCo; Project page: https://momopusheen.github.io/LoCo/
专题命中 可控生成 :inpainting(title,abstract);text-to-image(abstract);diffusion(abstract);分类 cs.CV
Comments Project Page: https://me.kiui.moe/intex/
专题命中 可控生成 :image synthesis(title,abstract);image generation(abstract);diffusion(abstract);分类 cs.CV
专题命中 可控生成 :image generation(title,abstract);text-to-image(abstract);diffusion(abstract);分类 cs.CV
Comments Hongsuk Choi and Isaac Kasahara have eqaul contributions. 19 pages, 15 figures, 3 tables
专题命中 可控生成 :diffusion(title,abstract);image generation(abstract);text-to-image(abstract);分类 cs.CV
Comments NeurIPS 2023
专题命中 可控生成 :diffusion(title,abstract);image generation(abstract);image synthesis(abstract);分类 cs.CV
Comments Accepted by 3DV24
专题命中 可控生成 :diffusion(title,abstract);text-to-image(abstract);image editing(abstract);分类 cs.CV
GeoDiff-SAR II: 3D-Driven Foundation Diffusion Models for SAR Generation via Decoupled Control
专题命中 可控生成 :diffusion(title,title_cn);image generation(abstract)
AI总结 本文提出GeoDiff-SAR II,一种基于3D模型引导的解耦框架,用于通过解耦控制生成合成孔径雷达图像,通过物理基础的几何-电磁线索实现对关键成像参数的可控生成。
Comments 23 pages,14 figures
通过组合并行令牌预测实现可控图像生成
机构 * Durham University(杜伦大学)
专题命中 可控生成 :image generation(title,abstract);text-to-image(abstract);diffusion(abstract)
AI总结 本文提出一种理论指导的离散生成过程组合方法,通过掩码生成(吸收扩散)实现多输入条件的精确组合,相比现有方法在误差率和FID上均有显著提升,并实现高效的文本到图像生成控制。
Comments 8 pages + references, 7 figures, accepted to CVPR Workshops 2026 (LoViF). arXiv admin note: substantial text overlap with arXiv:2405.06535
面向边缘计算的语义感知缓存用于高效图像生成
专题命中 可控生成 :image generation(title,abstract);text-to-image(abstract);diffusion(abstract)
AI总结 CacheGenius通过语义感知缓存和调度算法,在边缘计算中高效生成图像,减少41%的延迟和48%的计算成本。
专题命中 可控生成 :diffusion(title,abstract);text-to-image(abstract);分类 cs.CV、cs.GR
Comments Project page at https://anylens-diffusion.github.io/
专题命中 可控生成 :diffusion(title);image synthesis(title);分类 cs.CV
Journal ref Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 2024, pp. 5354-5363
Pro-DG:基于过程扩散引导的建筑立面生成
机构 * Warsaw University of Technology(华沙技术大学) ; Akces NCBR ; Imperial College London(伦敦帝国理工学院) ; New Jersey Institute of Technology(新泽西理工学院)
专题命中 可控生成 :diffusion(title,abstract);inpainting(abstract);分类 cs.CV、cs.GR
AI总结 本文提出Pro-DG方法,通过层次化过程规则生成控制图,实现建筑立面的真实感生成,支持结构编辑如楼层复制和窗户重排,经评估验证其在保持建筑身份和精确编辑方面的优越性。
Comments 17 pages, 15 figures, Computer Graphics Forum 2026 Journal Paper
具有风格引导的舞动生成
专题命中 可控生成 :diffusion(title,abstract);inpainting(abstract);分类 cs.CV、cs.MM
AI总结 本文提出Style-Guided Motion Diffusion方法,通过整合Transformer架构与风格调制模块,实现风格引导的舞蹈生成,提升舞蹈生成的可控性和风格一致性。
Journal ref Machine Intelligence Research, 2026
秩序并非布局:图像生成中的秩序到空间偏差
机构 * Renmin University of China, China(中国人民大学) ; Huazhong Agricultural University, China(华中农业大学) ; Huazhong University of Science and Technology, China(华中科技大学) ; Jiangnan University, China(江南大学) ; Nanchang University, China(南昌大学)
专题命中 可控生成 :image generation(title,abstract);text-to-image(abstract);分类 cs.CV、cs.MM
AI总结 本文研究了图像生成中因文本实体顺序导致的空间布局偏差问题,提出OTS-Bench进行量化评估,并通过微调和干预策略减少该偏差。
机构 * Yale University(耶鲁大学) ; Brown University(布朗大学) ; Chung-Ang University(Chung-Ang 大学)
专题命中 可控生成 :diffusion(title,abstract);text-to-image(abstract);分类 cs.CV、cs.MM
机构 * SECE, Peking University(北京大学SECE学院) ; The Chinese University of Hong Kong(香港中文大学) ; ARC Lab, Tencent(腾讯ARC实验室) ; The Hong Kong University of Science and Technology(香港科学与技术大学) ; GVC Lab, Great Bay University(Great Bay大学GVC实验室)
专题命中 可控生成 :image editing(title,abstract);diffusion(abstract);分类 cs.CV、cs.MM
Comments Project Webpage: https://liyaowei-stu.github.io/project/BlobCtrl/ This version presents a major update with rephrased writing. Accepted to SIGGRAPH Asia 2025
机构 * NVIDIA ; University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) ; University of Toronto(多伦多大学) ; Vector Institute(向量研究所)
专题命中 可控生成 :diffusion(title,abstract);image editing(abstract);分类 cs.CV、cs.GR
Comments International Conference on Computer Vision (ICCV) 2025, Project Website: https://research.nvidia.com/labs/toronto-ai/WeatherWeaver/
专题命中 可控生成 :image generation(title,abstract);diffusion(abstract);分类 cs.CV、cs.GR
Comments 10 pages, 7 figures, in Proceedings of CAADRIA2025
专题命中 可控生成 :image generation(title,abstract);inpainting(abstract);分类 cs.CV、cs.GR
Comments 18 pages, 15 figures, project page: https://anywheremultiagent.github.io, Accepted at AAAI 2025
专题命中 可控生成 :image generation(title,abstract);diffusion(abstract);分类 cs.CV、cs.GR
Comments ACM Transactions on Applied Perception (ACM Symposium on Applied Perception 2024)
专题命中 可控生成 :diffusion(title,abstract);image synthesis(abstract);分类 cs.CV、cs.GR
专题命中 可控生成 :image generation(title);text-to-image(abstract);diffusion(abstract);分类 cs.CV、cs.MM
Comments Accepted by ACM MM24
专题命中 可控生成 :image generation(title,abstract);image synthesis(abstract);分类 cs.CV、cs.MM
Comments arXiv admin note: substantial text overlap with arXiv:2012.03308
专题命中 可控生成 :image generation(title,abstract);image synthesis(abstract);分类 cs.CV、cs.MM
Comments CVPR 2021. Code: https://github.com/weihaox/TediGAN Data: https://github.com/weihaox/Multi-Modal-CelebA-HQ Video: https://youtu.be/L8Na2f5viAM
专题命中 可控生成 :diffusion(title,abstract);image generation(abstract);分类 cs.CV
Comments Project page: https://lixirui142.github.io/unicon-diffusion/
专题命中 可控生成 :image generation(title);text-to-image(title)
Comments 4 pages, 2 figures, to be published in IEEE International Conference on Sensing, Communication, and Networking, Workshop on Semantic Communication for 6G (SC6G-SECON23)
超越《星夜》:面向艺术家基础的文本到图像生成的捷径感知控制状态规划
机构 * Jilin University(吉林大学) ; Adobe(奥多比公司) ; University of Wisconsin(威斯康星大学)
专题命中 可控生成 :image generation(title,abstract);image synthesis(abstract);分类 cs.CV
AI总结 该研究针对艺术家基础的文本到图像生成存在的捷径偏差问题,提出 Atelier 框架,结合 ArtIntentBench 基准测试,提升了风格保真度与结构保留度,减少了捷径替换。
Comments 47 pages, 13 figures, including appendices. Kuan Xing and Ye Wang contributed equally