InfiniCity: Infinite-Scale City Synthesis
专题命中 文生图 :image synthesis(abstract);分类 cs.CV、cs.GR
视觉与机器人
图像生成、文生图、图像编辑、扩散模型和可控生成。
专题命中 文生图 :image synthesis(abstract);分类 cs.CV、cs.GR
专题命中 文生图 :text-to-image(abstract);分类 cs.CV、cs.MM
Comments Accepted at ACM Multimedia 2023
专题命中 文生图 :image synthesis(abstract);分类 cs.CV、cs.GR
专题命中 文生图 :text-to-image(abstract);分类 cs.CV、cs.MM
专题命中 文生图 :image synthesis(abstract);分类 cs.CV、cs.GR
Comments Project page: https://junxuan-li.github.io/megane/
专题命中 文生图 :image synthesis(abstract);分类 cs.CV、cs.GR
Comments ECCV 2022 paper. 14 pages of main content, 4 pages of references, and 11 pages of appendix
专题命中 文生图 :image synthesis(abstract);分类 cs.CV、cs.GR
Comments accepted to ECCV2022; code available at http://github.com/zhuhao-nju/mofanerf
专题命中 文生图 :image synthesis(abstract);分类 cs.CV、cs.GR
Comments Accepted to NeurIPS 2020
Journal ref Advances in Neural Information Processing Systems 33 (NeurIPS 2020), 9841-9850
专题命中 文生图 :image synthesis(abstract);分类 cs.CV、cs.GR
Comments 14 pages including refences
专题命中 文生图 :text-to-image(abstract);分类 cs.CV、cs.MM
专题命中 文生图 :image synthesis(abstract);分类 cs.CV、cs.GR
专题命中 文生图 :image synthesis(abstract);分类 cs.CV、cs.GR
Comments 16 pages, 10 figures
专题命中 文生图 :image synthesis(abstract);分类 cs.CV、cs.GR
Comments ICCV 2021 (oral); Project page: https://infinite-nature.github.io/; Video: https://www.youtube.com/watch?v=oXUf6anNAtc
专题命中 文生图 :image synthesis(abstract);分类 cs.CV、cs.GR
Comments SIGGRAPH Asia 2021 Technical Paper. Code: https://github.com/yizhiwang96/deepvecfont ; Homepage: https://yizhiwang96.github.io/deepvecfont_homepage/
专题命中 文生图 :image synthesis(abstract);分类 cs.CV、cs.GR
Comments Accepted to UIST2021. Project page: https://sites.google.com/view/deepmannequin/home
专题命中 文生图 :image synthesis(abstract);分类 cs.CV、cs.GR
专题命中 文生图 :image synthesis(abstract);分类 cs.CV、cs.GR
Comments Project Page: https://grail.cs.washington.edu/projects/vid2actor/ Supplementary Video: https://youtu.be/Zec8Us0v23o
专题命中 文生图 :image synthesis(abstract);分类 cs.CV、cs.GR
Comments Leonardo / SIGGRAPH 2020 Art Papers
Journal ref Leonardo, Volume 53, Issue 4, August 2020
专题命中 文生图 :image synthesis(abstract);分类 cs.CV、cs.GR
专题命中 文生图 :text-to-image(abstract);分类 cs.CV、cs.MM
专题命中 文生图 :image synthesis(abstract);分类 cs.CV、cs.GR
Comments Published at SIGGRAPH Asia 2018 (ACM Transactions on Graphics). Project page with codes, pretrained models, and human model lists is at http://kanamori.cs.tsukuba.ac.jp/projects/relighting_human/
专题命中 文生图 :image synthesis(abstract);分类 cs.CV、cs.GR
Comments Symposium on Geometry Processing 2019
Journal ref Computer Graphics Forum 38 (5), 2019
专题命中 文生图 :image synthesis(abstract);分类 cs.CV、cs.MM
Comments CVPR-19 Workshop on Computer Vision: Challenges and Opportunities for Privacy and Security (CV-COPS 2019)
专题命中 文生图 :image synthesis(abstract);分类 cs.CV、cs.GR
Comments In NeurIPS, 2018. Code, models, and more results are available at https://github.com/NVIDIA/vid2vid
专题命中 文生图 :text-to-image(abstract);分类 cs.CV;diffusion(comments)
Comments Accepted to DICTA 2022, released 11000+ environmental scene images generated by Stable Diffusion and 1000+ images generated by DALLE-2
SynerMedGen:通过任务对齐协同医疗多模态理解与生成
机构 * The University of Hong Kong, Hong Kong, China(香港大学)
专题命中 文生图 :image synthesis(abstract);分类 cs.CV
AI总结 SynerMedGen通过任务对齐实现医疗多模态理解与生成的协同,提出生成对齐理解任务和两阶段训练策略,实现零样本性能和跨数据集泛化,释放大规模数据集支持进一步研究。
Comments Accepted by ICML 2026
EviRank:用于多模态图像重排序的结构化相关性证据
机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) ; Tencent Yuanbao(腾讯元宝) ; The University of Hong Kong(香港大学)
专题命中 文生图 :text-to-image(abstract);分类 cs.CV
AI总结 针对现有多模态图像重排序器的不足,提出EviRank将查询解析为结构化证据包,通过证据条件验证实现重排序,在五个基准上达SOTA,蒸馏学生模型保留超90%能力且成本更低。
新概念必须进入之处:统一多模态模型中的入口门跨任务可用性
机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) ; Columbia University(哥伦比亚大学) ; CUHK(香港中文大学) ; University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
专题命中 文生图 :text-to-image(abstract);分类 cs.CV
AI总结 该研究通过分离统一多模态模型的理解与生成任务方向,发现跨任务可用性取决于概念绑定的入口层,提出的对齐目标可在极低损失下实现概念跨任务迁移。
Comments 27 pages, 10 figures
JoLT:用于上下文引导的高分辨率分块生成的联合潜在轨迹
机构 * Obvious Research(奥布弗西斯研究公司) ; Sorbonne Université(索邦大学)
专题命中 文生图 :text-to-image(abstract);分类 cs.CV
AI总结 本文提出JoLT方法,通过联合去噪LR与HR潜在图像生成高分辨率图像,其生成的图像细节丰富、视觉效果佳,优于竞争基线,为艺术创作提供新方向。
Comments 25 pages, 10 figures, 7 tables. Accepted at the AI4VA Workshop at ECCV 2026. Project page: https://obvious-research.github.io/jolt/
歧义去了哪里?探究多模态模型如何解释多义词
机构 * Princeton University(普林斯顿大学)
专题命中 文生图 :text-to-image(abstract);分类 cs.CV
AI总结 该研究对比17个文本到图像模型和15个文本生成模型,发现多模态模型生成图像的词义多样性低于文本,揭示了基础模型在不同模态间意义表达的迁移 gap。
Comments Oral Presentation, Sci-FM Workshop @ COLM 2026