arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

2025-12-03 至 2025-12-03 共收录 2 信号源:cs.CV, cs.GR, cs.MM

1. 文生图 2 篇

2512.02794 2025-12-03 cs.CV 88%

PhyCustom: Towards Realistic Physical Customization in Text-to-Image Generation

PhyCustom: 向文本到图像生成中的真实物理定制迈进

Fan Wu, Cheng Chen, Zhoujie Fu, Jiacheng Wei, Yi Xu, Deheng Ye, Guosheng Lin

机构 * Nanyang Technological University(南洋理工大学) Goertek Alpha Labs(歌尔声学Alpha实验室) Tencent(腾讯)

专题命中 文生图 :text-to-image(title,abstract);image generation(title);diffusion(abstract);分类 cs.CV

AI总结 PhyCustom通过引入两个新正则化损失,提升文本到图像生成中对物理概念的真实定制能力。

Comments codes:https://github.com/wufan-cse/PhyCustom

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02161 2025-12-03 cs.CV 83%

FineGRAIN: Evaluating Failure Modes of Text-to-Image Models with Vision Language Model Judges

FineGRAIN:通过视觉语言模型法官评估文本到图像模型的故障模式

Kevin David Hayes, Micah Goldblum, Vikash Sehwag, Gowthami Somepalli, Ashwinee Panda, Tom Goldstein

机构 * University of Maryland(马里兰大学) Columbia University(哥伦比亚大学) Sony AI(索尼人工智能)

专题命中 文生图 :text-to-image(title,abstract);image generation(abstract);分类 cs.CV

AI总结 FineGRAIN通过视觉语言模型评估文本到图像模型的故障模式,揭示属性保真度和物体表示的系统性错误,强调了针对性基准测试对生成模型可靠性的重要性。

Comments Accepted to NeurIPS 2025 Datasets and Benchmarks Track

详情

展开后加载摘要…

URL PDF HTML 收藏