arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

2026-02-03 至 2026-02-03 共收录 17 信号源:cs.CV, cs.GR, cs.MM

1. 文生图 17 篇

2601.11651 2026-02-03 cs.CV cs.AI cs.CY 88%

Aesthetics as Structural Harm: Algorithmic Lookism Across Text-to-Image Generation and Classification

审美作为结构性伤害:跨文本到图像生成和分类的算法审美主义

Miriam Doh, Aditya Gulati, Corinna Canali, Nuria Oliver

机构 * Université Libre de Bruxelles(布鲁塞尔自由大学) Weizenbaum Institute(魏泽曼研究所) ELLIS Alicante(阿利坎特ELLIS)

专题命中 文生图 :text-to-image(title,abstract);image generation(title);diffusion(abstract);分类 cs.CV

AI总结 本文揭示了生成式AI在文本到图像生成和性别分类中系统性地将面部吸引力与积极属性关联,导致性别偏见和审美歧视。

Comments 22 pages, 15 figures; v2 - fix typo

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13418 2026-02-03 cs.CV 88%

Reinforcement Learning Meets Masked Generative Models: Mask-GRPO for Text-to-Image Generation

强化学习遇见掩码生成模型:Mask-GRPO用于文本到图像生成

Yifu Luo, Xinhao Hu, Keyu Fan, Haoyuan Sun, Zeyu Chen, Bo Xia, Tiantian Zhang, Yongzhe Chang, Xueqian Wang

机构 * Tsinghua University(清华大学)

专题命中 文生图 :text-to-image(title,abstract);image generation(title);diffusion(abstract);分类 cs.CV

AI总结 本文提出Mask-GRPO,通过重新定义转移概率和多步骤决策问题,改进文本到图像生成模型,实现优于现有方法的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.08396 2026-02-03 cs.CV 88%

CoDi: Subject-Consistent and Pose-Diverse Text-to-Image Generation

CoDi: 一致主体和多姿态文本到图像生成

Zhanxin Gao, Beier Zhu, Liang Yao, Jian Yang, Ying Tai

机构 * Nanjing University(南京大学) Nanyang Technological University(南洋理工大学) Vipshop

专题命中 文生图 :text-to-image(title,abstract);image generation(title);diffusion(abstract);分类 cs.CV

AI总结 CoDi通过两阶段策略实现文本到图像生成中主体一致性和姿态多样性的平衡,提升视觉表现和性能。

Comments Accepted by ICLR2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01370 2026-02-03 cs.CV 88%

PointT2I: LLM-based text-to-image generation via keypoints

PointT2I: 通过关键点生成基于大语言模型的文本到图像生成

Taekyung Lee, Donggyu Lee, Myungjoo Kang

机构 * Interdisciplinary Program in Artificial Intelligence, Seoul National University(人工智能跨学科项目,首尔国立大学) Department of Mathematical Sciences, Seoul National University(数学科学系,首尔国立大学)

专题命中 文生图 :image generation(title,abstract);text-to-image(title,abstract);分类 cs.CV

AI总结 PointT2I通过大语言模型生成关键点并指导图像生成,无需微调即可实现基于文本提示的准确姿态对齐图像生成。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01382 2026-02-03 cs.CV cs.LG 85%

PromptRL: Prompt Matters in RL for Flow-Based Image Generation

PromptRL: 流基于图像生成中强化学习中提示的重要性

Fu-Yun Wang, Han Zhang, Michael Gharbi, Hongsheng Li, Taesung Park

机构 * The Chinese University of Hong Kong, Hong Kong(香港中文大学) Meta Superintelligence Labs, USA(Meta超智能实验室)

专题命中 文生图 :image generation(title,abstract);text-to-image(abstract);image editing(abstract);分类 cs.CV

AI总结 PromptRL通过整合语言模型作为可训练的提示精修代理,提升流基于图像生成中强化学习的性能和效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01756 2026-02-03 cs.CV 83%

Mind-Brush: Integrating Agentic Cognitive Search and Reasoning into Image Generation

Mind-Brush: 将代理认知搜索与推理整合到图像生成中

Jun He, Junyan Ye, Zilong Huang, Dongzhi Jiang, Chenjue Zhang, Leqi Zhu, Renrui Zhang, Xiang Zhang, Weijia Li

机构 * Sun Yat-sen University(中山大学) Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) MMLab, The Chinese University of Hong Kong(香港中文大学MMLab)

专题命中 文生图 :image generation(title,abstract);text-to-image(abstract);分类 cs.CV

AI总结 Mind-Brush通过整合代理认知搜索与推理,提升图像生成模型对复杂知识推理和动态适应能力,实现性能显著跃升。

Comments 36 pages, 24 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02295 2026-02-03 cs.CV 83%

Data-Driven Loss Functions for Inference-Time Optimization in Text-to-Image

面向文本到图像推理时间优化的数据驱动损失函数

Sapir Esther Yiflach, Yuval Atzmon, Gal Chechik

机构 * Bar-Ilan University(巴伊兰大学) NVIDIA(英伟达)

专题命中 文生图 :text-to-image(title,abstract);diffusion(abstract);分类 cs.CV

AI总结 本文提出 Learn-to-Steer 框架,通过学习数据驱动的损失函数提升文本到图像模型的空间推理能力,显著提高多个基准测试的准确性。

Comments Project page is at https://learn-to-steer-paper.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03199 2026-02-03 cs.CL 82%

Beyond Content: How Grammatical Gender Shapes Visual Representation in Text-to-Image Models

超越内容:语法性别如何塑造文本到图像模型中的视觉表示

Muhammed Saeed, Shaina Raza, Ashmal Vayani, Muhammad Abdul-Mageed, Ali Emami, Shady Shehata

专题命中 文生图 :text-to-image(title,abstract);image generation(abstract)

AI总结 研究揭示语法性别对文本到图像模型中视觉表示的显著影响,通过跨语言实验展示不同语法性别对性别表示的系统性影响。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02227 2026-02-03 cs.CV 79%

Show, Don't Tell: Morphing Latent Reasoning into Image Generation

展示,而非告知:将潜在推理转化为图像生成

Harold Haodong Chen, Xinxiang Yin, Wen-Jie Shu, Hongfei Zhang, Zixin Zhang, Chenfei Liao, Litao Guo, Qifeng Chen, Ying-Cong Chen

专题命中 文生图 :image generation(title);text-to-image(abstract);分类 cs.CV

AI总结 LatentMorph通过隐式潜在推理提升文本到图像生成的效率和效果,减少推理时间与标记消耗,同时提高与人类直觉的一致性。

Comments Code: https://github.com/EnVision-Research/LatentMorph

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01193 2026-02-03 cs.CL cs.CV 77%

Bridging Lexical Ambiguity and Vision: A Mini Review on Visual Word Sense Disambiguation

弥合词汇歧义与视觉:关于视觉词义消歧的简要综述

Shashini Nilukshi, Deshan Sumanathilaka

机构 * School of Computing Informatics(计算与信息学学院) Institute of Technology(技术研究所)

专题命中 文生图 :image generation(abstract);text-to-image(abstract);diffusion(abstract);分类 cs.CV

AI总结 本文综述了视觉词义消歧的发展,探讨了对比模型和LLM在解决词汇歧义中的作用,并指出未来发展方向。

Comments 2 figures, 2 Tables, Accepted at IEEE TIC 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22400 2026-02-03 cs.CV 77%

Closing the Safety Gap: Surgical Concept Erasure in Visual Autoregressive Models

弥合安全鸿沟:视觉自回归模型中的手术概念擦除

Xinhao Zhong, Yimin Zhou, Zhiqi Zhang, Junhao Li, Yi Sun, Bin Chen, Shu-Tao Xia, Xuan Wang, Ke Xu

机构 * Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳)) Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院) Jilin University(吉林大学) Peng Cheng Laboratory(鹏城实验室) Department of Computer Science and Technology, Tsinghua University(清华大学计算机科学与技术系)

专题命中 文生图 :image generation(abstract);text-to-image(abstract);diffusion(abstract);分类 cs.CV

AI总结 本文提出S-VARE方法,通过过滤交叉熵损失和保留损失实现VAR模型中的精确概念擦除,有效解决安全性和生成质量之间的平衡问题。

Comments Accepted toICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02171 2026-02-03 cs.CV 74%

Lung Nodule Image Synthesis Driven by Two-Stage Generative Adversarial Networks

由双阶段生成对抗网络驱动的肺结节图像合成

Lu Cao, Xiquan He, Junying Zeng, Chaoyun Mai, Min Luo

专题命中 文生图 :image synthesis(title);分类 cs.CV

AI总结 本文提出双阶段生成对抗网络TSGAN,通过解耦肺结节形态结构与纹理特征,提升合成数据多样性及检测模型性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00092 2026-02-03 cs.LG cs.AI cs.CL cs.CV 70%

Interpreting and Controlling Model Behavior via Constitutions for Atomic Concept Edits

通过宪法进行原子概念编辑来解释和控制模型行为

Neha Kalibhat, Zi Wang, Prasoon Bajpai, Drew Proud, Wenjun Zeng, Been Kim, Mani Malek

机构 * Google DeepMind(谷歌DeepMind)

专题命中 文生图 :image generation(abstract);text-to-image(abstract);分类 cs.CV

AI总结 该研究提出通过原子概念编辑学习模型宪法,以解释和控制模型行为,实验证明其在提升模型成功率方面效果显著。

Journal ref Twenty-Ninth Annual Conference on Artificial Intelligence and Statistics (AISTATS 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00849 2026-02-03 cs.LG cs.AI cs.NA math.NA 67%

RMFlow: Refined Mean Flow by a Noise-Injection Step for Multimodal Generation

RMFlow:通过噪声注入步骤细化均流以实现多模态生成

Yuhao Huang, Shih-Hsin Wang, Andrea L. Bertozzi, Bao Wang

机构 * Department of Mathematics and Scientific Computing and Imaging (SCI) Institute University of Utah(数学与科学计算及成像学院(SCI)院,犹他大学) Department of Mathematics, UCLA(数学系,加州大学洛杉矶分校)

专题命中 文生图 :image generation(abstract);text-to-image(abstract)

AI总结 RMFlow通过引入噪声注入步骤,改进均流模型,实现高效多模态生成,仅需单次功能评估即可达到接近最先进的性能。

Comments Accepted to ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12946 2026-02-03 cs.CY cs.AI cs.CL cs.CV cs.LG 57%

AI-generated data contamination erodes pathological variability and diagnostic reliability

AI生成的数据污染侵蚀病理变异性与诊断可靠性

Hongyu He, Shaowen Xiang, Ye Zhang, Yingtao Zhu, Jin Zhang, Hao Deng, Emily Alsentzer, Yun Liu, Qingyu Chen, Kun-Hsing Yu, Andrew Marshall, Tingting Chen, Srinivas Anumasa, Daniel Ebner, Dean Ho, Kee Yuan Ngiam, Ching-Yu Cheng, Dianbo Liu

专题命中 文生图 :image synthesis(abstract);分类 cs.CV

AI总结 AI生成的数据污染导致病理变异性下降和诊断可靠性降低,需政策强制的人类监督以防止医疗数据生态系统的退化。

Comments *Corresponding author: Dianbo Liu (dianbo@nus.edu.sg)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04236 2026-02-03 cs.CV 57%

Scaling Sequence-to-Sequence Generative Neural Rendering

序列到序列生成神经渲染的扩展

Shikun Liu, Kam Woh Ng, Wonbong Jang, Jiadong Guo, Junlin Han, Haozhe Liu, Yiannis Douratsos, Juan C. Pérez, Zijian Zhou, Chi Phung, Tao Xiang, Juan-Manuel Pérez-Rúa

机构 * Meta AI

专题命中 文生图 :image synthesis(abstract);分类 cs.CV

AI总结 Kaleido通过统一的解码器-only rectified flow transformer实现了序列到序列生成神经渲染的扩展,能够在无显式3D表示的情况下生成多视角视图,并在多个视图设置中达到与场景优化方法相当的性能。

Comments Published at ICLR 2026. Project Page: https://shikun.io/projects/kaleido

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00221 2026-02-03 eess.IV cs.CV 57%

Benchmarking Vanilla GAN, DCGAN, and WGAN Architectures for MRI Reconstruction: A Quantitative Analysis

对MRI重建中Vanilla GAN、DCGAN和WGAN架构进行基准测试:一项定量分析

Humaira Mehwish, Hina Shakir, Muneeba Rashid, Asarim Aamir, Reema Qaiser Khan

专题命中 文生图 :image synthesis(abstract);分类 cs.CV

AI总结 本研究通过定量分析比较了Vanilla GAN、DCGAN和WGAN在MRI重建中的性能,发现DCGAN和WGAN在图像质量和准确性方面表现更优。

Comments 20 pages

Journal ref Edelweiss Applied Science and Technology January 2026

详情

展开后加载摘要…

URL PDF HTML 收藏