ELODIN: Naming Concepts in Embedding Spaces
专题命中 文生图 :text-to-image(abstract);image synthesis(abstract);分类 cs.CV、cs.GR
Comments Added quantitative data, fixed formatting issues
视觉与机器人
图像生成、文生图、图像编辑、扩散模型和可控生成。
专题命中 文生图 :text-to-image(abstract);image synthesis(abstract);分类 cs.CV、cs.GR
Comments Added quantitative data, fixed formatting issues
专题命中 文生图 :image generation(abstract);text-to-image(abstract);分类 cs.CV、cs.MM
专题命中 文生图 :image generation(abstract);text-to-image(abstract);分类 cs.CV、cs.MM
Comments This paper has 29 pages with 22 figures, including rich supplementary information. Project page is at \url{https://classifier-as-generator.github.io/}
专题命中 文生图 :text-to-image(abstract);image synthesis(abstract);分类 cs.CV、cs.GR
Comments Accepted at ICCC 2022
专题命中 文生图 :image generation(abstract);text-to-image(abstract);分类 cs.CV、cs.MM
Comments ACM MM 2021 (Industrial Track). Code: https://github.com/researchmm/generate-it
专题命中 文生图 :inpainting(abstract);image synthesis(abstract);分类 cs.CV、cs.GR
Journal ref Sensors 2021, 21, 4032
专题命中 文生图 :image generation(abstract);image synthesis(abstract);分类 cs.CV、cs.GR
Comments Project page: https://choyingw.github.io/works/Voice2Mesh/index.html
专题命中 文生图 :image generation(abstract);image synthesis(abstract);分类 cs.CV、cs.GR
机构 * Department of CSE University of Texas at Arlington Arlington, Texas, USA(计算机科学与工程系 乌德勒支理工大学 美国德克萨斯州阿灵顿)
专题命中 文生图 :image synthesis(title)
专题命中 文生图 :text-to-image(title)
Comments Upcoming Publication, AIES 2025
机构 * Institute of Software, Chinese Academy of Sciences(中国科学院软件研究所) ; University of California at Merced(加州大学默塞德分校) ; University of Chinese Academy of Sciences(中国科学院大学) ; The University of Sydney(悉尼大学)
专题命中 文生图 :text-to-image(title)
Comments 14 pages,8 figures,4 tables
专题命中 文生图 :text-to-image(title)
Comments 15 pages, 5 figures. To be published as full paper in the Proceedings of the European Conference on Information Retrieval (ECIR) 2025
专题命中 文生图 :text-to-image(title)
专题命中 文生图 :image synthesis(title)
Journal ref Geophysical Journal International, 204, 1179-1190 (2016)
用于多模态理解与生成的解耦式视觉-语言系统
机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) ; University of Chinese Academy of Sciences(中国科学院大学) ; Peng Cheng Laboratory(鹏城实验室) ; Li Auto(理想汽车)
专题命中 文生图 :image generation(abstract);text-to-image(abstract);分类 cs.CV
AI总结 本研究提出多模态大语言模型Libra的解耦式视觉-语言架构,通过开关注意力与FFN模块实现自模态与跨模态解耦,在理解与生成任务上均取得优异性能。
面向物理保真的科学图表生成
机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) ; Shanghai Jiao Tong University(上海交通大学)
专题命中 文生图 :image generation(abstract);text-to-image(abstract);分类 cs.CV
AI总结 针对现有文本到图像生成模型生成的科学图表存在物理错误的问题,提出Princigram生成器,通过结构化物理思维链实现物理保真的科学图表生成,在相关基准上验证了其有效性。
无辜的面板,仇恨的故事:评估与检测多轮视觉故事生成中的仇恨意图
机构 * CISPA Helmholtz Center for Information Security(CISPA亥姆霍兹信息安全中心) ; Hewlett Packard Enterprise(惠普企业)
专题命中 文生图 :image generation(abstract);text-to-image(abstract);分类 cs.CV
AI总结 本研究针对多轮视觉故事生成的组级仇恨意图问题,构建了专用评估数据集,发现现有审核系统存在漏检,提出互补防御方法,强调安全需适配视觉叙事的发展。
Comments 16 pages, 5 figures, 10 tables
危害并非普遍存在:迫切需要针对特定社区的毒性检测
机构 * Microsoft Research Cambridge, UK(英国剑桥微软研究院) ; Microsoft Paris, France(法国巴黎微软)
专题命中 文生图 :image generation(abstract);text-to-image(abstract);分类 cs.CV
AI总结 研究指出文本到图像生成的通用毒性检测方法不能保护边缘化社区,主张特定社区毒性检测(CTD)。通过与专家合作制定指南,利用图像数据集实验表明现有模型表现不佳,基于提示的方法和参数高效微调可提升性能,但CTD性能仍远低于通用检测,需持续研究。
Comments 18 pages, under review
CLIP 所知道但无法表达的:从冻结的中间特征中恢复否定信息
机构 * Purdue University(普渡大学)
专题命中 文生图 :image generation(abstract);text-to-image(abstract);分类 cs.CV
AI总结 对比视觉语言模型如CLIP对否定不敏感,研究发现是表征坍缩所致。提出轻量级事后校正系统PeakPatch,在不改变预训练权重下恢复否定信号,通过特定网络提取信号、预测偏差向量等,经实验验证其有效性及泛化性。
SuperFlow: 通过实时强化学习训练流匹配模型
专题命中 文生图 :image generation(abstract);text-to-image(abstract);分类 cs.CV
AI总结 SuperFlow通过实时强化学习优化流模型训练,提升文本-图像生成效率和质量。
Comments This article is withdrawn because it was submitted to arXiv without obtaining the consent of all listed authors
SPQR:良性模型适应下安全对齐的多维基准测试
机构 * Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学) ; University of Waterloo(滑铁卢大学) ; Michigan State University(密歇根州立大学)
专题命中 文生图 :text-to-image(abstract);diffusion(abstract);分类 cs.CV
AI总结 研究文本到图像扩散模型安全对齐在良性微调下的稳定性,引入SPQR基准测试,通过单评分指标提供统一框架,经多方面分析确定安全对齐失败情况,为T2I安全对齐技术提供简洁全面的基准。
Comments 34 pages, 9 figures, 13 tables
交错思维中的模态隔离桥接:通过逐步强化监督模态转换
机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) ; Shanghai Jiaotong University(上海交通大学) ; Zhejiang University(浙江大学) ; University of Chinese Academy of Sciences(中国科学院大学)
专题命中 文生图 :image generation(abstract);text-to-image(abstract);分类 cs.CV
AI总结 提出MoTiF框架,通过反射式SFT和Flow-GRPO优化模态转换保真度,解决交错思维中图像与文本脱节的模态隔离问题,提升跨模态一致性和任务准确性。
Comments 22 pages, 5 figures, 6 tables
LEMUR 2:释放人工智能的神经网络多样性
机构 * University of Würzburg(维尔茨堡大学)
专题命中 文生图 :text-to-image(abstract);image synthesis(abstract);分类 cs.CV
AI总结 LEMUR 2旨在释放神经网络多样性,通过统一生成、评估和部署管道,利用多种方式生成超14000个不同架构及大量训练记录,采用特定管道进行自动部署和延迟基准测试,涵盖多模态任务,为LLM微调提供数据基础,推动相关新兴范式发展。
Comments 10 pages, 9 figures, 1 table
分布匹配变分自编码器
专题命中 文生图 :diffusion(abstract);image synthesis(abstract);分类 cs.CV
AI总结 本文提出DMVAE,通过分布匹配约束将编码器的潜在分布与任意参考分布对齐,发现基于自监督学习的分布在图像生成中表现优异。
Comments ICML2026
思维形状:通过视觉思维链进行渐进式物体组装
机构 * School of Science and Engineering, The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)科学与工程学院) ; School of Data Science, The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)数据科学学院) ; Sun Yat-sen University(中山大学) ; The Hong Kong University of Science and Technology, Guangzhou(香港科学与技术大学(广州)) ; Shenzhen Future Network of Intelligence Institute (FNii-Shenzhen)(深圳未来网络智能研究所(FNii-Shenzhen)) ; Guangdong Provincial Key Laboratory of Future Networks of Intelligence, CUHK(SZ)(广东省未来网络智能重点实验室,CUHK(SZ))
专题命中 文生图 :image generation(abstract);text-to-image(abstract);分类 cs.CV
AI总结 提出Shape-of-Thought (SoT)框架,通过视觉思维链在渲染2D域中逐步组装形状,解决文本到图像生成中的组合结构约束问题,在组件计数和结构拓扑上显著优于直接生成。
Comments ICML2026
RUB: 评估未学习模型中的残留知识
机构 * Electrical and Computer Engineering University of Alberta(电气与计算机工程大学阿尔伯塔大学)
专题命中 文生图 :text-to-image(abstract);image synthesis(abstract);分类 cs.CV
AI总结 提出鲁棒未学习原则及统一基准RUB,通过未学习映射攻击(UMA)检测残留信息,揭示现有方法在对抗评估下的脆弱性。
Journal ref Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, June 2026, pages 8550-8559
CheXGenBench:合成胸片保真度、隐私和实用性的统一基准
机构 * University of Edinburgh(爱丁堡大学) ; Samsung AI Center, Cambridge(剑桥三星AI中心)
专题命中 文生图 :image generation(abstract);text-to-image(abstract);分类 cs.CV
AI总结 提出CheXGenBench,首个统一评估框架,同时衡量合成胸片生成模型的保真度、隐私风险和下游实用性,涵盖11种前沿T2I模型,揭示当前模型在长尾分布、隐私风险和下游多模态任务中的局限。
Comments Published in Transactions of Machine Learning Research (06/2026)
Journal ref Transactions on Machine Learning Research (2026)
模态强制实现可扩展的空间生成
机构 * Carnegie Mellon University(卡内基梅隆大学) ; World Labs
专题命中 文生图 :image generation(abstract);text-to-image(abstract);分类 cs.CV
AI总结 提出Modality Forcing方法,通过为每个模态分配独立噪声水平,实现单DiT的联合图像-深度生成,利用稀疏深度数据训练,继承T2I预训练的可扩展性,在深度估计上取得竞争性能。
跨模态掩码组合概念建模以增强视觉-语言组合性
机构 * MoE Key Laboratory of Brain-inspired Intelligent Perception and Cognition, University of Science and Technology of China(中国科学技术大学,教育部脑启发智能感知与认知重点实验室) ; Independent Researcher(独立研究员)
专题命中 文生图 :image generation(abstract);text-to-image(abstract);分类 cs.CV
AI总结 提出MACCO框架,通过掩码一个模态的组合概念并从另一模态完整上下文重建,增强视觉-语言模型的组合理解能力,在五个基准上显著提升。
Comments Accepted to ACL 2026 Main Conference, 25 pages
ARM: 一种具有统一离散表示的自回归大型多模态模型
机构 * Shanghai Key Lab of Intelligent Information Processing, Fudan University(复旦大学上海智能信息处理重点实验室) ; School of Computer Science, Fudan University(复旦大学计算机科学技术学院) ; Shanghai Collaborative Innovation Center of Intelligent Visual Computing(上海智能视觉计算协同创新中心) ; Youtu Lab, Tencent(腾讯优图实验室) ; Meta AI ; Shanghai AI Laboratory(上海人工智能实验室)
专题命中 文生图 :image generation(abstract);text-to-image(abstract);分类 cs.CV
AI总结 提出ARM模型,通过离散语义视觉分词器将图像映射为紧凑token序列,结合自回归建模和强化学习,统一实现图像理解、生成和编辑,并提升任务性能与跨任务协同。
Comments technical report