Generative Model via Quantile Assignment
通过分位数分配生成模型
专题命中 文生图 :diffusion(abstract);image synthesis(abstract)
AI总结 NeuroSQL通过隐式学习低维潜在表示,无需辅助网络,实现高效稳定的合成数据生成。
视觉与机器人
图像生成、文生图、图像编辑、扩散模型和可控生成。
通过分位数分配生成模型
专题命中 文生图 :diffusion(abstract);image synthesis(abstract)
AI总结 NeuroSQL通过隐式学习低维潜在表示,无需辅助网络,实现高效稳定的合成数据生成。
RMFlow:通过噪声注入步骤细化均流以实现多模态生成
机构 * Department of Mathematics and Scientific Computing and Imaging (SCI) Institute University of Utah(数学与科学计算及成像学院(SCI)院,犹他大学) ; Department of Mathematics, UCLA(数学系,加州大学洛杉矶分校)
专题命中 文生图 :image generation(abstract);text-to-image(abstract)
AI总结 RMFlow通过引入噪声注入步骤,改进均流模型,实现高效多模态生成,仅需单次功能评估即可达到接近最先进的性能。
Comments Accepted to ICLR 2026
人工智能进展:面向创意产业的综述
机构 * Visual Information Laboratory, University of Bristol, Bristol, UK(布里斯托大学视觉信息实验室)
专题命中 文生图 :text-to-image(abstract);diffusion(abstract)
AI总结 本文综述了自2022年以来人工智能在创意产业中的进展,探讨了生成式AI、大语言模型和扩散模型等技术对创意生产流程的影响,并分析了人类与AI协作的新趋势及面临的挑战。
Comments This is an updated review of our previous paper (see https://doi.org/10.1007/s10462-021-10039-7), and has been accepted by Artificial Intelligence Review journal
多语言到多模态(M2M):通过单语文本解锁新语言
机构 * Amazon(亚马逊)
专题命中 文生图 :image generation(abstract);text-to-image(abstract)
AI总结 M2M通过单语文本学习多语言多模态对齐,实现多语言文本到图像检索的零样本迁移
Comments EACL 2026 Findings accepted. Camera-ready version
拓扑视角下的最优多模态嵌入空间
专题命中 文生图 :image generation(abstract);text-to-image(abstract)
AI总结 本文通过拓扑数据分析比较CLIP和CLOOB的嵌入空间,揭示其模态差距驱动因素和维度坍缩的影响,为多模态模型优化提供新视角。
Comments This manuscript contains substantive technical inaccuracies and an incomplete treatment of the stated topic. Subsequent developments and a reassessment of the problem indicate that the scope and framing of the work do not adequately reflect the current state of research, and the analysis is therefore incomplete and outdated
Beam-Brainstorm: 一种生成式特定地点波束成形方法
专题命中 文生图 :diffusion(abstract);image synthesis(abstract)
AI总结 本文提出了一种生成式特定地点波束成形方法,通过联合结构建模和自定义扩散模型,实现高效且高质量的用户特定波束生成。
通过链式推理和任务指令提示降低版权侵权风险
机构 * Munich RE(慕尼黑RE)
专题命中 文生图 :image generation(abstract);text-to-image(abstract)
AI总结 本文通过链式推理和任务指令提示结合负提示和提示重写,降低生成图像的版权侵权风险,并评估不同模型复杂度下的效果。
探索过渡匹配的设计空间
机构 * FAIR at Meta(Meta 的 FAIR)
专题命中 文生图 :text-to-image(abstract);diffusion(abstract)
AI总结 本文研究了过渡匹配中头部模块的设计,发现MLP头部结合特定时间加权和高频采样器在生成质量、训练和推理效率上达到最佳效果。
近似高斯映射用于生成性图像隐写术
专题命中 文生图 :diffusion(abstract);image synthesis(abstract)
AI总结 本文提出近似高斯映射用于生成性图像隐写术,通过调节噪声尺度和方差嵌入秘密信息,并通过两阶段解耦优化策略提升鲁棒性和安全性。
Comments 13 pages
关于将EEG信号与生成式AI连接的综述:从图像和文本到更广泛的领域
机构 * School of Information, The University of Texas at Austin, Austin, TX, USA(信息学院,德克萨斯大学奥斯汀分校,奥斯汀,德克萨斯州)
专题命中 文生图 :diffusion(abstract);image synthesis(abstract)
AI总结 本综述探讨了将EEG信号转化为图像、文本和音频的生成式AI方法,分析了当前技术趋势与挑战,为未来研究提供参考。
机构 * Leibniz Supercomputing Centre(莱比锡超算中心) ; Stanford University(斯坦福大学)
专题命中 文生图 :image generation(abstract);text-to-image(abstract)
Comments 21 pages, 8 figures. Full preprint version; shorter version in preparation
机构 * Norwegian University of Science and Technology, Department of Computer Science(挪威科学技术大学计算机科学系)
专题命中 文生图 :diffusion(abstract);image synthesis(abstract)
机构 * University of Maryland, College Park(马里兰大学彭福分校) ; Queen Mary University of London(伦敦玛丽女王大学) ; Google DeepMind(谷歌DeepMind)
专题命中 文生图 :image generation(abstract);text-to-image(abstract)
专题命中 文生图 :text-to-image(abstract);diffusion(abstract)
Comments Finding unfinished issue in this work , still refining
专题命中 文生图 :image generation(abstract);text-to-image(abstract)
机构 * ETH Zurich(苏黎世联邦理工学院) ; Princeton University(普林斯顿大学)
专题命中 文生图 :image generation(abstract);text-to-image(abstract)
Comments NeurIPS 2025 (Spotlight); Code available at https://github.com/smonsays/scale-compositionality
机构 * Apple(苹果公司) ; Sapienza(拉维尼恩大学)
专题命中 文生图 :image generation(abstract);text-to-image(abstract)
Comments NeurIPS 2025
机构 * Carnegie Mellon University(卡内基梅隆大学) ; University of Rochester(罗切斯特大学)
专题命中 文生图 :image generation(abstract);text-to-image(abstract)
机构 * School of Mathematics University of Edinburgh(数学学院爱丁堡大学) ; School of Mathematical and Computer Science Heriot-Watt University(数学与计算机科学学院赫瑞斯大学)
专题命中 文生图 :text-to-image(abstract);diffusion(abstract)
机构 * Sea AI Lab(海智实验室) ; SUTD(新加坡科技设计大学) ; NUS(国立大学) ; NTU(南洋理工大学) ; University of Waterloo(滑铁卢大学)
专题命中 文生图 :image generation(abstract);text-to-image(abstract)
机构 * Fudan University(复旦大学)
专题命中 文生图 :text-to-image(abstract);image synthesis(abstract)
Comments Accepted by Findings of EMNLP2025
专题命中 文生图 :text-to-image(abstract);diffusion(abstract)
专题命中 文生图 :text-to-image(abstract);diffusion(abstract)
Comments 13 pages
机构 * School of Computing(计算学院) ; National University of Singapore(新加坡国立大学)
专题命中 文生图 :text-to-image(abstract);diffusion(abstract)
Comments 23 pages
专题命中 文生图 :image generation(abstract);text-to-image(abstract)
机构 * Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院,清华大学) ; ARC Lab, Tencent PCG(腾讯PCG ARC实验室) ; The Chinese University of Hong Kong(香港中文大学) ; The University of Hong Kong(香港大学)
专题命中 文生图 :text-to-image(abstract);diffusion(abstract)
Comments Code: https://github.com/TencentARC/MindOmni
专题命中 文生图 :text-to-image(abstract);diffusion(abstract)
Comments 6 pages, 4 figures, IEEE ICIP 2025
专题命中 文生图 :text-to-image(abstract);diffusion(abstract)
机构 * School of Mathematical Sciences, Peking University(北京大学数学科学学院) ; School of Electronic, Information and Electrical Engineering, Shanghai Jiao Tong University(上海交通大学电子信息与电气工程学院)
专题命中 文生图 :image generation(abstract);text-to-image(abstract)
机构 * University of California, Los Angeles(加州大学洛杉矶分校)
专题命中 文生图 :image generation(abstract);text-to-image(abstract)