arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

期刊&会议

NeurIPS

Conference on Neural Information Processing Systems · 会议 · Machine Learning

至 收录 17318
2508.10900 2026-06-23 cs.CV 版本更新

Quantum Visual Fields with Neural Amplitude Encoding

量子视觉场与神经振幅编码

Shuteng Wang, Christian Theobalt, Vladislav Golyanik

机构 * MPI for Informatics, SIC(马克斯·普朗克信息研究所,科学信息中心)

AI总结 提出一种基于神经振幅编码和全纠缠量子电路的量子隐式神经表示架构QVF,用于2D图像和3D几何场学习,在量子硬件模拟器上优于现有量子方法并与经典基线竞争。

Comments NeurIPS 2025; 19 pages, 13 figures and four tables; project page: https://4dqv.mpi-inf.mpg.de/QVF/

URL PDF HTML 收藏
2502.15376 2026-06-23 cs.LG cond-mat.mes-hall 版本更新

Learning Chern Numbers of Topological Insulators with Gauge Equivariant Neural Networks

利用规范等变神经网络学习拓扑绝缘体的陈数

Longde Huang, Oleksandr Balabanov, Hampus Linander, Mats Granath, Daniel Persson, Jan E. Gerken

机构 * Department of Mathematical Sciences, Chalmers University of Technology and University of Gothenburg(数学科学系,查尔姆斯理工大学和哥德堡大学) Department of Physics, Stockholm University, AlbaNova University Center(物理系,斯德哥尔摩大学,阿尔巴诺瓦大学中心) VERSES AI Research Lab, Los Angeles, USA(VERSES AI研究实验室,美国洛杉矶) Department of Physics, University of Gothenburg(物理系,哥德堡大学)

AI总结 本文提出利用规范等变网络预测多带拓扑绝缘体的陈数,通过引入新的规范等变归一化层和通用逼近定理,证明模型能泛化至非平凡陈数样本。

Journal ref Advances in Neural Information Processing Systems 38, 147997-148026, 2026

URL PDF HTML 收藏
2606.19882 2026-06-19 cs.CV cs.LG 新提交

Multimodal Concept Bottleneck Models

多模态概念瓶颈模型

Tongqing Shi, Ge Yan, Tuomas Oikarinen, Tsui-Wei Weng

机构 * UC San Diego(加州大学圣地亚哥分校)

AI总结 提出多模态概念瓶颈模型(MM-CBM),利用双概念瓶颈层对齐图像和文本嵌入,实现可解释的零样本分类和图像检索,在四个基准上平均准确率提升高达51.26%。

Comments Present at NeurIPS 2025 Mechanistic Interpretability Workshop

URL PDF HTML 收藏
2510.19893 2026-06-19 cs.LG 版本更新

EQPO: Equitable Group Relative Policy Optimization for Clinical Reasoning

EQPO: 面向临床推理的公平群体相对策略优化

Shiqi Dai, Wei Dai, Jiaee Cheong, Paul Pu Liang

机构 * MIT(麻省理工学院) Harvard University(哈佛大学)

AI总结 提出EQPO分层强化学习方法,通过自适应重加权样本促进异质临床人群的均衡学习,在7个诊断基准上降低F1标准差43.9%,缩小预测公平差距27.2%。

Comments Accepted as Oral on NeurIPS 2025 GenAI4Health Workshop

URL PDF HTML 收藏
2606.18338 2026-06-18 cs.LG astro-ph.EP astro-ph.IM 新提交

ThousandWorlds: A benchmark for climate emulation of potentially habitable exoplanets

ThousandWorlds: 一个用于潜在宜居系外行星气候模拟的基准数据集

Edward T. Stevenson, Mei Ting Mak, Eric Wolf, Denis E. Sergeev, Tobi Hammond, N. J. Mayne, Miles Cranmer

机构 * University of Cambridge(剑桥大学) University of Oxford(牛津大学) University of Colorado Boulder(科罗拉多大学博尔德分校) University of Bristol(布里斯托大学) Purdue University(普渡大学) University of Exeter(埃克塞特大学)

AI总结 为加速系外行星气候模拟,提出ThousandWorlds基准数据集,包含五个全球气候模型的约1800次模拟,用于评估机器学习模拟器在低数据、多模拟器参数到场回归任务中的性能。

Comments 10 pages main text, 26 pages references/appendix, plus NeurIPS checklist. Data at https://doi.org/10.57967/hf/8695. Code at https://github.com/edstevenson/ThousandWorlds

URL PDF HTML 收藏
2606.17639 2026-06-18 cs.RO cs.CV 新提交

ERQA-Plus: A Diagnostic Benchmark for Reasoning in Embodied AI

ERQA-Plus:具身AI推理的诊断基准

Hong Yang, Basura Fernando

机构 * Centre for Frontier AI Research, Agency for Science, Technology and Research(新加坡科技研究局前沿人工智能研究中心) College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算与数据科学学院)

AI总结 提出ERQA-Plus基准,包含1766个基于机器人中心图像的问答实例,覆盖感知、动作、社交、导航和常识推理,用于诊断具身AI的推理能力。

Comments under review at NeurIPS

URL PDF HTML 收藏
2503.08038 2026-06-18 cs.LG cs.AI cs.CV 版本更新

Generalized Kullback-Leibler Divergence Loss

广义Kullback-Leibler散度损失

Jiequan Cui, Beier Zhu, Qingshan Xu, Zhuotao Tian, Xiaojuan Qi, Bei Yu, Hanwang Zhang, Richang Hong

机构 * Hefei University of Technology(合肥工业大学) University of Science and Technology of China(中国科学技术大学) Nanyang Technological University(南洋理工大学) The Chinese University of Hong Kong(香港中文大学) The University of Hong Kong(香港大学) Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳))

AI总结 本文提出广义KL散度损失,通过解耦KL损失为加权MSE和交叉熵损失,并引入非对称优化修正和类别全局信息,在对抗训练和知识蒸馏中取得SOTA性能。

Comments TPAMI 2026, extension of our NeurIPS paper "Decoupled Kullback-Leibler Divergence Loss". arXiv admin note: substantial text overlap with arXiv:2305.13948

URL PDF HTML 收藏
2606.17526 2026-06-17 cs.LG 新提交

MGUP: A Momentum-Gradient Alignment Update Policy for Stochastic Optimization

MGUP:一种用于随机优化的动量-梯度对齐更新策略

Da Chang, Ganzhao Yuan

机构 * Shenzhen Institute of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究院) Shenzhen University of Advanced Technology(深圳理工大学) Pengcheng Laboratory(鹏城实验室) University of Chinese Academy of Sciences(中国科学院大学)

AI总结 提出MGUP机制,通过按固定比例选择参数施加大步长、其余参数用小步长,增强动量优化器,理论保证收敛,实验表明提升训练效率与稳定性。

Comments Published in NeurIPS 2025

URL PDF HTML 收藏
2511.19162 2026-06-17 cs.IR cs.CY cs.HC cs.LG cs.MM 版本更新

BioArtlas: Computational Clustering of Multi-Dimensional Complexity in Bioart

BioArtlas:生物艺术中多维复杂性的计算聚类

Joonhyung Bae

机构 * Graduate School of Culture Technology(文化科技研究生院)

AI总结 本文提出BioArtlas,通过新型轴感知表示对81件生物艺术作品进行多维分析,揭示四种组织模式,并通过交互式网页界面提供分析与探索。

Comments Bae, J. BioArtlas: Computational Clustering of Multi-Dimensional Complexity in Bioart. In The Thirty-ninth Annual Conference on Neural Information Processing Systems Creative AI Track: Humanity

URL PDF HTML 收藏
2606.15899 2026-06-16 cs.CR cs.AI cs.HC cs.LG cs.MA 新提交

SkillVetBench: LLM-as-Judge for Multi-Dimensional Security Risk Evaluation in Open-Source LLM Agent Skills

SkillVetBench: 基于LLM评判的多维安全风险评估开源LLM智能体技能

Ismail Hossain, Sai Puppala, Md Jahangir Alam, Tanzim Ahad, Sajedul Talukder

机构 * SUPREME Lab, University of Texas at El Paso, Texas, USA(SUPREME实验室,德克萨斯理工大学埃尔帕索分校,德克萨斯州,美国)

AI总结 提出SkillVetBench,利用LLM作为评判器对开源LLM智能体技能进行多维安全风险评估,引入五维技能智能体风险评分(SARS)和CVSS v4.0向量分解,在78个恶意技能上实现零假阴性,22个良性技能上零假阳性。

Comments The main research paper is submitted to NeurIPS 2027, it is in under review

URL PDF HTML 收藏
2510.24987 2026-06-16 q-bio.QM cs.LG q-bio.GN

scMRDR: A scalable and flexible framework for unpaired single-cell multi-omics data integration

scMRDR:一种可扩展且灵活的无配对单细胞多组学数据整合框架

Jianle Sun, Chaoqi Liang, Ran Wei, Peng Zheng, Lei Bai, Wanli Ouyang, Hongliang Yan, Peng Ye

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Carnegie Mellon University(卡内基梅隆大学) The Chinese University of Hong Kong(香港中文大学) Guangzhou Laboratory(广州实验室)

AI总结 scMRDR通过β-VAE架构解耦细胞潜在表示,结合等距正则化、对抗目标和掩码重建损失,实现无配对多组学数据整合,有效提升大规模数据处理能力。

Comments Accepted at NeurIPS 2025 (Spotlight)

Journal ref Advances in Neural Information Processing Systems 38 (2025): 154538-154565

URL PDF HTML 收藏
2406.07277 2026-06-16 cs.CL cs.AI cs.MA

Speaking Your Language: Spatial Relationships in Interpretable Emergent Communication

说出你的语言:可解释的涌现交流中的空间关系

Olaf Lipinski, Adam J. Sobey, Federico Cerutti, Timothy J. Norman

机构 * University of Southampton(索姆塞特大学) The Alan Turing Institute(艾伦·图灵研究所) University of Brescia(布雷西亚大学)

AI总结 本文研究了智能体如何通过空间关系交流,展示了其能发展出表达观察部分关系的语言,实现90%以上的准确率,并证明该语言可被人类解读。

Comments Accepted at NeurIPS 2024. 18 pages, 3 figures

Journal ref In Advances in Neural Information Processing Systems (Vol. 37, pp. 140113-140137) 2024

URL PDF HTML 收藏
2606.14673 2026-06-15 cs.LG 新提交

Compressed Computation is (probably) not Computation in Superposition

压缩计算(可能)不是叠加计算

Jai Bhagat, Sara Molas-Medina, Giorgi Giglemiani, Stefan Heimersheim

机构 * Metamorphic Independent(独立研究者) UK AI Security Institute(英国人工智能安全研究所) Apollo Research

AI总结 通过分析压缩计算(CC)模型,发现其性能提升源于标签中的混合矩阵,而非真正的叠加计算,SNMF基线可复现其损失特征。

Comments Presented at the Mechanistic Interpretability Workshop at NeurIPS 2025

URL PDF HTML 收藏
2606.14108 2026-06-15 cs.LG cs.AI 新提交

Numbers Already Carry Their Own Embeddings

数字本身已携带其嵌入

Suhyun Bae, Donghun Lee

机构 * Department of Mathematics, Korea University(高丽大学数学系)

AI总结 提出无训练嵌入方法AOE,同时保留数字的实数值与p-adic模签名,实现即插即用并在代数组合基准上首次达到完美精度。

Comments Presented at the MATH-AI Workshop at NeurIPS 2025

URL PDF HTML 收藏
2412.03716 2026-06-15 cs.LG cs.CY 版本更新

A Water Efficiency Dataset for African Data Centers

非洲数据中心用水效率数据集

Noah Shumba, Opelo Tshekiso, Pengfei Li, Giulia Fanti, Shaolei Ren

机构 * Carnegie Mellon University(卡内基梅隆大学) Carnegie Mellon University Africa Kigali Rwanda(卡内基梅隆大学非洲分校,基亚利,卢旺达) Rochester Institute of Technology(罗切斯特理工学院) Rochester New York USA(罗切斯特,纽约州,美国) Carnegie Mellon University Pittsburgh Pennsylvania USA(卡内基梅隆大学匹兹堡,宾夕法尼亚州,美国) University of California, Riverside(加州大学河滨分校)

AI总结 构建首个结合天气与发电数据的非洲41国数据中心用水效率数据集,评估Llama-3-70B和GPT-4推理用水量,发现多数非洲国家用水低于全球平均。

Comments Accepted by NeurIPS 2024 Workshop on Tackling Climate Change with Machine Learning

URL PDF HTML 收藏
2606.12747 2026-06-12 cs.AI 新提交

Prefill Awareness in Large Language Models

大型语言模型中的预填充感知

Andy Wang, Parv Mahajan, David Demitri Africa, Alexandra Souly, Jordan Taylor, Robert Kirk

机构 * Constellation University of Wisconsin-Madison(威斯康星大学麦迪逊分校星座研究所) Constellation Georgia Institute of Technology(佐治亚理工学院星座研究所) UK AI Security Institute(英国人工智能安全研究所)

AI总结 研究大型语言模型能否识别并响应其助手消息被预填充或篡改,发现前沿模型具有显著预填充感知能力,可能影响安全评估方法。

Comments Submitted to NeurIPS 2026

URL PDF HTML 收藏
2509.03340 2026-06-12 cs.LG cs.AI cs.CE physics.comp-ph 版本更新

Equivariant Flow Matching for Symmetry-Breaking Bifurcation Problems

等变流匹配用于对称破缺分岔问题

Fleur Hendriks, Ondřej Rokoš, Martin Doškář, Marc G. D. Geers, Vlado Menkovski

机构 * Department of Mechanical Engineering, Eindhoven University of Technology(埃因霍温理工大学机械工程系) DIFFER – Dutch Institute for Fundamental Energy Research(荷兰基础能源研究所) Faculty of Civil Engineering, Department of Mechanics, Czech Technical University in Prague(布拉格捷克技术大学土木工程学院力学系) Department of Mathematics and Computer Science, Eindhoven University of Technology(埃因霍温理工大学数学与计算机科学系)

AI总结 针对非线性动力系统中对称破缺导致的多稳态共存问题,提出等变流匹配方法,结合等变架构与最优传输耦合机制,准确捕捉多模态分布和对称破缺分岔,优于非概率和变分方法。

Comments 9 pages, 7 figures including appendices. Accepted to Machine Learning and the Physical Sciences Workshop, NeurIPS 2025 (https://ml4physicalsciences.github.io/2025/). Repository with corresponding code: https://github.com/FHendriks11/bifurcationML/. Video explanation: https://www.youtube.com/watch?v=wsL3h17KtjY

URL PDF HTML 收藏
2606.11199 2026-06-11 cs.CL cs.AI cs.IR cs.LG 新提交

NightFeats @ MMU-RAGent NeurIPS 2025: A Context-Optimized Multi-Agent RAG System for the Text-to-Text Track

NightFeats @ MMU-RAGent NeurIPS 2025: 面向文本到文本轨道的上下文优化多智能体RAG系统

Quentin Fever, Naziha Aslam

机构 * NightFeats

AI总结 提出一种结构化多智能体RAG系统NightFeats,通过检索、策展和组合三阶段分解知识合成,引入时序语义重排序、矛盾协调和引用保留架构,在MMU-RAGent竞赛中超越商业基线。

Comments 5 pages, 1 figure, 1 table. NeurIPS 2025 Competition Track (MMU-RAGent). System developed October 2025

URL PDF HTML 收藏
2601.22725 2026-06-11 cs.CV cs.AI 版本更新

OpenVTON-Bench: A Large-Scale High-Resolution Benchmark for Controllable Virtual Try-On Evaluation

OpenVTON-Bench:用于可控虚拟试穿评估的大规模高分辨率基准

Jin Li, Tao Chen, Kai Wen, Siqi Yin, Shuai Jiang, Weijie Wang, Jingwen Luo, Chenhui Wu

机构 * Renxing Intelligence, Hangzhou, China Hangzhou Dianzi University, Hangzhou, China(杭州电子科技大学)

AI总结 提出OpenVTON-Bench,包含约10万对高分辨率图像,通过DINOv3聚类和Gemini描述构建,并设计多模态评估协议,沿五个维度衡量试穿质量,与人类判断高度一致。

Comments Under review for the NeurIPS 2026 Datasets and Benchmarks Track

URL PDF HTML 收藏
2507.11688 2026-06-11 cs.LG 版本更新

Composing Linear Layers from Irreducibles

从不可约元组合线性层

Travis Pence, Daisuke Yamada, Vikas Singh

机构 * University of Wisconsin-Madison(威斯康星大学麦迪逊分校)

AI总结 提出用Clifford代数将线性层分解为双向量(几何基元)的组合,仅需O(log^2 d)参数,在LLM注意力投影中匹配强基线性能。

Comments 35 Pages, 11 Tables, 6 Figures, Appearing in NeurIPS 2025

Journal ref Advances in Neural Information Processing Systems 38 (2025)

URL PDF HTML 收藏
2512.03077 2026-06-11 cs.CY cs.AI 版本更新

Irresponsible AI: big tech's influence on AI research and associated impacts

不负责任的人工智能:大型科技公司对AI研究的影响及相关影响

Alex Hernandez-Garcia, Alexandra Volokhova, Ezekiel Williams, Dounia Shaaban Kabakibo, Mélisande Teng

机构 * Big Tech(大科技公司)

AI总结 本文指出大型科技公司对AI研究的不成比例影响推动了不负责任的AI发展,并加剧了环境和社会负面影响,呼吁研究者通过集体行动加以抵制。

Comments Presented as a spotlight oral at the International Conference on Machine Learning 2026 (Position Paper Track). First version presented at NeurIPS 2025 Workshop on Algorithmic Collective Action

URL PDF HTML 收藏
2509.16456 2026-06-11 cs.AI 版本更新

GPO: Learning from Critical Steps to Improve LLM Reasoning

GPO:从关键步骤中学习以改进大语言模型推理

Jiahao Yu, Zelei Cheng, Xian Wu, Xinyu Xing

机构 * Department of Computer Science Northwestern University(计算机科学系西北大学) AI Foundations Capital One(人工智能基础资本 one) Meta AI

AI总结 提出引导式关键优化(GPO)微调策略,通过识别推理轨迹中的关键步骤并优先学习,显著提升大语言模型的多步推理能力。

Comments 39th Conference on Neural Information Processing Systems (NeurIPS 2025)

URL PDF HTML 收藏
2510.08073 2026-06-11 cs.CV cs.LG 版本更新

Physics-Driven Spatiotemporal Modeling for AI-Generated Video Detection

物理驱动的时空建模用于AI生成视频检测

Shuhai Zhang, ZiHao Lian, Jiahao Yang, Daiyuan Li, Guoxuan Pang, Feng Liu, Bo Han, Shutao Li, Mingkui Tan

机构 * South China University of Technology(华南理工大学) University of Science and Technology of China(中国科学技术大学) Key Laboratory of Big Data and Intelligent Robot, Ministry of Education(教育部大数据与智能机器人重点实验室) Pazhou Lab(琶洲实验室) University of Melbourne(墨尔本大学) Hunan University(湖南大学) Hong Kong Baptist University(香港 Baptist大学)

AI总结 提出基于概率流守恒的物理驱动AI生成视频检测范式,通过归一化时空梯度(NSG)统计量捕捉物理异常,结合预训练扩散模型估计NSG,并利用最大均值差异(MMD)进行检测,在Recall和F1-Score上分别提升16.00%和10.75%。

Comments Accepted at NeurIPS 2025 spotlight

URL PDF HTML 收藏
2510.02660 2026-06-11 cs.HC cs.AI 版本更新

When Researchers Say Mental Model/Theory of Mind of AI, What Are They Really Talking About?

当研究人员谈论AI的心理模型/心智理论时,他们究竟在说什么?

Xiaoyun Yin, Elmira Zahmat Doost, Shiwen Zhou, Garima Arya Yadav, Jamie C. Gorman

机构 * Center for Human, Artificial Intelligence, and Robot Teaming(人类、人工智能与机器人协同中心)

AI总结 本文指出当前AI心智理论研究混淆了行为预测与真实认知,提出应转向人机交互中的互惠心智理论框架。

Comments This work have been accepted in CogInterp @ NeurIPS 2025

URL PDF HTML 收藏
2606.11130 2026-06-10 cs.LG 新提交

Robust Regression of General ReLUs with Queries

一般ReLU的鲁棒回归与查询

Ilias Diakonikolas, Daniel M. Kane, Mingchen Ma

机构 * University of Wisconsin-Madison(威斯康星大学麦迪逊分校) University of California, San Diego(加利福尼亚大学圣迭戈分校)

AI总结 针对高斯分布下一般ReLU的平方损失鲁棒回归,提出首个高效查询算法,使用d polylog(1/ε)+Õ(min{1/p,1/ε})个标签查询达到O(opt)+ε误差,并证明查询复杂度近最优。

Comments Appeared at NeurIPS 2025

URL PDF HTML 收藏
2606.09940 2026-06-10 cs.LG cs.AI 新提交

Interactions Between Crosscoder Features: A Compact Proofs Perspective

交叉编码器特征间的交互:一个紧凑证明的视角

Dmitry Manning-Coe, Thomas Read, Anna Soligo, Oliver Clive-Griffin, Chun-Hei Yip, Rajashree Agrawal, Jason Gross

机构 * Anthony J. Leggett Institute for Condensed Matter Theory(安东尼·J·莱格特凝聚态理论研究所) MATS UK AI Security Institute (AISI)(英国人工智能安全研究所) Imperial College London(帝国理工学院伦敦分校) Goodfire University of Cambridge(剑桥大学) Theorem Labs(定理实验室)

AI总结 本文从紧凑证明角度形式化交叉编码器特征交互,提出交互度量并应用于计算稀疏性、语义聚类和检测休眠代理。

Comments Accepted at the NeurIPS 2025 Workshop on Mechanistic Interpretability

URL PDF HTML 收藏
2510.04514 2026-06-10 cs.AI cs.CE cs.CL cs.CV stat.ME 版本更新

ChartAgent: A Multimodal Agent for Visually Grounded Reasoning in Complex Chart Question Answering

ChartAgent: 一种用于复杂图表问答中视觉基础推理的多模态智能体

Rachneet Kaur, Nishan Srishankar, Zhen Zeng, Sumitra Ganesh, Manuela Veloso

机构 * J.P. Morgan AI Research(摩根大通人工智能研究)

AI总结 提出ChartAgent框架,通过迭代分解查询为视觉子任务并利用图表专用视觉工具(如绘制注释、裁剪区域)进行空间域推理,在ChartBench和ChartX上取得最先进性能,尤其对无标注图表提升显著。

Comments Accepted at ACL 2026 (Main Conference). Also presented as an oral paper at the NeurIPS 2025 Multimodal Algorithmic Reasoning Workshop (https://marworkshop.github.io/neurips25/)

URL PDF HTML 收藏
2511.02603 2026-06-10 cs.CL 版本更新

CGES: Confidence-Guided Early Stopping for Efficient and Accurate Self-Consistency

CGES:面向高效准确自一致性的置信引导早停方法

Ehsan Aghazadeh, Ahmad Ghasemi, Hedyeh Beyhaghi, Hossein Pishro-Nik

机构 * University of Massachusetts Amherst(马萨诸塞大学阿姆赫斯特分校)

AI总结 提出贝叶斯框架CGES,通过自适应停止采样减少自一致性推理调用次数,在5个推理基准上平均减少58%调用且精度损失仅0.4个百分点。

Comments Extended version. A preliminary version was accepted at the Efficient Reasoning Workshop @ NeurIPS 2025. Code: https://github.com/EhsanAghazadeh/cges

URL PDF HTML 收藏
2310.05264 2026-06-10 cs.LG cs.CV 版本更新

The Emergence of Reproducibility and Generalizability in Diffusion Models

扩散模型中可重复性与泛化性的出现

Huijie Zhang, Jinfan Zhou, Yifu Lu, Minzhe Guo, Peng Wang, Liyue Shen, Qing Qu

机构 * CIFAR-10 dataset(CIFAR-10数据集)

AI总结 研究发现扩散模型在相同初始噪声和确定性采样器下,不同模型输出高度相似,且这种可重复性在记忆和泛化两种训练模式下均存在,对训练效率、模型隐私等有重要启示。

Comments NeurIPS Diffusion Model Workshop 2023 (best paper award), the Forty-first International Conference on Machine Learning (ICML 2024)

URL PDF HTML 收藏
2606.09156 2026-06-09 cs.CV 新提交

OmniGen-AR: AutoRegressive Any-to-Image Generation

OmniGen-AR: 自回归任意到图像生成

Junke Wang, Xun Wang, Qiushan Guo, Peize Sun, Weilin Huang, Zuxuan Wu, Yu-Gang Jiang

机构 * Institute of Trustworthy Embodied AI, Fudan University(复旦大学可信具身人工智能研究所) Shanghai Collaborative Innovation Center of Intelligent Visual Computing(上海智能视觉计算协同创新中心) Bytedance Seed(字节跳动Seed) The University of Hong Kong(香港大学)

AI总结 提出统一自回归框架OmniGen-AR,通过共享视觉分词器和解耦因果注意力,支持文本、空间信号和视觉上下文等多种条件输入,在多项基准上达到最优或竞争性能。

Comments Accepted by NeurIPS

URL PDF HTML 收藏