arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SenseNova-U1.5:迈向原生统一视觉智能

SenseNova-U1.5: Towards Native Unified Visual Intelligence

Haiwen Diao, Jiahao Wang, Chenjing Ding, Hanming Deng, Jiangnan Chen, Ruixi Zhang, Ruohui Wang, Wenwen Tong, Xiangyu Fan, Yubo Wang, Yue Zhu, Yuwei Niu, Zhengqi Bai, Zhiqian Lin, Zhitao Yang, Zhongang Cai, Bo Yang, Chen Feng, Chengguang Lv, Guangjia Liu, Guanlin Wang, Hanyu Zhang, Haojia Yu, Hongcan Xiao, Hongli Wang, Huan Wu, Huaping Zhong, Jian Fang, Jianan Fan, Jiaqi Li, Jiefan Lu, Jing Zuo, Jingcheng Ni, Junxiang Xu, Linjun Dai, Mutian Xu, Peishen Yan, Penghao Wu, Ruijie Mao, Ruisi Wang, Shihao Bai, Shuang Yang, Shuya Yang, Shuyan Zheng, Silei Wu, Siying Li, Tao Chu, Tianbo Zhong, Tongxi Zhou, Weichao Luo, Weichen Fan, Wenhao Jia, Wenjie Gao, Xiangli Kong, Yan Li, Yang Yong, Zimo Wen, Zixuan Qian, Wenxiu Sun, Ruihao Gong, Quan Wang, Lewei Lu, Lei Yang, Ziwei Liu, Dahua Lin

arXiv 2609.11929首次发表:更新:

AI 中文总结

SenseNova-U1.5作为8B-MoT原生统一多模态模型,通过无编码器架构和专家蒸馏,在图像生成与编辑上取得显著提升,并验证了多模态理解向视觉规划迁移的可行性。

AI 中文摘要

我们发布了SenseNova-U1.5,一个8B-MoT原生统一多模态模型,能够在无编码器、无VAE的架构中理解、推理并生成视觉内容。我们通过空间连贯的补丁重建强化其视觉接口,并利用精心策划的生成与编辑数据、改进的任务表述、结构化提示增强以及高达4K的原生分辨率来扩展其训练。在后训练阶段,我们针对视觉美学、双语文本渲染、信息图生成和图像编辑优化了专门的专家模型,并通过多专家在线策略蒸馏整合它们的能力。在广泛的评估中,SenseNova-U1.5显著提升了图像保真度、文本渲染、复杂构图、多参考编辑和交错生成,同时改进了指令跟随并保持了主体身份、几何形状和未修改区域。尽管其生成数据中结构化格式的暴露有限,SenseNova-U1.5仍能有效泛化到长、复杂和结构化的视觉指令,进一步证明多模态理解可以迁移到视觉规划和创作。这些发现共同表明,原生统一建模是迈向在完全端到端框架内感知、推理和创造的系统的有前景路径。我们将开源训练代码,包括监督微调、强化学习和在线策略蒸馏。

英文摘要

We launch SenseNova-U1.5, an 8B-MoT native unified multimodal model that understands, reasons about, and generates visual content within an encoder-free and VAE-free architecture. We strengthen its visual interface through spatially coherent patch reconstruction and scale its training with carefully curated generation and editing data, improved task formulation, structural prompt enhancement, and native resolutions of up to 4K. For post-training, we optimize specialized experts for visual aesthetics, bilingual text rendering, infographic generation, and image editing, and consolidate their capabilities through multi-expert on-policy distillation. Across extensive evaluations, SenseNova-U1.5 largely advances image fidelity, text rendering, complex composition, multi-reference editing, and interleaved generation, while improving instruction following and preserving subject identity, geometry, and unmodified regions. Despite limited exposure to structured formats in its generation data, SenseNova-U1.5 generalizes effectively to long, complex, and structured visual instructions, further proving that multimodal understanding can transfer to visual planning and creation. Together, these findings position native unified modelling as a promising path towards systems that perceive, reason and create within a fully end-to-end framework. We will open-source training code, including supervised fine-tuning, reinforcement learning, and on-policy distillation.

CommentsProject page: https://github.com/OpenSenseNova/SenseNova-U1

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑