arXivDaily arXiv每日学术速递 周一至周五更新

大厂专区

MiniMax

2026-03-09 至 2026-03-09 共收录 1
2512.13687 2026-03-09 cs.CV

Towards Scalable Pre-training of Visual Tokenizers for Generation

面向生成任务的视觉分词器可扩展预训练

Jingfeng Yao, Yuda Song, Yucong Zhou, Xinggang Wang

机构 * Huazhong University of Science and Technology(华中科技大学) MiniMax

AI总结 VTP通过联合优化图像-文本对比、自监督和重建损失,提升视觉分词器的生成性能和扩展性。

Comments Our pre-trained models are available at https://github.com/MiniMax-AI/VTP

详情

展开后加载摘要…

URL PDF HTML 收藏