AI 中文总结
提出ProFiT,一种轻量级流匹配蛋白质结构分词器,通过简单训练策略高效学习语义表示,无需手动对齐,重建质量与泛化能力媲美大型分词器,即插即用,适应多种下游任务。
AI 中文摘要
作为蛋白质模态与离散建模之间的桥梁,蛋白质结构分词仍主要依赖于针对特定下游任务定制且需大量训练数据的复杂训练目标,这阻碍了其向更广泛的应用场景迁移。为解决此问题,我们提出ProFiT,一种轻量级流匹配分词器。通过鼓励健康码本利用的简单训练策略,ProFiT能够高效训练,无需任何手动语义对齐即可自然学习到语义上有意义的表示,同时实现与规模大得多的分词器相当或更优的重建质量和泛化能力。我们在多种设置下进行了广泛评估,证明ProFiT是一种即插即用的分词器,可适应多样化的下游任务。本研究进一步揭示了流匹配分词器范式的巨大潜力。我们的代码在此https URL公开提供。
英文摘要
As the bridge between protein modality and discrete modeling, protein structure tokenization still largely relies on heavily engineered training objectives tailored to specific downstream tasks and large training datasets, which hinders its transfer to broader application scenarios. To address this issue, we propose ProFiT, a lightweight flow matching tokenizer. With simple training strategies that encourage healthy codebook utilization, ProFiT can be trained efficiently and naturally learns semantically meaningful representations without any manual semantic alignment, while achieving reconstruction quality and generalization that match or surpass those of substantially larger tokenizers. We conduct extensive evaluations across a wide range of settings and demonstrate that ProFiT is a plug-and-play tokenizer adaptable to diverse downstream tasks. This study further reveals the significant potential of the flow matching tokenizer paradigm. Our code is publicly available at https://github.com/QDKStorm/ProFiT.