AI 中文总结
研究针对抗体表达排序中标记数据稀缺问题,提出基于偏好的学习框架,结合定量表达与弱监督,通过改进DPO适用于蛋白质语言模型,在多样数据集上评估,该方法优于基线,为抗体可表达性优化提供可扩展方案。
AI 中文摘要
抗体表达排序在抗体设计中至关重要,但标记实验数据稀缺严重阻碍其建模。为此,我们提出统一的基于偏好的学习框架,将稀缺的定量表达数据与免疫数据中的大规模弱阳性监督相结合。通过引入联合掩码对数似然近似和基于IMGT的比对,使直接偏好优化(DPO)适用于蛋白质语言模型,能对可变长度序列进行高效训练。在1254个标记序列和400万个未标记骆驼科来源抗体的多样内部数据集上评估,我们的方法在多数指标上持续优于基线。结果表明偏好学习能有效从弱监督中学习,为数据受限环境下的抗体可表达性优化提供了可扩展解决方案。项目页面:此https URL。
英文摘要
Antibody expression ranking is a critical task in antibody design, yet its modelling is severely hindered by the scarcity of labeled experimental data. To address this, we propose a unified preference-based learning framework that integrates scarce quantitative expression data with large-scale weak positive supervision from immunization data. We adapt Direct Preference Optimization (DPO) to protein language models by introducing a union-masked log-likelihood approximation and IMGT-based alignment, enabling efficient training on variable-length sequences. Evaluating on a diverse internal dataset of 1254 labeled sequences and 4 million unlabeled camelid-derived antibodies, we show that our method consistently outperforms baselines on most metrics. Our results demonstrate that preference learning can effectively learn from weak supervision, providing a scalable solution for antibody expressibility optimization in data-constrained settings. Project page: https://kisoji-biotechnology-inc.github.io/Preference-Expression-Ranking/.
CommentsAccepted at ICML 2026