arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过功能感知掩码学习任务特异性抗体表征

Learning Task-Specific Antibody Representations via Function-Aware Masking

Ayan Goel, Thomas A. Walton, Amirali Aghazadeh

arXiv 2609.00518首次发表:更新:

发表机构

School of Electrical and Computer Engineering, Georgia Institute of Technology; School of Computer Science, Georgia Institute of Technology(佐治亚理工学院电气与计算机工程学院; 佐治亚理工学院计算机科学学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出功能感知掩码算法,通过对齐特定功能先验的掩码位置塑造抗体表征空间,提升了结构、CDR等相关任务性能,还开发了混合掩码策略平衡多维度目标。

AI 中文摘要

通过掩码语言建模(MLM)预训练的抗体专用语言模型可学习对下游序列设计和属性预测任务至关重要的表征,但预训练过程中很少利用掩码本身作为归纳偏置来源。虽然优先对互补决定区(CDRs)进行掩码可改善结合相关预测,但抗体在多种功能上具有多样的生物先验。本文提出功能感知掩码,这是一类预训练算法,其掩码位置与特定功能先验(如来自IMGT注释或结构预测的先验)对齐,以塑造学习到的表征空间。研究表明,这些专用掩码策略显著提升了对应任务的性能,在结构相关任务上最高提升14%,在CDR相关任务上最高提升5.9倍。为进一步提升跨多个功能维度的性能,开发了整合多种先验的混合掩码策略,平衡结合、结构和生物物理目标的重构。结果表明,合理的掩码放置是一种无参数机制,可在抗体语言模型训练中施加功能归纳偏置。

英文摘要

Antibody-specific language models pretrained via masked language modeling (MLM) learn representations that are critical for downstream sequence design and property prediction tasks. Yet, the corruption process itself is rarely leveraged as a source of inductive bias during pretraining. While preferentially masking complementarity-determining regions (CDRs) improves binding-related predictions, antibodies possess diverse biological priors over a variety of functions. Herein, we introduce function-aware masking, a family of pretraining algorithms that align mask placement with specific functional priors (e.g., from IMGT annotations or structure predictions) to shape the learned representation space. We show that these specialist masking strategies significantly improve performance on their respective objectives, yielding up to a 14% gain on structure-related tasks and up to a 5.9x improvement on CDR-related tasks. To further improve performance across multiple functional axes, we develop hybrid masking strategies that integrate multiple priors, balancing reconstruction over binding, structural, and biophysical objectives. Our results demonstrate that informed mask placement provides a parameter-free mechanism for imposing functional inductive biases in antibody language model training.

CommentsAccepted to MLCB 2026 (Oral)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑