arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

进化课程学习改进生物序列建模

Evolutionary Curriculum Learning Improves Biological Sequence Modeling

Richard Zhu, Kento Nishi

arXiv 2608.00697首次发表:更新:

AI 中文总结

该研究提出进化课程学习(ECL),将其应用于VAE模型的蛋白质变异预测与RNA序列生成任务,提升了下游任务性能,证实进化距离是生物序列建模中排序训练课程的有效归纳偏置。

AI 中文摘要

基于多序列比对(MSA)训练的变分自编码器(VAE)已成为强大的生物序列生成模型,应用范围涵盖疾病变异预测到功能性RNA设计。然而,标准生物VAE训练将所有序列视为可交换的,忽略了组织同源序列的丰富进化结构——从进化上接近的到高度分化的序列。我们提出进化课程学习(ECL),这是一种利用该结构的训练策略,通过遵循幂律扩展计划,逐步向模型呈现与采样锚点进化距离递增的序列。将其应用于两种架构不同的VAE模型和两个生物领域(使用EVE的蛋白质变异效应预测、使用RfamGen的RNA家族序列生成),ECL在每种配置的5个随机种子下均提升了下游任务性能。对于p53,Mean ClinVar分类AUROC从0.981升至0.989;对于PTEN,ECL在每个种子下均达到1.000,而基线模型不稳定(均值0.905,低至0.54)。对于RNA,ECL提升了所有三个测试家族的平均协方差模型比特得分,且在15次训练运行中的12次超过了匹配种子的基线,但由于仅三个家族,该效应在家族层面无法确定为显著。消融实验表明,按进化距离逐步扩展采样序列,优于固定大小邻域采样和均匀随机采样。因此,进化距离是生物序列建模中排序训练课程的有用归纳偏置。

英文摘要

Variational autoencoders (VAEs) trained on multiple sequence alignments (MSAs) have emerged as powerful generative models for biological sequences, with applications ranging from disease variant prediction to functional RNA design. However, standard biological VAE training treats all sequences as exchangeable, ignoring the rich evolutionary structure that organizes homologous sequences from evolutionarily close to highly divergent. We propose Evolutionary Curriculum Learning (ECL), a training strategy that exploits this structure by progressively exposing the model to sequences of increasing evolutionary distance from sampled anchors, following a power-law expansion schedule. Applied to two architecturally distinct VAE models and two biological domains--protein variant effect prediction with EVE and RNA family sequence generation with RfamGen--ECL improves downstream task performance across five random seeds per configuration. Mean ClinVar classification AUROC rises from 0.981 to 0.989 for p53; for PTEN, ECL attains 1.000 in every seed whereas the baseline is unstable (mean 0.905, falling as low as 0.54). For RNA, ECL raises mean covariance-model bit scores on all three families tested and exceeds its seed-matched baseline in 12 of 15 training runs, though with only three families the effect cannot be established as significant at the family level. Ablation experiments show that progressively expanding the sampled sequences by evolutionary distance outperforms fixed-size neighborhood sampling in addition to uniform random sampling. Evolutionary distance is therefore a useful inductive bias for ordering the training curriculum in biological sequence modeling.

CommentsPublished in ICML 2026 SPIGM Workshop

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑