AI 中文总结
该研究提出知识引导的掩码策略 VESTIGE,针对基因组Transformer微调,在古DNA重建任务中,相比标准MLM显著提升性能,且原理可适用于多种降解序列处理场景。
AI 中文摘要
标准掩码语言模型(MLM)微调对每个 token 位置应用统一的掩码概率,假设重建难度与位置无关。但当降解过程具有特征且集中在可预测位置时,该假设不成立:在损伤峰值位点,模型性能可能低于频率匹配的随机预测器。我们提出 VESTIGE,这是一种无参数、可直接替换标准 MLM 整理器的方法,其将掩码分布与经验测量的逐位置腐败轮廓对齐。我们将其应用于古 DNA(aDNA)重建,其中胞嘧啶脱氨产生位置依赖的 C→T / G→A 梯度,该梯度通过 mapDamage2 逐位置量化。我们重新缩放使平均 C/G 掩码率等于 15%——与标准 MLM 相同——从而将空间重新分布作为唯一变量,在对猛犸 CDS 语料库(两个样本,七个基因)进行两次 DNABERT-2 运行时,模型、数据、随机种子和超参数均保持固定。在六个末端区域宽度和 626 对窗口上,VESTIGE 在所有宽度下均优于标准 MLM(差值为 +4.18 至 +10.35 个百分点,所有 p 值均 < 10^-8),将验证交叉熵降低 13%(3.274 对比 3.757),且在损伤放大至真实 PMD 率的 10-30 倍时,所有六个重建(三个基因)的 ESMFold 重建 TM 评分均 > 0.95。1D CNN 生物安全分类器返回 AUC = 0.935,清除了 98.2% 的重建窗口,剩余 1.76% 归因于参考基因组特征,而非重建伪影。该原理与领域无关:任何可测量的位置或上下文特定腐败轮廓——FFPE、亚硫酸氢盐、宏基因组或纳米孔——均可直接替代 PMD 阵列,使 VESTIGE 成为处理降解或噪声序列输入的智能系统的知识引导训练例程。
英文摘要
Standard masked-language-model fine-tuning applies a uniform masking probability across every token position, assuming reconstruction difficulty is position-agnostic. When the degradation process is characterised and concentrated at predictable positions, this assumption fails: at peak damage sites the model can underperform a frequency-matched random predictor. We introduce VESTIGE, a parameter-free, drop-in replacement for the standard MLM collator that aligns the masking distribution with an empirically measured per-position corruption profile. We apply it to ancient DNA (aDNA) reconstruction, where cytosine deamination produces a position-dependent C-to-T / G-to-A gradient quantified per-position by mapDamage2. Rescaling so the mean C/G masking rate equals 15% - identical to standard MLM - isolates spatial redistribution as the sole variable, with model, data, seed, and hyperparameters held fixed across both DNABERT-2 runs on a mammoth CDS corpus (two specimens, seven genes). Across six terminal-zone widths and 626 paired windows, VESTIGE leads standard MLM at every width (Delta = +4.18 to +10.35 pp, all p < 10^-8), cuts validation cross-entropy by 13% (3.274 vs. 3.757), and yields ESMFold reconstructions with TM-score > 0.95 across all six reconstructions (three genes) even under damage amplified 10-30x beyond authentic PMD rates. A 1D CNN biosecurity classifier returns AUC = 0.935 and clears 98.2% of reconstructed windows, the 1.76% remainder attributable to reference-genome features, not reconstruction artefacts. The principle is domain-agnostic: any measurable position- or context-specific corruption profile - FFPE, bisulfite, metagenomic, or nanopore - substitutes directly for the PMD array, making VESTIGE a knowledge-guided training routine for intelligent systems operating on degraded or noisy sequence inputs.
Comments18 pages, 9 figures, 6 tables