AI 中文总结
该研究开发GenDA模型,发现其微调后变异预测性能优于同规模自回归模型,但功能生成未达预期,表明三者需单独验证。
AI 中文摘要
双向离散扩散模型天然适用于基因组建模,因为它能从两端侧翼重建缺失序列。我们开发了GenDA(基因组密度优化吸收扩散模型),其额外假设是熵引导的跨度放置会将重建压力集中在组成复杂的区域,从而同时改善下游变异效应预测和功能序列生成。我们的结果仅部分支持这一前提:经过监督微调后,2.02亿参数的GenDA模型在合并的ClinVar单核苷酸变异(SNV)上达到的AUROC为0.774,比规模相近的自回归模型高出0.103;然而,匹配的随机跨度变异达到0.777,这未提供熵引导导致ClinVar性能提升的证据。更意外的是,GenDA在零样本功能修复压力测试中失败:在启动子、增强子、外显子边界和内含子边界上,它未能始终优于精确保留3元组(3-mer)组成但打乱天然间隙的对照模型。该失败在50至500碱基对(bp)的间隙时已存在,尽管增强子的性能退化在更长间隙时会加剧。诊断结果确定了几个边界条件:熵测量的是局部序列复杂度而非功能重要性;1元组(1-mer)分词限制了物理上下文;训练跨度上限为300 bp;以及高绝对AlphaGenome保真度可与负对照归一化修复共存。这些结果表明,微调后的强变异预测、合理的损坏先验和功能生成是不同的主张,需要单独验证。
英文摘要
Bidirectional discrete diffusion model appears naturally suited to genomic modeling because it can reconstruct missing sequence from both flanks. We developed GenDA (Genomic Density-optimized Absorbing Diffusion) under the additional hypothesis that entropy-guided span placement would concentrate reconstruction pressure on compositionally complex regions, improving both downstream variant-effect prediction and functional sequence generation. Our results only partially support this premise. After supervised fine-tuning, the 202M-parameter GenDA model reaches a pooled ClinVar SNV AUROC of 0.774, exceeding a similarly scaled autoregressive model by 0.103. However, a matched random-span variant reaches 0.777, providing no evidence that entropy guidance causes the ClinVar improvement. More unexpectedly, GenDA fails a zero-shot functional inpainting stress test: across promoters, enhancers, exon boundaries, and intron boundaries, it does not consistently outperform a control that shuffles the native gap while exactly preserving 3-mer composition. Failure is already present for 50--500-bp gaps, although enhancer degradation worsens at longer gaps. Diagnostics identify several boundary conditions: entropy measures local sequence complexity rather than functional importance; 1-mer tokenization limits physical context; training spans are capped at 300 bp; and high absolute AlphaGenome fidelity can coexist with negative control-normalized restoration. These results show that strong fine-tuned variant prediction, a plausible corruption prior, and functional generation are distinct claims that require separate validation.