Motif-Vocab:面向基因组语言模型的统计校准转录因子身份分词器
Motif-Vocab: StatisticallyCalibrated Transcription-Factor-Identity Tokenization forGenomic Language Models
查看机构详情
- Washington University in St. Louis(华盛顿大学圣路易斯分校)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
针对基因组语言模型分词缺乏调控先验的问题,提出Motif-Vocab,通过统计校准基序匹配生成TF身份令牌,在多数下游任务上优于随机对照,提供可解释的归纳偏置。
中文摘要 AI 辅助
分词是基因组语言模型中的一个核心设计选择,然而大多数脱氧核糖核酸(DNA)分词器使用字符、固定长度的k-mer或基于频率的子词,而没有显式利用关于DNA结合调控因子特异性的先验信息。我们提出了Motif-Vocab,一种基于生物信息的分词器,它扫描DNA双链以寻找统计校准的基序匹配,生成转录因子(TF)身份令牌,并对未匹配的序列应用核苷酸、k-mer或字节对编码(BPE)。基序特异的零分布将不同长度和简并性的位置权重矩阵(PWM)置于共同的显著性尺度上;确定性重叠规则使表示具有可复现性。在二十亿碱基对上的受控双向编码器表示从变换器(BERT)预训练中,真实基序库在55项范围内下游任务中的54项上优于随机化基序对照。在源自DART-Eval任务2的基序不相交识别任务上,TF特异性令牌相比位置匹配的通用基序令牌将宏F1提高了0.040,相比匹配的无基序分词器提高了0.033(95%自助法置信区间:0.027--0.038)。基序令牌还获得更强的归因,并产生比打乱对照更大的遮挡效应。密集的无基序分词器仍然是强大的通用基线,包括在五项任务的BERT-base面板上的近乎平局。因此,Motif-Vocab并非通用的精度替代品;它是一种针对基序敏感基因组建模的定向、可解释的归纳偏置。
英文摘要
Tokenization is a central design choice in genomic language models, yet most deoxyribonucleic acid (DNA) tokenizers use characters, fixed-length k-mers, or frequency-derived subwords without explicitly using prior information about the specificity of DNA-binding regulatory factors. We introduce Motif-Vocab, a biologically informed tokenizer that scans both DNA strands for statistically calibrated motif matches, emits transcription-factor (TF) identity tokens, and applies nucleotide, $k$-mer, or byte-pair encoding (BPE) to unmatched sequence. Motif-specific null distributions put position-weight matrices (PWMs) of different lengths and degeneracy on a common significance scale; deterministic overlap rules make the representation reproducible. In controlled Bidirectional Encoder Representations from Transformers (BERT) pretraining on two billion base pairs, real motif libraries outperform randomized-motif controls on 54 of 55 in-scope downstream tasks. On a motif-disjoint recognition task derived from DART-Eval Task 2, TF-specific tokens improve macro-F1 by 0.040 over a position-matched generic motif token and by 0.033 over a matched no-motif tokenizer (95\% bootstrap confidence interval: 0.027--0.038). Motif tokens also receive stronger attribution and produce larger occlusion effects than shuffled controls. Dense no-motif tokenizers remain strong general-purpose baselines, including a near-tie on the five-task BERT-base panel. Thus, Motif-Vocab is not a universal accuracy replacement; it is a targeted, interpretable inductive bias for motif-sensitive genomic modeling.