arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

恢复率在不同转录因子间不可比较:归因评估的机会校正

Recovery Rates Are Not Comparable Across Transcription Factors: Chance Correction for Attribution Evaluation

Hyunkyung Han, Min Jung Kim

arXiv 2609.16271首次发表:更新:

发表机构

Yonsei University College of Medicine; Research Institute of Radiologic Science(延世大学医学院; 放射科学研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究指出基因组序列模型归因评估中恢复率未经机会校正不可比较,提出闭式机会水平和校正分数,并证明现有评估方法可能失效。

AI 中文摘要

基因组序列模型的归因方法通常通过它们恢复已知基序的程度,或通过删除证据时预测的退化程度来评估。如果没有随机情况下的值,这两种分数都无法解释,并且通常不会与随机值进行比较。我们表明,这种遗漏不是精度问题,而是有效性问题。连续基序重叠的均匀机会水平为\(L/(N-L+1)\);在UniBind中的268个转录因子中,该值范围从0.0118到0.0427,仅由基序长度和窗口大小决定的3.6倍差异。对于两个因子,机会水平本身的引导区间不重叠,因此它们的原始恢复率是不可比较的量。对此进行校正消除了已发表的五个因子的三类分类:一个被报告为分辨率失败的因子获得了第二高的校正值,领先于两个阳性对照之一,而两个被报告为完全失败的因子落在机会水平或低于机会水平。我们进一步表明,基于扰动的评估可能无法满足其自身的前提条件:对于一个因子,完全掩蔽的输入仍然得分高于决策边界,并且曲线在掩蔽位置数量上不是单调的,因此其下面积不是忠实度的度量。我们提供了闭式机会水平、机会校正分数以及在任何归因计算之前运行的两个筛选。

英文摘要

Attribution methods for genomic sequence models are commonly evaluated by how much of a known motif they recover, or by how a prediction degrades as evidence is deleted. Neither score is interpretable without the value it would take by chance, and neither is routinely reported against one. We show that this omission is not a matter of precision but of validity. The uniform chance level for contiguous motif overlap is \(L/(N-L+1)\); across 268 transcription factors in UniBind it ranges from 0.0118 to 0.0427, a 3.6-fold spread determined by motif length and window size alone. For two factors the bootstrap intervals of the chance levels themselves do not overlap, so their raw recovery rates are not comparable quantities. Correcting for this dissolves a published three-way classification of five factors: a factor reported as a resolution failure attains the second-highest corrected value, ahead of one of the two positive controls, and two reported as complete failures fall at or below chance. We further show that perturbation-based evaluation can fail its own precondition: for one factor a fully masked input still scores above the decision boundary, and the curve is not monotone in the number of masked positions, so the area under it is not a measure of faithfulness. We provide chance levels in closed form, a chance-corrected score, and two screens that run before any attribution is computed.

Comments8 pages, 7 figures, 8 tables. Code and data: https://github.com/puckradi/gauge (archived at https://doi.org/10.5281/zenodo.21594374)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑