arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.24953q-bio.QMcs.LG

超越Token:探究学习到的蛋白质表征中的高阶上位效应

Beyond Tokens: Probing Higher-Order Epistasis in Learned Protein Representations

Maryam Rahimimovassagh, Ivan Garibay, Niloofar Yousefi

首次发表
浏览论文内容

中文总结 AI 辅助

本研究引入ORBIT框架,分析蛋白质表征中的高阶上位效应,发现RIT可提升成对Token可及性,更深MLP能改善部分性能,揭示了传统指标未体现的表征层面变化。

中文摘要 AI 辅助

蛋白质适应度景观包含非线性相互作用,其中突变效应取决于其他残基。我们引入ORBIT(一种相互作用变换的阶数分辨基准测试框架),该框架将相互作用存在性、表征可及性与功能恢复分离。ORBIT首先在具有已知相互作用阶数的合成景观上验证基于Walsh的诊断方法,随后在FLIP 2-vs-rest设置下分析实验测量的GB1适应度景观。我们比较了岭回归、标准MLP、独立Token、非线性独立Token及残差交互Token化(RIT)。在20个配对训练随机种子中,主要的两层隐藏层对比发现,FLIP测试R²、三阶或四阶功能恢复、最终层三阶或四阶可及性方面,架构间无显著差异;但相对于独立Token对照,RIT在Token阶段显著提高了成对可及性(ΔA_tok,2=0.2468,d_z=1.67,经Holm校正的p=1.14×10^-5),未检测到下游高阶优势。预先指定的深度/容量分析显示,更深的MLP改善了FLIP预测、三阶功能恢复及最终层三阶可及性;四阶可及性相对于浅层MLP也有所提升,但绝对保留R²仍低于零。因此ORBIT揭示了传统预测指标所掩盖的表征层面变化,并区分了早期感知交互的编码与下游非线性容量构建的高阶结构。

英文摘要

Protein fitness landscapes contain nonlinear interactions in which mutation effects depend on other residues. We introduce ORBIT, an Order-Resolved Benchmarking of Interaction Transformations framework that separates interaction presence, representation accessibility, and functional recovery. ORBIT first validates Walsh-based diagnostics on synthetic landscapes with known interaction order, then analyzes the experimentally measured GB1 fitness landscape under the FLIP 2-vs-rest setting. We compare ridge regression, a standard MLP, independent tokens, nonlinear independent tokens, and Residual Interaction Tokenization (RIT). Across 20 paired training seeds, the primary two-hidden-layer comparison found no significant architecture differences in FLIP test R^2, third- or fourth-order functional recovery, or final-layer third- or fourth-order accessibility. However, RIT significantly increased pairwise accessibility at the token stage relative to both independent-token controls (Delta A_tok,2 = 0.2468, d_z = 1.67, Holm-adjusted p = 1.14 x 10^-5), without a detectable downstream higher-order advantage. A pre-specified depth/capacity analysis showed that deeper MLPs improved FLIP prediction, third-order functional recovery, and final-layer third-order accessibility; fourth-order accessibility also improved relative to the shallow MLP but remained below zero in absolute held-out R^2. ORBIT therefore reveals representation-level changes hidden by conventional prediction metrics and distinguishes early interaction-aware encoding from higher-order structure constructed by downstream nonlinear capacity.

发表机构

  • University of Central Florida(中佛罗里达大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑