可表示但未学习:编码秩与交互预测下限
Representable but Unlearned: Encoding Rank and the Interaction-Prediction Floor
查看机构详情
- University of Central Florida(中佛罗里达大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本研究通过计算编码等价类与对比设计可达空间,提出无需拟合的秩检查方法,揭示固定编码下预测误差下限,并区分编码能力与模型实际表现。
中文摘要 AI 辅助
输入编码可以限制预测器能够联合复现的测量对比,即使没有单个对比被迫消失。我们从编码器的等价类和固定对比设计中计算可达到的对比空间,无需标签、损失或拟合模型;将记录的对比投影到该空间上,为这些类上的任何无限制解码器提供了一个经验误差下限。在一个140矩形siRNA交互面板上,图神经网络的仅训练特征掩码将165个端点状态合并为90个类,并将140个交互对比的秩削减至72。由此产生的下限为0.009980,占拟合模型交互平方误差的14.6%;拟合模型达到0.068335,略差于预测无交互的对照。至少恢复三个化学列可恢复满秩。在不使用掩码的情况下重新拟合完全消除了下限,但在报告协议下交互MSE仅改善了0.000017,且恢复的列仍不在任何训练输入中。在一个发布的RNA剪接预测器上,其编码在测量状态上是单射的,相同的计算返回完整设计秩1,986和恰好为零的下限。这些结果区分了编码所允许的与拟合模型所达到的;它们并未确定剩余误差的限制因素。秩检查无需拟合,并限制了在固定编码下任何数量的训练所能恢复的内容。项目仓库可在该https URL获取。
英文摘要
Input encodings can restrict which measured contrasts a predictor can jointly reproduce, even when no single contrast is forced to vanish. We compute the attainable contrast space from an encoder's equivalence classes and a fixed contrast design, without labels, loss, or a fitted model; projecting the recorded contrasts onto that space gives an empirical error floor for any unrestricted decoder on those classes. On a 140-rectangle siRNA interaction panel, a graph neural network's training-only feature mask merges 165 endpoint states into 90 classes and cuts the rank of the 140 interaction contrasts to 72. The resulting floor is 0.009980, which is 14.6% of the fitted model's interaction squared error; the fitted model reaches 0.068335, slightly worse than a control predicting no interaction at all. A minimum of three restored chemistry columns recovers full rank. Refitting without the mask removes the floor entirely, yet interaction MSE improves by only 0.000017 under the reported protocol, and the restored columns remain absent from every training input. On a released RNA-splicing predictor, whose encoding is injective on the measured states, the same computation returns the full design rank of 1,986 and a floor of exactly zero. These results separate what an encoding permits from what a fitted model achieves; they do not identify what limits the remaining error. The rank check needs no fits and bounds what any amount of training under a fixed encoding can recover. The project repository is available at https://github.com/shadi97kh/REPRESENTABLE-BUT-UNLEARNED.