AI 中文总结
该研究针对通用LLM酶分类中EC编号二至四级准确率近乎为零的问题,提出无训练诊断基准EC-Reason-Bench,分解四个维度评估,发现外部知识优先于推理等关键结论。
AI 中文摘要
酶功能预测是一种分层的、知识密集型的蛋白质功能分类形式。现有基准暴露出一个异常现象:通用大语言模型(LLM)通常能正确预测EC编号的第一级(粗粒度),但当要求预测完整EC编号时,其第二至第四级的准确率几乎降至零,而专用模型和工具仍能正常使用。我们提出EC-Reason-Bench,一种无训练的诊断评估协议,旨在回答两个问题:为什么通用LLM在EC编号预测上得分几乎为零,以及在不更新任何权重的情况下能恢复多少损失。我们将酶分类能力分解为四个可单独测量的正交维度:输出结构、外部知识、推理结构和推理鲁棒性。我们使用推理时方法测试每个维度,对比共享的零样本基准,重现之前报道的近零性能。对多个强推理LLM的实验得出四个主要发现:第一,外部知识具有决定性,必须先于推理:统一的闭卷性能随开卷访问急剧上升,缩小了模型间的差距;第二,在闭卷设置中,级联思维和思维链(Chain-of-Thought)是否有帮助或有害取决于模型的弃权(不执行)倾向;第三,一旦有证据可用,最佳LLM设置的综合得分与简单投票最近检索邻居的EC编号无法区分;这种平局是平均的人为结果,它隐藏了对抗证据集上的巨大增益与多功能酶上的同等巨大损失。因此,对证据的推理是作为冲突邻居的仲裁者,而非知识的来源,且没有单一数字的排行榜能体现这一点;第四,准确率服从同源性可用性定律。
英文摘要
Enzyme function prediction is a hierarchical, knowledge-intensive form of protein function classification. Existing benchmarks expose an anomaly: general LLMs often get the coarse first level right, yet once asked for a complete EC number their accuracy at levels two through four drops to almost zero, while specialized models and tools stay usable. We propose EC-Reason-Bench, a training-free, diagnostic evaluation protocol built to answer two questions: why general LLMs score close to nothing on EC number prediction, and how much of that loss can be recovered without updating a single weight. We break enzyme classification ability into four orthogonal levers that can each be measured on their own: output structure, external knowledge, reasoning structure, and reasoning robustness. We test each lever with an inference-time method against a shared zero-shot baseline reproducing previously reported near-zero performance. Experiments with several strong reasoning LLMs yield four main findings. First, external knowledge is decisive and must precede reasoning: uniformly low closed-book performance rises sharply with open-book access, narrowing model gaps. Second, in closed-book settings, whether cascading and chain-of-thought help or hurt depends on a model's tendency to abstain. Third, once evidence is available the aggregate score of the best LLM setting is indistinguishable from simply voting the EC numbers of the nearest retrieved neighbors; that tie is an artifact of averaging, and it hides a large gain on adversarial evidence set against an equally large loss on multi-functional enzymes. Reasoning over evidence therefore acts as an arbiter of conflicting neighbors rather than as a source of knowledge, and no single-number leaderboard can see it. Fourth, accuracy obeys a law of homology availability.