发表机构
Harvard University; Massachusetts Institute of Technology(哈佛大学; 麻省理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过扰动输入和证据冲突测试六个生物推理模型,发现Evo2和ESM3对BioReason系列贡献有限,而其他模型受益于生物输入,揭示后训练策略未确保基础模型表示有效参与推理。
AI 中文摘要
生物推理模型通过后训练将大语言模型与生物基础模型表示及生物文本相连接。它们的基准准确率被视为大语言模型对这些输入进行推理的证据。我们在六个生物推理模型上,针对DNA、蛋白质和单细胞任务检验了这一假设。我们在保持查询和其他输入不变的情况下扰动一个生物输入,构建证据冲突,将一种基因组、蛋白质或细胞的基础模型表示与另一种的文本配对,对语言模型接收的表示拟合线性探针,并分析推理轨迹与生物输入的对应关系。Evo2和ESM3对BioReason和BioReason-Pro在评估任务上的性能贡献甚微。打乱DNA序列几乎不改变BioReason疾病预测准确率,在证据冲突中,这两个模型在97.9%和99.7%的情况下遵循文本。基于Evo2和ESM3表示训练的线性探针能预测任务目标,因此这些基础模型编码了与任务相关的信息,但对BioReason和BioReason-Pro的整体性能提升有限。相比之下,基础模型输入对ChatNT、Prot2Text-V2和CellWhisperer的性能有贡献,基因句子中的差异表达基因对Cell2Sentence-Scale的性能有贡献。在BioReason-Pro的SFT和RL检查点以及42个BioReason检查点中,准确率的提升并不意味着生物输入的性能贡献更大。BioReason推理轨迹错误陈述核苷酸变化,而BioReason-Pro推理轨迹在证据冲突下描述了最终预测中省略的功能。我们发现,当前的后训练策略并不能确保基础模型表示对任务性能做出贡献。
英文摘要
Biological reasoning models use post-training to connect LLMs to biological foundation model representations and biological text. Their benchmark accuracy is taken as evidence that LLMs reason over these inputs. We test this assumption in six biological reasoning models across DNA, protein, and single-cell tasks. We perturb one biological input while holding the query and other inputs fixed, construct evidence conflicts that pair the foundation model representation of one genome, protein, or cell with the text of another, fit linear probes to the representations the language model receives, and analyze reasoning traces against the biological inputs. Evo2 and ESM3 contribute little to BioReason and BioReason-Pro performance on the evaluated tasks. Shuffling the DNA sequence barely changes BioReason disease prediction accuracy, and in evidence conflicts the two models follow the text in 97.9% and 99.7% of cases. Linear probes trained on the Evo2 and ESM3 representations predict the task targets, so these foundation models encode information relevant to the task, but provide limited overall performance improvement to BioReason and BioReason-Pro. In contrast, foundation model inputs contribute to ChatNT, Prot2Text-V2, and CellWhisperer performance, and differentially expressed genes in the gene sentence contribute to Cell2Sentence-Scale performance. Across SFT and RL checkpoints of BioReason-Pro and 42 BioReason checkpoints, increases in accuracy do not imply greater performance contributions from biological inputs. BioReason traces misstate nucleotide changes, while BioReason-Pro traces describe functions omitted from final predictions under evidence conflicts. We find that current post-training strategies do not ensure that foundation model representations contribute to task performance.