arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

学习蛋白质折叠能否泛化到更广泛的推理?

Does Learning Protein Folding Generalize to Broader Reasoning?

Yong Liu, Zhanpeng Shi, Yizhou Dang, Zhongyue Zhang, Xiaoliang Shi, Zhijian Wei, Shuangjia Zheng

arXiv 2609.38879首次发表:更新:

发表机构

Shanghai Jiao Tong University; Fudan University; Shanghai Innovation Institute; Northeastern University(上海交通大学; 复旦大学; 上海创新研究院; 东北大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过蛋白质折叠任务构建FoldingCorpus数据集和Fold2Reason后训练方法,利用离散结构与连续几何双信号,显著提升模型在空间、图、科学及通用推理等10项基准上的表现,证明结构密集科学数据可系统增强语言模型的广泛推理能力。

AI 中文摘要

大型语言模型高度依赖人类文本,而人类文本往往传达表面答案,而非其背后的空间和结构逻辑。蛋白质折叠是一个天然的测试平台,因为一个已解析的结构可以产生数千条可精确检验的空间和拓扑陈述。我们提出疑问:学习折叠蛋白质能否教会通用模型可复用的推理能力?为回答这一问题,我们构建了FoldingCorpus,一个源自蛋白质的问答数据集,以及Fold2Reason,一种通过两种互补信号在其上进行后训练的配方:通过模型原生语言头预测的离散结构答案,以及从相同共享表示解码的连续三维几何。在FoldBench上,Fold2Reason的结构预测得分是Qwen3.5-9B的2.7至3.5倍。除了蛋白质结构预测,它还提升了涵盖空间、图、科学和通用推理的所有10个基准的性能,将宏平均准确率从45.09%提高到48.33%(+3.23个百分点),在所有10个基准上均获得正向收益,而由随机、合成和打乱结构构建的匹配对照组则产生显著较小或负向的收益。我们的工作表明,非语言、结构密集的科学数据能够系统地提升语言模型中的广泛推理能力,使一个已解决的科学问题成为后训练监督的实用来源。

英文摘要

Large language models rely heavily on human text, which often conveys surface answers rather than the spatial and structural logic behind them. Protein folding is a natural testbed, because one solved structure yields thousands of exactly checkable spatial and topological statements. We ask: can learning to fold proteins teach general models reusable reasoning capabilities? To answer this, we build FoldingCorpus, a protein-derived question-answer dataset, and Fold2Reason, a recipe that post-trains on it through two complementary signals: discrete structural answers predicted via the model's native language head, and continuous 3D geometry decoded from the same shared representations. On FoldBench, Fold2Reason achieves structure prediction scores 2.7 to 3.5 times those of Qwen3.5-9B. Beyond protein structure prediction, it improves performance on all 10 benchmarks spanning spatial, graph, scientific, and general reasoning, raising macro-average accuracy from 45.09% to 48.33% (+3.23 pp), with positive gains on all 10 benchmarks, while matched controls built from random, synthetic, and shuffled structure yield substantially smaller or negative gains. Our work shows that non-linguistic, structure-dense scientific data can systematically improve broad reasoning in language models, making a solved scientific problem a practical source of post-training supervision.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑