发表机构
XScience Lab; Wenge AI(X科学实验室; 文阁人工智能)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究科学计算经验巩固问题,引入SciConsolidate方法,通过对比成败归纳跨任务过程,经开发验证门选择,用失败告知等扩展数据,克服抽象执行差距,在SciCode上实验,为科学计算建立经验到能力途径及自我改进起点。
AI 中文摘要
大语言模型日益用于解决科学计算任务,但一个问题的可执行反馈很少能转化为后续问题的持久能力。我们研究科学计算经验巩固,即把经过验证的运行时经验转化为可转移的过程知识并持续改进模型。此设置面临两个挑战:轨迹衍生工件可能编码特定于源的修复而非跨任务计算机制;较弱的目标模型可能无法实施有效的抽象过程。我们引入SciConsolidate,它对比成功与失败以归纳跨任务过程,通过开发验证门进行选择,并使用失败告知、无答案查询合成来扩展巩固数据。由于目标模型可能无法直接执行这些抽象,更强的模型将其转化为可执行代码监督用于标准的无过程监督微调;匹配的无过程教师分支隔离过程指导的价值。在SciCode上,运行时过程注入使Qwen3.6 - 27B在子步骤/主要问题点上提高了+3.85/+6.26,但Qwen3.5 - 9B几乎没有总体主要问题增益,证明了抽象执行差距。经过过程指导的具体化后,9B学生模型在无过程部署下比无过程监督微调控制提高了+3.89/+6.25,比原始9B模型提高了+5.62/+11.25。这些结果为科学计算建立了从经验到能力的途径,并为扩展自我改进的科学辅助提供了实用起点。
英文摘要
Large language models increasingly solve scientific-computing tasks, but executable feedback from one problem rarely becomes durable capability on subsequent problems. We study scientific-computing experience consolidation: converting verified runtime experience into transferable procedural knowledge and persistent model improvement. This setting presents two challenges: trajectory-derived artifacts may encode source-specific repairs rather than cross-task computational mechanisms; and a weaker target model may be unable to operationalize an otherwise valid abstract procedure - an abstraction-execution gap. We introduce SciConsolidate, which contrasts verified successes and failures to induce cross-task procedures, selects them through a development-validation gate, and uses failure-informed, answer-free query synthesis to expand the consolidation data without requiring pre-existing reference answers. Because the target model may not directly execute these abstractions, a stronger model concretizes them into executable code supervision for standard, procedure-free SFT; a matched no-procedure teacher branch isolates the value of procedural guidance. On SciCode, runtime procedure injection improves Qwen3.6-27B by +3.85/+6.26 sub-step/main-problem points, but yields almost no aggregate main-problem gain for Qwen3.5-9B, providing operational evidence of the abstraction-execution gap. After procedure-guided concretization, the 9B student improves under procedure-free deployment by +3.89/+6.25 points over the no-procedure SFT control and by +5.62/+11.25 over the original 9B model. These results establish an experience-to-capability pathway for scientific computing and provide a practical starting point for scaling self-improving scientific assistance.