SciWalker:利用算子图与执行反馈合成科学编码问题
SciWalker: Synthesizing Scientific Coding Problems with Operator Graphs and Execution Feedback
浏览论文内容
中文总结 AI 辅助
针对科学编码训练数据稀缺问题,提出SciWalker框架,通过算子图采样与执行反馈合成高质量问题,构建8178个跨5领域32子域的问题,并用GSPO强化学习提升Qwen3.5-9B的SciCode准确率9.9个百分点。
中文摘要 AI 辅助
提升大型语言模型(LLMs)的科学编码能力需要高质量的训练数据。然而,此类数据仍然稀缺,因为手动编写真实问题成本高昂且耗时,而系统性地覆盖多样化的科学领域和算法组合仍具挑战性。为解决这一问题,我们引入了SciWalker,一个通过算子链采样和执行反馈来合成科学编码问题的框架。该框架将科学库接口与操作模式相结合以实例化算子,将它们组织成算子图,并采样算子链作为计算工作流线索。在这些线索的引导下,我们采用LLMs生成具有科学依据的问题陈述、参考解决方案和测试,并使用执行反馈迭代修复失败的生成结果。通过将结构化工作流组合与验证和质量审查相结合,SciWalker实现了可扩展的任务生成,同时促进了科学依据性、计算多样性和可执行性。利用该框架,我们构建了8,178个高质量问题,涵盖5个科学领域和32个子领域。为评估其训练效用,我们使用GSPO算法对Qwen3.5-9B进行强化学习。该训练将SciCode子问题准确率提高了9.9个百分点,从29.3%提升至39.2%,并在科学代码生成、代码修复和推理基准上均有所提升。SciWalker的代码可在以下网址获取:此https URL。
英文摘要
Improving the scientific coding capabilities of large language models (LLMs) requires high-quality training data. However, such data remain scarce because manually authoring realistic problems is costly and time-consuming, while systematically covering diverse scientific domains and algorithmic combinations remains challenging. To address this, we introduce SciWalker, a framework for synthesizing scientific coding problems through operator-chain sampling and execution feedback. The framework combines scientific library interfaces with operation modes to instantiate operators, organizes them into operator graphs, and samples operator chains as computational workflow cues. Guided by these cues, we adopt LLMs to generate scientifically grounded problem statements, reference solutions, and tests, with failed generations iteratively repaired using execution feedback. By combining structured workflow composition with verification and quality review, SciWalker enables scalable task generation while promoting scientific grounding, computational diversity, and executability. Using this framework, we construct 8,178 high-quality problems spanning 5 scientific domains and 32 subdomains. To evaluate their training utility, we conduct reinforcement learning on Qwen3.5-9B using the GSPO algorithm. This training improves SciCode subproblem accuracy by 9.9 percentage points, from 29.3% to 39.2%, with gains across scientific code generation, code repair, and reasoning benchmarks. The code for SciWalker is available at https://github.com/lichenx1/SciWalker.
发表机构
- The Hong Kong University of Science and Technology(香港科技大学)
- Shanghai AI Laboratory(上海人工智能实验室)
- Tsinghua University(清华大学)
- Nanyang Technological University(南洋理工大学)
机构由 AI 辅助整理,请以论文原文为准。