发表机构
The Chinese University of Hong Kong, Shenzhen; Nanyang Technological University; Centre for Biomedical Data Science, Duke-NUS Medical School; Zhejiang University; Jurisprudence Research Association, China Law Society(香港中文大学(深圳); 南洋理工大学; 杜克-新加坡国立大学医学院生物医学数据科学中心; 浙江大学; 中国法学会法理学研究会)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
ScienceClaw提出固定参数程序自进化框架,统一任务求解、科学验证与程序更新,并通过跨23学科的基准评估持续改进能力。
AI 中文摘要
大型语言模型智能体正在加速科学自动化,然而,经过验证的执行很少能转化为持久的程序级改进,且现有评估未能在自然科学和社会科学的顺序任务中检验这一过程。我们将ScienceClaw形式化为固定参数的程序自进化,统一了任务求解、科学验证和程序更新。ScienceClaw-Eval涵盖23个学科,通过顺序流和独立重置评估来衡量科学正确性、进化收益、保留率、跨数据集迁移和进化成本。我们的框架通过多轮交互修复可执行工作流,将经过重新执行验证的失败-成功轨迹转化为关联的Skill和Operator候选,并且仅当源任务重放再现修复且独立科学任务有所改进时,才保留更新。代码可在该https URL获取。
英文摘要
Large language model agents are accelerating scientific automation, yet verified executions rarely become persistent program-level improvements, and existing evaluations do not examine this process across sequential tasks in both the natural and social sciences. We formalize ScienceClaw as fixed-parameter program self-evolution that unifies task solving, scientific verification, and program updates. ScienceClaw-Eval spans 23 disciplines and measures scientific correctness, evolutionary gain, retention, cross-dataset transfer, and evolution cost through sequential streams and independent reset evaluation. Our framework repairs executable workflows through multi-turn interaction, converts re-execution-verified failure--success trajectories into linked Skill and Operator candidates, and retains an update only when source-task replay reproduces the repair and independent scientific tasks improve. Code is available at https://github.com/beita6969/ScienceClaw.
Comments28 pages