arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向回合级智能体强化学习的科学发现环境扩展

Scaling Scientific Discovery Environments for Turn-Level Agentic RL

Yucheng Xu, Keyi Zhang, Yuyang Yu, Min Zhang, Shiyuan Meng, Pei Chu, Zhongying Tu

arXiv 2607.28990首次发表:更新:

发表机构

Shanghai AI Lab(上海人工智能实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出可扩展框架SciDisco,结合SciThèque与DiscoPO,训练出的SciDisco-14B在假设驱动的科学数据分析基准上实现了当前最优性能。

AI 中文摘要

大型语言模型智能体在数据驱动的科学发现任务中展现出良好能力,这类智能体与执行环境交互并生成统计结论。长时程科学分析因缺乏针对真实科学数据的过程监督环境而受限。本文提出SciDisco,这是一个用于在可过程验证环境中训练科学发现智能体的可扩展框架。SciThèque将假设、数据集、隐藏证据图和验证器编译为任务环境,可在交互过程中检查分析进展。基于有向无环图(DAG)的轨迹合成利用这些环境构建经验证器筛选的多回合演示。DiscoPO则将环境作为训练信号来源,为生成可验证分析证据的动作分配回合级信用。实验显示,SciDisco-14B在假设驱动的科学数据分析基准上达到了当前最优性能。

英文摘要

Large language model agents have shown promising capabilities in data-driven scientific discovery tasks, where an agent interacts with an execution environment and produces a statistical claim. Long-horizon scientific analysis remains constrained by the lack of process supervised environments over real-world scientific data. This paper introduces SciDisco, a scalable framework for training Scientific Discovery agents in process-verifiable environments. SciThèque compiles hypotheses, datasets, hidden evidence graphs, and verifiers into task environments where analytical progress can be checked during interaction. DAG-grounded trajectory synthesis uses these environments to construct verifier-filtered multi-turn demonstrations. DiscoPO then uses the environment as the source of training signal, assigning turn-level credit to actions that produce verifiable analytical evidence. Experiments show that SciDisco-14B reaches state-of-the-art on hypothesis-driven scientific data analysis benchmarks.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑