发表机构
The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究自主数据科学智能体依赖试错流程效率低的问题,提出数据科学世界模型概念及DSWorld框架,含多种技术。构建数据集并引入优化策略,实验显示其加速智能体训练和推理,在转换预测任务上超基线。
AI 中文摘要
尽管自主数据科学智能体在数据理解和决策方面能力较强,但仍严重依赖涉及昂贵计算的试错工作流程。这一瓶颈促使人们开发能够在实际执行前预测数据科学操作效果的模型。本文引入数据科学世界模型的概念,通过根据当前工作流程状态和候选操作预测环境状态转换来对数据科学执行环境进行建模。我们进一步提出了DSWorld,这是一个实用框架,结合了结构化状态构建、成本感知路由、轻量级实际执行以及用于昂贵操作的基于大语言模型的模拟器。为支持训练,我们构建了一个8K规模的转换轨迹数据集,并引入了反思世界模型优化,这是一种用于改进转换预测的误差感知强化学习策略。实验表明,DSWorld在保持竞争力的同时,将基于强化学习的智能体训练速度提高了约14倍,将基于搜索的推理速度提高了约3至6倍,并且在转换预测任务上比最强的大语言模型基线高出35.6%。代码可在该https网址获取。
英文摘要
Despite strong capabilities in data understanding and decision-making, autonomous data science agents still heavily rely on trial-and-error workflows that involve expensive computation. This bottleneck motivates models that can anticipate the effects of data science operations before real execution. In this paper, we introduce the concept of Data Science World Model, which model the data science execution environment by predicting environment state transitions conditioned on current workflow states and candidate operations. We further propose DSWorld, a practical framework that combines structured state construction, cost-aware routing, lightweight real execution, and an LLM-based simulator for expensive operations. To support training, we construct an 8K-scale transition trajectory dataset and introduce Reflective World Model Optimization, an error-aware reinforcement learning strategy for improving transition prediction. Experiments show that DSWorld accelerates RL-based agent training by approximately $14\times$ and search-based inference by approximately $3$-$6\times$ while maintaining competitive performance, and outperforms the strongest LLM baseline by 35.6% on transition prediction tasks. The code is available at https://anonymous.4open.science/r/DSWorld.