arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SciDataSailor:深度科学数据探索

SciDataSailor: Deep Scientific Data Exploring

Jiyong Rao, Yicheng Qiu, Chi Zhang, Chunfeng Song, Runkai Zhao

arXiv 2607.28098首次发表:更新:

AI 中文总结

该研究针对科学数据仓库交互难题,提出SciDataSailor框架,以带特定机制的MCTS实现轨迹合成,构建了微调模型与含千余任务的评估基准,推动LLM智能体的科学数据探索能力。

AI 中文摘要

科学数据集通常被组织为包含异构且相互依赖文件的分层仓库,这使得对其进行检查、整合和分析需要大量人力并依赖领域专业知识。尽管大型语言模型(LLM)智能体在规划、推理和工具使用方面已取得显著进展,但现有研究在很大程度上忽略了它们通过可执行环境与真实科学数据资产交互的能力。我们提出深度科学数据探索这一智能体任务范式,其中智能体在仓库中导航、解释异构文件和模式、执行分析、整合跨文件证据并基于执行的观察结果生成结论。为实现这一范式,我们提出SciDataSailor框架,该框架通过平衡广泛探索与针对性利用来合成工具交互轨迹。SciDataSailor将轨迹合成为蒙特卡洛树搜索(MCTS),并带有四个特定任务机制:难度分层的探索种子、双反馈首玩紧迫性、分层策略到工具的动作生成以及熵引导的分支。使用该框架,我们构建了用于监督微调的SciDataSailor-SFT-2K和用于评估的SciDataSailor-Bench,后者包含627个元信息摘要任务和586个科学问答任务,覆盖生命科学、地球科学和物理科学领域的27个数据集。

英文摘要

Scientific datasets are commonly organized as hierarchical repositories containing heterogeneous and interdependent files, making their inspection, integration, and analysis labor-intensive and reliant on domain expertise. Although large language model (LLM) agents have advanced substantially in planning, reasoning, and tool use, existing research has largely overlooked their ability to interact with real scientific data assets through executable environments. We introduce Deep Scientific Data Exploration, an agentic task paradigm in which agents navigate repositories, interpret heterogeneous files and schemas, execute analyses, integrate cross-file evidence, and produce conclusions grounded in executed observations. To operationalize this paradigm, we present SciDataSailor, a framework for synthesizing tool-interactive trajectories by balancing broad exploration with targeted exploitation. SciDataSailor instantiates trajectory synthesis as Monte Carlo Tree Search (MCTS) with four task-specific mechanisms: difficulty-stratified exploration seeds, dual-feedback first-play urgency, hierarchical strategy-to-tool action generation, and entropy-guided branching. Using this framework, we construct SciDataSailor-SFT-2K for supervised fine-tuning and SciDataSailor-Bench for evaluation, with the latter comprising 627 meta-information summarization tasks and 586 scientific question-answering tasks across 27 datasets spanning the life, earth, and physical sciences.

Comments63 pages, 10 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑