arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DataSense-Bench:迈向AI科学家的第一步

DataSense-Bench: The First Step Toward an AI Scientist

Yudi Zhang, Mingyu Cao, Lu Yin, Mykola Pechenizkiy, Shiwei Liu

arXiv 2610.12190首次发表:更新:

发表机构

Eindhoven University of Technology; University of Surrey; ELLIS Institute Tübingen; Max Planck Institute for Intelligent Systems; Tübingen AI Center(埃因霍温理工大学; 萨里大学; 图宾根ELLIS研究所; 马克斯·普朗克智能系统研究所; 图宾根人工智能中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究推出DataSense-Bench基准,评估AI智能体选择训练数据的能力,实验发现智能体数据选择增益有限、排序能力不稳定,且对数据信号的价值解读存在差异。

AI 中文摘要

随着关于递归自我改进(RSI)和通用人工智能(AGI)的说法不断增多,我们提出一个简单问题:前沿AI模型是否具备数据感知能力,即它们能否可靠地为训练选择合适的数据?我们推出DataSense-Bench,通过机器学习中数据选择与性能预测这一基础问题来研究该能力。我们要求AI智能体选择并排序可用于微调小型LLM模型的候选训练子集,允许智能体检查数据、编写并执行分析代码、运行模型前向传播,但不允许训练模型或访问实际评估任务。随后我们在每个选定子集上微调基础模型,并在标准化协议下评估其训练后性能。我们在终端问题解决和工具使用场景中实例化该基准,分别从OpenThoughts-Agent和EnvScaler中选择轨迹,并在TBLite和BFCL上进行评估。接着我们从两个互补维度评估智能体:排名最高子集的训练后性能,反映识别高价值训练数据的能力;以及排序准确性,反映预测所选子集相对性能的能力。实验中,智能体相较于随机选择的增益有限,无法可靠地对所选组进行排序,且排序能力在不同任务间无法保持一致:Astra在全部3次工具使用运行中识别出最佳组,但在全部3次终端运行中仅识别出1次。对两个任务的执行轨迹分析显示,智能体常使用相似的数据信号,但对其训练价值的解读存在差异。

英文摘要

As claims about recursive self-improvement (RSI) and artificial general intelligence (AGI) proliferate, we ask a simple question: do frontier AI models have a sense of data, i.e., can they reliably select the right data for training? We introduce DataSense-Bench to study this capability through the fundamental problem of data selection and performance forecasting in machine learning. We ask AI agents to select and rank candidate training subsets that can be used to fine-tune a small LLM model. Agents are allowed to inspect the data, write and execute analysis code, and run model forward passes, but can not train the model or access the actual evaluation tasks. We then fine-tune the base model on each selected subset and evaluate its post-training performance under a standardized protocol. We instantiate the benchmark in terminal problem solving and tool use, selecting trajectories from OpenThoughts-Agent and EnvScaler and evaluating on TBLite and BFCL, respectively. We then evaluate the agents along two complementary dimensions: the post-training performance of the top-ranked subset, reflecting the ability to identify high-value training data, and ranking accuracy, reflecting the ability to predict the relative performance of the selected subsets. In our experiments, selection gains over random selection are limited; agents do not reliably rank their selected groups, and ranking ability does not hold consistently across tasks: Astra identifies the best group in all three tool-use runs but in only one of three terminal runs. Analysis of execution traces on both tasks shows that agents often use similar data signals while interpreting their training value differently.

Comments25 pages. Project page: https://datasense-bench.github.io/ . Code: https://github.com/DataSense-Bench/DataSense-Bench

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑