arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.40097cs.CL

AutoDataBench:一个用于加速自动研究的数据中心测试平台

AutoDataBench: A Data-centric Testbed for Accelerating Auto Research

Ruifeng Yuan, Yizhi Li, Yaxin Du, Fengyu Cai, Yiqi Liu, Hou Pong Chan, Chenghua Lin, Yun Chen, Jian Yang, Bryan Dai, Pinyan Lu, Chenghao Xiao

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出AutoDataBench,一个隔离评估数据智能的受控测试平台,通过三个优化任务验证LLM改进数据的能力,并证明其轨迹可提升下游编码性能。

中文摘要 AI 辅助

现有的自动研究基准测试常常将多种改进来源纠缠在一起,包括训练框架、超参数、计算预算和数据,这使得难以将某个前沿智能体优于另一个智能体的表现归因于特定的研究能力。在本工作中,我们隔离并系统评估了数据智能:智能体理解、操作和改进塑造模型能力的数据的能力。我们引入了AutoDataBench,一个基于数据智能概念框架的受控测试平台,该框架涵盖数据诊断、数据组织和数据构建,通过三个高度策划的优化任务实例化,同时保持非数据因素固定。在工具使用、检索和知识注入方面,我们评估了前沿LLM在特定任务资源预算下通过迭代实验改进训练数据的能力。除了优化性能,我们问:LLM是否理解其数据干预的作用?我们将训练前的预测与观察到的结果进行比较,以寻找数据效应推理的证据,而不仅仅是试错,并探索迭代反馈是否帮助LLM更好地理解训练数据的变化如何影响模型性能。最后,我们展示了重用AutoDataBench轨迹进行中期训练能提高下游编码性能,突显了其在评估数据智能和生成高质量训练数据方面的价值。代码和资源可在以下https URL获取。

英文摘要

Existing auto-research benchmarks often entangle multiple sources of improvement, including training frameworks, hyperparameters, compute budgets, and data, making it difficult to attribute why one frontier agent outperforms another to specific research capabilities. In this work, we isolate and systematically evaluate Data Intelligence: an agent's ability to understand, manipulate, and improve the data that shapes model capabilities. We introduce AutoDataBench, a controlled testbed built on a conceptual framework of data intelligence spanning data diagnosis, data organization, and data construction, instantiated through three highly curated optimization tasks while holding non-data factors fixed. Across tool use, retrieval, and knowledge injection, we evaluate frontier LLMs' ability to improve training data through iterative experimentation under task-specific resource budgets. Beyond optimization performance, we ask: do LLMs understand what their data interventions do? We compare predictions made before training with observed outcomes to seek evidence of data-effect reasoning beyond trial and error, and explore whether iterative feedback helps LLMs better understand how changes to training data affect model performance. Finally, we show that reusing AutoDataBench trajectories for mid-training improves downstream coding performance, highlighting its value in both evaluating data intelligence and generating high-quality training data. Code and resources are available at https://github.com/AutoDataBench/AutoDataBench.

发表机构

  • The Hong Kong Polytechnic University(香港理工大学)
  • IQuest Research
  • Shanghai Jiaotong University(上海交通大学)
  • TU Darmstadt(达姆施塔特工业大学)
  • The University of Manchester(曼彻斯特大学)
  • University of Macau(澳门大学)
  • Shanghai University of Finance and Economics(上海财经大学)
  • Beihang University(北京航空航天大学)

机构由 AI 辅助整理,请以论文原文为准。

↑