发表机构
HKUST; SLAI; NTU(香港科技大学; SLAI; 南洋理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
AutoDataBench通过逐样本验收标准评估智能体自主编写训练数据的能力,发现当前智能体虽能生成合格任务但效率低下,为递归自我改进中的数据合成提供了直接衡量基准。
AI 中文摘要
近期语言模型能力的提升更多来自数据而非架构。前沿实验室和数据公司生产可验证的智能体任务,这些任务随后通过监督微调和强化学习转化为模型能力。然而,这条生产线仍依赖人工劳动和人机协作。自动化任务创建将使数据生产随计算力而非专家人数扩展,能扩展到更多领域,并实现递归自我改进(RSI)的关键一步。当前评估智能体编写此类任务能力的方法,是衡量模型在基于智能体产出的数据训练后的表现。这与数据行业的常见实践不符,在数据行业中,数据是逐样本交付的,每个样本需根据一组标准被接受,而非直接投入训练。现有评估均未询问单个任务是否满足数据管道的验收标准。因此,我们引入AutoDataBench。给定一个原始基准任务和目标模型尝试该任务的记录,智能体必须为同一套件编写一个新任务,该任务需在有效性、新颖性、难度和行为覆盖方面满足实际验收标准。在三个可执行智能体任务的基准上,在默认的45分钟时间预算下,我们评估的所有智能体得分均未超过100分中的20分。给予最强智能体四倍的时间可大幅提高其得分,而一个可用任务的成本几乎保持不变。当前智能体可以编写所需质量的训练任务,但效率不高。AutoDataBench直接衡量了智能体自主数据合成能力:一次一个工件,根据生产管道会应用的标准进行评判,且无需训练运行。代码和数据可在该https URL获取。
英文摘要
Recent gains in language model capability have come more from data than from architecture. Frontier labs and data companies produce verifiable agentic tasks, which supervised finetuning and reinforcement learning then turn into capability.This production line still rests on human labour and on human-in-the-loop collaboration. Automating task creation would let data production scale with compute rather than with expert headcount, would extend to more domains, and would enable a key step in recursive self-improvement (RSI). Current evaluations of an agent's ability to write such tasks measure how a model performs after training on what the agent produced. That does not match common practice in the data industry, where data is delivered sample by sample and each sample is accepted against a set of criteria rather than put straight into training. No existing evaluation asks whether an individual task meets the acceptance criteria of a data pipeline. We therefore introduce AutoDataBench. Given an original benchmark task and a record of the target model attempting it, an agent must write a new task for the same suite that meets practical acceptance standards on validity, novelty, difficulty and behavioural coverage. Across three benchmarks of executable agent tasks, no agent we evaluate scores above 20 out of 100 at the default time budget of 45 minutes. Giving the strongest agent four times as long improves its score substantially, while the cost of one usable task stays almost unchanged. Current agents can write training tasks of the required quality, but not efficiently. AutoDataBench provides a direct measure of an agent's capacity for autonomous data synthesis: one artifact at a time, judged against the criteria a production pipeline would apply, and without a training run. Code and data are available at https://github.com/StarDewXXX/AutoDataBench.