arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RSIBench-Data:用于递归自我改进的以数据为中心的研究基准测试

RSIBench-Data: Benchmarking Data-Centric Research for Recursive Self-Improvement

Fanqing Meng, Lingxiao Du, Qiguang Chen, Ziqi Zhao, Haocheng Lu, Mengkang Hu, Michael Qizhe Shieh

arXiv 2607.25886首次发表:更新:

发表机构

Evolvent AI; National University of Singapore(埃沃尔文特人工智能公司; 新加坡国立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究递归自我改进中以数据为中心的研究,引入RSIBench-Data基准测试,让LLM智能体迭代修订训练数据策略,评估四个前沿智能体在六个基准测试中的表现,发现其虽有能力但改进不一致,该基准测试提供了可测量的测试平台。

AI 中文摘要

递归自我改进需要将模型失败的证据转化为更好的模型。以数据为中心的训练后研究需要诊断能力差距、设计和验证训练数据策略,并从检查点反馈中学习。现有的基准测试将研究决策与优化、服务、评估和系统实现纠缠在一起,掩盖了智能体的研究能力。我们引入了RSIBench-Data,这是一个针对以数据为中心的研究人员的大型语言模型(LLM)智能体的受控基准测试,具有固定的训练后堆栈。智能体迭代地为固定目标模型修订训练数据策略;训练和服务使用Tinker支持的服务,官方评估通过Harbor和E2B沙盒运行,并且各智能体的预算固定。我们在软件工程、终端使用、科学问答和数学等六个基准测试中评估了四个前沿智能体。智能体展示了核心的以数据为中心的研究能力:在58.33%的设置中,它们通过根据反馈改进策略来改进首次有效尝试。然而,改进并不一致。在最佳观察分数之后继续进行的搜索中,78.26%以得分更低的最终尝试结束,而其余的仅恢复相同的峰值。因此,即使后期修订失败,强大的候选者也可能在运行早期或中期出现。轨迹分析在更强的运行中识别出四种模式:准确的假设、基于验证的监督、行为对齐的数据和强检查点的保留。这些发现表明,当前的智能体可以做出有用的以数据为中心的发现,但尚不能将反馈转化为一致的改进。RSIBench-Data为递归自我改进所需的研究能力提供了一个可测量、可审计的测试平台。我们在此https URL上开源了我们的代码。

英文摘要

Recursive self-improvement requires turning evidence of model failures into better models. Data-centric post-training research entails diagnosing capability gaps, designing and validating training-data strategies, and learning from checkpoint feedback. Can LLM agents automate this loop? Existing benchmarks entangle research decisions with optimization, serving, evaluation, and systems implementation, obscuring agents' research capability. We introduce RSIBench-Data, a controlled benchmark of LLM agents as data-centric researchers with a fixed post-training stack. Agents iteratively revise training-data strategies for a fixed target model; training and serving use Tinker-backed services, official evaluation runs through Harbor and E2B sandboxes, and budgets are fixed across agents. We evaluate four frontier agents on six benchmarks across software engineering, terminal use, scientific question answering, and mathematics. Agents demonstrate core data-centric research capabilities: in 58.33\% of settings, they improve upon the first valid attempt by refining strategies from feedback. However, improvement is inconsistent. Among searches continuing after the best observed score, 78.26\% end with a lower-scoring final attempt, while the rest only recover the same peak. A strong candidate may therefore appear early or midway through a run even as later revisions fail. Trajectory analysis identifies four patterns in stronger runs: accurate hypotheses, validation-grounded supervision, behavior-aligned data, and preservation of strong checkpoints. These findings suggest that current agents can make useful data-centric discoveries but cannot yet translate feedback into consistent improvements. RSIBench-Data provides a measurable, auditable testbed for the research capabilities required for recursive self-improvement. We open-source our code at https://github.com/evolvent-ai/RSIBench-Data.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑