arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.12788cs.AI

ARAC:针对端到端研究的自动研究对齐性与完整性基准测试

ARAC: Benchmarking Auto-Research's Alignment and Completeness on End-to-End Researchs

Jiale Cui, Yueyao Yuan, Kaixi Zhong, Xiaogang Xu, Jiafei Wu, Zhe Liu

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出ARAC-Bench基准框架,通过学术认知技能系统与三阶段诊断协议评估自动研究的对齐性等,对11个SOTA框架评估发现存在显著差距,且与博士候选人排名相关性强,为自主研究系统提供工具与奖励信号。

中文摘要 AI 辅助

自动研究领域的快速发展带来了一项基础性评估挑战:我们如何衡量其研究轨迹与人类研究行为的对齐性、逻辑连贯性及进化完整性?本文提出了自动研究的对齐性与完整性基准测试框架ARAC-Bench,这是一种模拟研究者的评估框架,其评估目标从匹配最终结果转向复现高质量的人类研究过程。该框架通过两个协同组件运行:学术认知技能系统,它首次将隐性的评审专家专业知识转化为分阶段校准的可量化评分标准;以及三阶段能力诊断协议,它在严格的模块化约束下将研究过程分解为三个可追踪、相互独立的维度:提案、实验和综合。对11个SOTA框架的系统评估显示,最高对齐分数仅为100分中的67.9,这表明在模拟严谨的人类研究方法方面存在显著差距。针对博士候选人排名的验证显示出0.8141的强相关性,证实ARAC-Bench可靠地反映了研究者真正重视的维度。ARAC-Bench不仅提供了细粒度的诊断工具,还为训练下一代自主研究系统提供了可扩展的奖励信号。

英文摘要

The rapid advancement of Auto-Research has surfaced a fundamental evaluation challenge: how can we measure the alignment, logical coherence, and evolutionary completeness of its research trajectory with human research behavior? We propose Auto-Research's Alignment and Completeness, ARAC-Bench: a Researcher-Mimicking Evaluation framework that shifts the objective from matching final answers to reproducing high-quality human research processes. The framework operates through two synergistic components: the Academic Cognition Skills system, which is the first to transforms implicit reviewer expertise into stage-calibrated, quantifiable rubrics; and a three-stage capability diagnostic protocol, which decomposes the research process under strict modular constraints into three traceable, mutually independent dimensions: Proposal, Experiment, and Synthesis. Systematic evaluation of 11 SOTA frameworks yields a best alignment score of only 67.9 of 100, revealing a significant gap in simulating rigorous human methodology. Validation against Ph.D. Candidates rankings shows a strong correlation of 0.8141, confirming that ARAC-Bench reliably reflects the dimensions researchers truly value. ARAC-Bench provides not only a fine-grained diagnostic tool but also a scalable reward signal for training the next generation of autonomous research systems.

补充信息

↑