自动化研究如何被评估?基准与评估实践的调查
How Is Automated Research Evaluated? A Survey of Benchmarks and Evaluation Practices
浏览论文内容
中文总结 AI 辅助
本调查针对分散的自动化研究评估,综述六大评估目标,对比评估设计要素,总结三项经验,识别缺口并给出建议,助力读者梳理评估、选基准及设计后续研究。
中文摘要 AI 辅助
自动化研究系统支持文献综合、构思、实验、写作和同行评审,但其评估分散在不同任务、基准和研究中,难以直接比较。我们从评估设计与证据的角度综述相关文献,涵盖六个目标:文献综合、研究构思、可执行工作流、学术写作与交流、自动同行评审及端到端研究。我们对比任务构建、证据来源、评估者与评分流程,以说明不同设计所评估的能力。我们的综合研究强调三个反复出现的经验:输出检查、流程检查与人类研究提供互补信息;评估者校准与所评估的属性相关;资源预算与尝试选择是解读性能比较的关键。我们识别出诊断性评估设计及支持证据中的记录缺口,并将这些对比转化为特定评估场景的报告与审计建议。该调查帮助读者梳理现有评估、选择合适基准并设计后续研究。
英文摘要
Automated research systems support literature synthesis, ideation, experiments, writing, and peer review, but their evaluation is dispersed across tasks, benchmarks, and studies that are difficult to compare directly. We review this literature from the perspective of evaluation design and evidence, covering six targets: literature synthesis, research ideation, executable workflows, scholarly writing and communication, automatic peer review, and end-to-end research. We compare task construction, evidence sources, evaluators, and scoring procedures to explain the capabilities assessed by different designs. Our synthesis highlights three recurring lessons: output checks, process checks, and human studies provide complementary information; evaluator calibration is specific to the property being assessed; and resource budgets and attempt selection are integral to interpreting performance comparisons. We identify diagnostic evaluation designs and documented gaps in supporting evidence, and translate these comparisons into reporting and audit recommendations for specific evaluation settings. The survey helps readers navigate existing evaluations, select appropriate benchmarks, and design subsequent studies.
发表机构
- University of Macau(澳门大学)
机构由 AI 辅助整理,请以论文原文为准。