arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.24252cs.AIcs.SE

SA-Bench:评估基于大语言模型的论文复现中的语义对齐

SA-Bench: Evaluating Semantic Alignment in LLM-Based Paper Reproduction

Xue Hu, Zewei Pan, Zeli Su, Zhou Liu, Wentao Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

本研究推出SA-Bench基准,评估LLM复现论文时的语义对齐,发现现有模型复现准确率低,需优化脚手架以提升语义规范验证能力。

中文摘要 AI 辅助

大语言模型(LLM)智能体可生成论文复现代码,但常产生科学上不忠实的实现,我们将这种失败模式定义为语义漂移,即生成的代码与论文规范存在隐秘偏差。我们推出SemanticAlign-Bench(SA-Bench),这是一个诊断基准,涵盖2025年ICLR、ICML和NeurIPS的30篇论文。针对每篇论文,我们将其规范分解为原子化且可验证的实现声明,称为语义对齐单元(SAU),并沿四个诊断维度评估复现库:数值漂移、方法漂移、协议漂移和顺序漂移。我们在五个机器学习领域共构建1491个SAU,评估12种生成器配置(4种模型×3种脚手架)。即便最强配置(Claude+PaperCoder)的平均SAU得分仅为1.0分中的0.301,360次评估的总平均得分仅为0.221。失败分类显示,智能体尝试了大部分要求但实现错误,实现不匹配和桩函数占零分声明的大多数。我们的分析进一步表明,针对可执行性优化的脚手架对科学复现的帮助有限;缩小差距需要优先考虑语义规范验证的脚手架。该基准、注释和评估流程公开可用。

英文摘要

LLM agents can generate paper reproduction code, yet often produce scientifically unfaithful implementations. We define this failure mode as semantic drift, where generated code silently diverges from the paper's specifications. We introduce SemanticAlign-Bench(SA-Bench), a diagnostic benchmark covering 30 papers from ICLR, ICML and NeurIPS 2025. For each paper, we decompose its specifications into atomic and verifiable implementation claims, which we call Semantic Alignment Units (SAUs) and evaluate repositories along four diagnostic dimensions spanning numerical, methodological, protocol and ordering drift. In total, we construct 1,491 SAUs across five ML domains and evaluate 12 generator configurations (4 models $\times$ 3 scaffolds). Even the strongest configuration (Claude+PaperCoder) achieves a mean SAU score of only 0.301 out of 1.0, with an overall mean of 0.221 across 360 evaluations. A failure taxonomy reveals that agents attempt most requirements but implement them incorrectly, with implementation mismatch and stubs accounting for the majority of zero-scored claims. Our analysis further indicates that scaffolds optimized for executability provide limited leverage for scientific reproduction; narrowing the gap requires scaffolds that prioritize semantic specification verification. The benchmark, annotations and evaluation pipeline are publicly available.

发表机构

  • Beihang University(北京航空航天大学)
  • Shanghai Jiao Tong University(上海交通大学)
  • Minzu University of China(中央民族大学)
  • Peking University(北京大学)
  • Zhongguancun Academy(中关村科学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑