arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.29270cs.CLcs.AI

SHADOWBENCH:迈向自动形式化中语义对齐的可靠自动评估

SHADOWBENCH: Toward Reliable Automatic Evaluation of Semantic Alignment in Autoformalization

  • Electronics and Telecommunications Research Institute(电子通信研究院)
  • Seoul National University(首尔大学)
  • University of Maryland, College Park(马里兰大学帕克分校)
  • Amazon Web Services(亚马逊网络服务)

机构由 AI 辅助整理,请以论文原文为准。

Hojae Han, Jongyoon Kim, Sanghyeok Park, Dongwook Cheon, Yeachan Park, Myeong Jae Jeon, Sunjong Choe, Soonho Kong, Wonseok Hur, Seung-won Hwang, Donghoon Hyeon

AI总结:

针对自动形式化评估的缺陷,该研究提出SA-Pass方法并构建ShadowBench基准,实验显示其与专家判断一致性达98.8%,且被用作ICML 2026 AI4Math挑战赛第4赛道基准。

AI中文摘要:

自动形式化将非正式数学定理转换为Lean等证明助手的代码。当前评估指标的核心挑战在于,其可能接受类型正确但对齐错误的陈述,或拒绝表述不同的正确陈述。受Pass@k启发,我们提出SA-Pass(语义对齐Pass),该方法使用称为“影子(shadows)”的辅助陈述来表征目标陈述。生成的陈述仅在满足编译、蕴含每个影子(正向检查)且被影子的合取所蕴含(反向检查)时,才能获得满分。我们将SA-Pass实例化为ShadowBench,这是一个包含178道研究生至研究级问题的Lean 4全自动形式化基准,涵盖8个数学领域。配备Numina-Lean-Agent的Claude Code(Opus 4.8)达到61.8%的编译率和11.2%的SA-Pass得分。在6种智能体配置生成的输出中,SA-Pass与专家判断的二元一致性达98.8%。ShadowBench的早期版本曾作为2026年ICML AI4Math挑战赛第4赛道的基准。

英文摘要:

Autoformalization translates informal mathematical theorems into code for proof assistants such as Lean. A central challenge is that current evaluation metrics can accept type-correct but misaligned statements or reject correct statements written in a different formulation. Inspired by Pass@$k$, we propose SA-Pass (*Semantic Alignment Pass*), which tests formal statements using auxiliary statements called *shadows* that characterize the intended statement. A generated statement receives full credit only when it compiles, implies each shadow (forward check), and is implied by their conjunction (backward check). We instantiate SA-Pass in ShadowBench, a Lean 4 full autoformalization benchmark of 178 postgraduate- to research-level problems spanning eight mathematical areas. Claude Code (Opus 4.8) with Numina-Lean-Agent reaches $61.8\%$ compile rate and $11.2\%$ SA-Pass. Across outputs generated by six agentic configurations, SA-Pass achieves $98.8\%$ binary agreement with expert judgments. An early version of ShadowBench served as the benchmark for Track 4 of the ICML 2026 AI4Math Challenge.

补充信息

↑