arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

测试翻译:一种基于证明器的自然语言到一阶逻辑评估指标

Quizzing the Translation: A Prover-Grounded Evaluation Metric for NL$\rightarrow$FOL

Pu Suo, Ali Emami

arXiv 2609.33612首次发表:更新:

发表机构

Emory University(埃默里大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出基于定理证明器的SIV评估指标,通过正反探针检测自然语言到一阶逻辑翻译中的内容丢失与过度断言,显著优于现有指标。

AI 中文摘要

对自然语言问题进行符号推理的标准流程是将问题翻译成一阶逻辑,然后调用定理证明器。翻译步骤是瓶颈:把“every”换成“some”,后续的所有推理都会出错。然而,当今的评估指标往往给更糟糕的翻译打更高的分,因为BLEU、BERTScore和Smatch++奖励的是表面重叠,而最严重的错误恰好保留了这种表面重叠。我们提出了SIV,它从目标公式中衍生出两类探针,并使用定理证明器来验证候选翻译是否满足每个探针。正向探针是候选翻译必须蕴含的陈述,用于检测丢失内容的翻译;对比探针是候选翻译不得蕴含的陈述,用于检测比原文断言更多的翻译。在一个受控的扰动FOLIO翻译池上,错误的严重程度解释了SIV分数方差的80%,而任何先前指标最多只能解释17%。在一个不相交的池上的六类错误中,SIV在超过99%的配对中将参考翻译排在扰动候选之上。由于每个探针都标注了其测试的内容,失败模式也提供了带标签的错误轨迹,以宏F1分数0.638恢复了扰动类别,几乎是仅基于分数的基线的两倍。在434个人工审核的真实LLM翻译上,SIV取得了最高的AUC,独特地检测并分级了专家标注的重大错误,并且在词汇表外的翻译上选择弃权(不执行),而不是错误打分。

英文摘要

A standard pipeline for symbolic reasoning over natural-language problems translates them into first-order logic and invokes a theorem prover. The translation step is the bottleneck: swap "every" for "some" and every inference that follows is corrupted. Yet today's metrics often score more broken translations higher than less broken ones, because BLEU, BERTScore, and Smatch++ reward surface overlap that the worst errors happen to preserve. We introduce SIV, which derives two kinds of probes from the target formula and uses a theorem prover to verify the candidate translation against each. Positive probes are statements the candidate must entail, which detect translations that drop content; contrastive probes are statements the candidate must not entail, which detect translations that assert more than the original. On a controlled pool of perturbed FOLIO translations, the severity of the error accounts for 80% of SIV's score variance, compared with at most 17% for any prior metric. Across six error classes on a disjoint pool, SIV scores the reference above the perturbed candidate in over 99% of pairs. Because each probe is labeled with what it tests, the failure pattern also supplies a labeled error trace, recovering the perturbation class at macro-F1 0.638, nearly double the score-only baseline. On 434 expert-audited real LLM translations, SIV attains the top AUC, uniquely detects and grades expert-labeled major errors, and abstains, rather than mis-scoring, on out-of-vocabulary translations.

CommentsAccepted to EMNLP 2026 (Main Conference). 18 pages, 3 figures, 10 tables. Code and data: https://github.com/pu-suo/siv-metric

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑