arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31927cs.IRcs.CLcs.LG

配方匹配,而非等价

Recipe-Matching, Not Equivalence

Ali Habibullah, Mohammad Alshiekh, Yazan Alshoibi, Salman Khan, Naeemullah Khan

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过对比 LLM 编写与计算机验证的配对,揭示 MathNet-Retrieve 基准中配方匹配带来的分数提升主要源于 LLM 生成而非真实能力,且该提升在真实重复项上消失,并伴随保留率下降。

中文摘要 AI 辅助

MathNet-Retrieve 要求检索器针对一个数学问题,找到陈述同一问题的文档。一个 LLM 在固定的提示词下编写每份黄金文档及其近似错误干扰项;LLM 评判者对其过滤。我们将此过程称为“配方”,在以此方式构建的配对上进行训练称为“配方匹配”,并询问该过程在基准声称测试的能力之外能带来多少分数提升。来自同一基础模型的两个模型,在行和设置上匹配,仅训练文件不同:一组配对是在基准发布的提示词下由另一供应商的 LLM 和评判者编写的,另一组配对是计算机代数验证的,全程无 LLM 参与。前者在简单层级上领先 45 个 R@1 点。通过非 LLM 释义对照,该差距的一半到三分之二源于配对本身由 LLM 编写:在无关提示词下的 LLM 重写,使用验证模型的负样本,恢复了 45 点中的 30 点和 22 点;使用相同负样本的回译几乎未恢复任何点。剩余的 15 到 25 点仅出现在基准自身提示词下,并在无生成器编写的真实重复项(同一问题两种语言)上消失。困难层级奖励配方的配对结构,即针对最小编辑近似错误的深度重写:仅 LLM 重写在该层级上得分为零,附加负样本解锁了该层级,且每个实现此效果的负样本都会损失跨语言点数;在该层级上得分最高的集合,对无 LLM 编写的近似错误的区分能力,差于带验证负样本的 LLM 重写。当模型以其自身方式构建的配对训练时,MELD 也会移动,且不损失保留率;在 SABER-Math 上,注册的攻击失败,唯一的增益来自其 LLM 编写的摘要,该增益虽小但在匹配预算下保持。仅在 MathNet-Retrieve 上,我们能够将基准分数上升而真实保留率下降的反转归因于训练文件的单次编辑。我们发布了无生成器的重复项评估、近似错误测试和三个训练模型。

英文摘要

MathNet-Retrieve asks a retriever to find, for a math problem, a document stating the same problem. An LLM under one fixed prompt writes each gold document and its near-miss distractors; LLM judges filter them. We call this procedure the "recipe", training on pairs built the same way "recipe-matching", and ask how much score it buys beyond the ability the benchmark claims to test. Two models from one base, matched in rows and settings, differ only in the training file: pairs written under the benchmark's published prompt by another vendor's LLM and judge, or computer-algebra-verified pairs with no LLM anywhere. The first leads by 45 R@1 points on the easy tier. By a non-LLM paraphrase control, half to two thirds of that gap comes from the pairs being LLM-written at all: LLM rewrites under two unrelated prompts, with the verified model's negatives, recover 30 and 22 of the 45 points; back-translations with the same negatives recover almost none. The remaining 15 to 25 points appear only under the benchmark's own prompt and vanish on real duplicates no generator wrote, the same problem in two languages. The hard tier rewards the recipe's pair structure, a deep rewrite against a minimal-edit near-miss: LLM rewrites alone score zero on it, attaching negatives unlocks it, and every negative that does so costs cross-language points; the sets scoring highest on it separate near-misses no LLM wrote worse than LLM rewrites with verified negatives. MELD also moves when a model trains on pairs built its way, without losing retention; on SABER-Math the registered attack fails, and the one gain, from its LLM-written summaries, is small but holds at a matched budget. Only on MathNet-Retrieve could we pin an inversion, benchmark score up and real retention down, to one edit of a training file. We release the generator-free duplicate evaluations, the near-miss test and three trained models.

发表机构

  • KAUST Academy(阿卜杜拉国王科技大学学院)
  • King Abdullah University of Science and Technology (KAUST)(阿卜杜拉国王科技大学)
  • Oxford Brookes University(牛津布鲁克斯大学)
  • Lady Margaret Hall, University of Oxford(牛津大学玛格丽特夫人学堂)

机构由 AI 辅助整理,请以论文原文为准。

↑