发表机构
Scale AI(Scale AI)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究重建了 ERG 的 MRS 生成与解析基准,评估 Claude 等大模型在双向任务上的表现,发现生成可达高 BLEU 但解析远逊于 ACE,表明生成分数不足以证明模型理解形式语义。
AI 中文摘要
英语资源语法(ERG)是一部手写的英语计算语法。给定一个句子,其处理器 ACE 会生成一种称为最小递归语义(MRS)的形式化意义表示:即句子谓词及其论元的图。该语法是双向的,也能将 MRS 转换回英语句子。\citet{hajdik2019} 利用 ERG 的树库构建了一个用于该生成任务(MRS 到文本)的基准,并训练了序列到序列模型来解决该任务。解析任务(文本到 MRS)可以在相同的句子上进行测试。我们重建了他们 10K 句子的测试集划分,并在两个方向上对两个大型语言模型——Claude Sonnet~4.5 和 Claude Opus~5——进行评分,与它们训练过的系统以及 ACE 进行对比,且不进行任何任务特定训练。给定一个 MRS 和三个示例,Opus 以 76.3 BLEU 写出句子,比在 72k 对(66.1 BLEU)上训练的系统高出 10 分,与在额外一百万对(77.2 BLEU)上训练的系统相当。Sonnet 得分为 65.7 BLEU,而让它在 ACE 自身的候选句子中进行选择则将其提升至 69.6,而一个保留 Opus 自身句子在候选中的池化评判者则增加了 0.6 分(77.0 BLEU)。然而,在解析方向上,模型远落后于 ACE:要求它们为相同句子生成 MRS 时,它们在图的谓词和论元上达到 57.2(Sonnet)和 65.5(Opus)的 F$_1$,而 ACE 为 91.0,并且在大约 1% 的句子上与黄金标准完全匹配。我们刻画了解析任务的失败模式,并得出结论:仅凭生成分数并不能表明模型理解形式化语义表示。
英文摘要
The English Resource Grammar (ERG) is a hand-written computational grammar of English. Given a sentence, its processor, ACE, produces a formal meaning representation called Minimal Recursion Semantics (MRS): a graph of the sentence's predicates and their arguments. The grammar is bidirectional and can also turn an MRS back into an English sentence. \citet{hajdik2019} used the ERG's treebank to build a benchmark for that generation task, MRS to text, and trained sequence-to-sequence models to solve it. The parsing task, text to MRS, can be tested on the same sentences. We reconstruct their 10K-sentence test split, and score two large language models, Claude Sonnet~4.5 and Claude Opus~5, in both directions against their trained systems and against ACE, with no task-specific training. Given an MRS and three examples, Opus writes the sentence at 76.3 BLEU, ten points above their system trained on 72k pairs (66.1 BLEU), and comparable to their system trained on a million extra pairs (77.2 BLEU). Sonnet scores 65.7 BLEU, and letting it choose among ACE's own candidate sentences lifts it to 69.6, while a pooled judge that keeps Opus's own sentence among the candidates adds 0.6 points (77.0 BLEU). In the parsing direction, however, the models fall far behind ACE: asked for the MRS of the same sentences, they reach 57.2 (Sonnet) and 65.5 (Opus) F$_1$ on the graph's predicates and arguments against 91.0 for ACE, and exact-match the gold on about 1\% of sentences. We characterize the failure modes for the parsing tasks, and conclude that a generation score alone does not show that models understand formal semantic representations.