arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.24885cs.CLcs.CY

复制上限:面向精选语料库的基于本体生成的输入暴露控制

The Copy Ceiling: An Input-Exposure Control for Ontology-Grounded Generation over Curated Corpora

  • DreamLab AI

机构由 AI 辅助整理,请以论文原文为准。

John J. O'Hare

AI总结:

本文提出暴露核算方法,通过复制上限基线评估基于图检索的语料库问答,区分暴露与恢复,揭示模型增益多源于上下文复制而非推理,并验证其作为评估控制的有效性。

AI中文摘要:

当语言模型通过基于图的检索从精选语料库中回答问题时,较大的基础增强提升并不能证明其对检索到的结构进行了推理:上下文可能已经暴露了黄金答案。我们提出了暴露核算方法,该方法将每个黄金条目分类为所展示的上下文是否暴露了它以及答案是否恢复了它。其标量参考是复制上限,即上下文逐字复制所达到的召回率;相对于复制增益的符号增益衡量了模型相对于这一确定性、无评判者基线的召回率。在十个模型中,无辅助召回率平均为0.26,基于基础的召回率为0.92,但相对于复制的增益均为负值(-0.067至-0.022)。在11360个黄金条目观测中,代表在十种模型下评估的1136个目标实例,只有三个未暴露条目获得词汇信用。一项分层的模型评判审计对423个观测进行了,采用对称的引文验证策略,估计97.1%的获得信用的条目断言的所需关系;所有三个未暴露的信用均未通过关系裁决。在脚手架未暴露的目标上,词汇恢复率从无辅助的0.121降至基于基础的0.004;裁决验证了92个无辅助信用中的71个,而三个基于基础的信用均未通过验证,未建立全帧关系恢复率。在图标题词汇之外改写问题将暴露率从0.964降至0.328,而缺失触发的回退仅在506个问题中的2个上激活。一项配对生产研究将评判质量提高了+0.27(汇总),但阴性对照未能在格式良好的语料库块之外建立内容特异性。这些结果支持将暴露核算作为语料库派生评估的常设控制。该核算区分了暴露条目的遗漏和超出暴露的恢复;它并不确定是否发生了推理。

英文摘要:

We built a node that grounds a replaceable language model in a maintained ontology corpus, then asked what its successful-looking evaluation could support. Across ten models, grounding raised target-name recall from 0.265 unaided to about 0.92. A copy baseline, the recall a verbatim copy of the shown context already achieves, scores 0.964, and every model sits 0.022 to 0.067 below it. Copying therefore scores higher on this limited recall measure, which does not assess whether answers are better. The comparison tests what a recall score establishes; it does not test whether reasoning occurred, because a reasoned answer and a copy score alike when the answer name is already in context. We report exposure accounting (four counts classifying each gold item by whether the context exposed it and the answer recovered it) and a model-judged audit of 423 sampled item observations. A separate paired production study found a model-judged quality gain of +0.27 [+0.11, +0.45] on a 0-5 scale. Operational studies found failures that recall alone would not show: rephrasing questions out of the graph's vocabulary cut exposure from 0.964 to 0.328, yet the absence-keyed fallback would have fired on only 2 of 506; and inserting extracted facts degraded judged pages in every arm, so that step was disabled. Five-arm controls show that any well-formed on-corpus block beats no context but do not establish that the specific content matters, and no matched comparison against flat-text retrieval was run. The corpus is public and largely LLM-generated, which establishes neither training exposure nor novelty. Each study has its own outcome measure. Where gold derives from the injected corpus, we recommend reporting the accounting beside quality judgements, not in place of them.

补充信息

↑