arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.12162cs.AIcs.CL

LLM 在草稿-验证-修订流水线中能否解决指示歧义?

Can LLMs in Draft-Verify-Revise Pipelines Resolve Deictic Ambiguity?

Obinna I. Ekekezie

首次发表
浏览论文内容

中文总结 AI 辅助

本研究探讨草稿-验证-修订流水线中LLM的指示歧义问题,通过合成数据集和e值检验,发现推理努力影响解析准确性,并建议明确指代对象以降低错误。

中文摘要 AI 辅助

草稿-验证-修订是一种常见的 LLM 编排模式,用于扩展推理时的计算量。一个 LLM 起草,第二个 LLM 对草稿进行批评并提供反馈,第三个 LLM 利用该反馈将草稿修订为最终输出。随着上下文在阶段之间级联,不同阶段的 LLM 可能对诸如“previous”之类的上下文相关表达式产生不同的解析。当这种情况发生时,该表达式会发生指示转移,即其指代对象的改变。这一现象通过一个包含 10 个基础示例的合成数据集进行了研究,每个示例在三种条件下呈现。在保持共享组件不变的情况下,这些条件变化了草稿阶段 LLM(助手)或验证阶段 LLM(评分者)是否正确解析了表达式,以及修订阶段 LLM(元评估者)需要多少独立推理来确定哪个解读是正确的。来自三个提供商的六个模型在 21 种推理努力配置下进行了测试,使用 e 值进行序贯检验,在一个主实验和一个消融实验中,消融实验从评分者的反馈中移除了错误分类标签。一个单独的 LLM 分析了元评估者对每个错误判决的陈述理由。平衡准确率(敏感性和特异性的未加权平均值)范围从 0.156(低于随机水平)到接近完美。GPT-5.2 从无推理时的 0.156 上升到其最高推理努力水平时的 0.942,而 Gemini 3 Pro 在每个水平上都保持在 0.94 以上。Gemini 3 Pro 在低推理努力下的得分超过了 GPT-5.2 在极高推理努力下的得分,而每次试验的成本约为后者的 5%。当元评估者出错时,它倾向于依赖表面线索而非操作性推理。实施草稿-验证-修订流水线的上下文工程师应警惕指示转移,并在每个阶段明确预期的指代对象。

英文摘要

Draft-verify-revise is a common LLM orchestration pattern for scaling inference-time compute. One LLM drafts, a second critiques the draft and provides feedback, and a third uses that feedback to revise the draft into the final output. As context cascades between stages, LLMs at different stages can resolve a context-dependent expression such as "previous" differently. When that happens, the expression undergoes a deictic shift, a change in what it refers to. This phenomenon was studied with a synthetic dataset of 10 base examples, each rendered in three conditions. Holding the shared components constant, the conditions varied whether the draft stage LLM (the assistant) or the verify stage LLM (the grader) resolved the expression correctly, and how much independent reasoning the revise stage LLM (the meta-evaluator) needed to determine which reading was correct. Six models from three providers were tested across 21 reasoning effort configurations using e-values for sequential testing, in a primary experiment and an ablation experiment that removed error classification labels from the grader's feedback. A separate LLM analyzed the meta-evaluator's stated rationale for each wrong verdict. Balanced accuracy (the unweighted mean of sensitivity and specificity) ranged from 0.156, below chance, to near-perfect. GPT-5.2 rose from 0.156 without reasoning to 0.942 at its highest reasoning effort level, while Gemini 3 Pro stayed above 0.94 at every level. Gemini 3 Pro at low reasoning effort outscored GPT-5.2 at xhigh reasoning effort for roughly 5% of the cost per trial. When the meta-evaluator erred, it tended to rely on surface cues rather than operational reasoning. Context engineers implementing draft-verify-revise pipelines should be wary of deictic shifts and make the intended referent explicit at each stage.

发表机构

  • Cambridge Health Alliance(剑桥健康联盟)
  • Harvard Medical School(哈佛医学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑