AI 中文总结
该研究探讨了作为评判者的决策模型中,结构化评判请求的呈现方式会影响其对候选答案的错误接受率,发现添加一个冒号的编辑会显著提升特定模型Jev的错误接受率,而GPT-6 Sol未出现此类问题。
AI 中文摘要
被指示对最终承诺进行评分的答案评判者,即使早期值与参考值匹配,也应拒绝明显错误的最终值。我们表明,根据结构化评判请求的呈现方式,给候选答案添加一个冒号就可能违反这一要求。数值参考可证实错误,配对干预可区分候选编辑与整合的呈现选择。在200个此前未使用过的DROP和GSM8K源集群上,在排序后的JSON键下,该编辑使Jev的错误接受率从1.0%升至26.0%(采用三个输出标签时),从3.0%升至26.5%(采用已发布的四标签评分指令时)。在插入式呈现下,两种候选变体均被拒绝。尽管Jev在两种评分配置和呈现方式下均达到控制阈值,但这些交互作用通过了预先指定的统计校正。大部分额外接受发生在被分配了更大数值误差的候选中。GPT-6 Sol未观察到线索条件下的错误接受,缺失响应未解决。结果表明,基本评判能力可与逻辑等价请求呈现下对固定候选编辑的显著不同脆弱性共存。参考感知语法和复合排序变化限制了该发现的操作范围,且未测量其内部原因。
英文摘要
An answer judge instructed to grade the final commitment should reject an explicitly wrong final value even when an earlier value matches the reference. We show that adding one colon to a candidate can violate this requirement depending on the presentation of the structured judging request. Numeric references certify the error, and paired interventions distinguish the candidate edit from the integration's presentation choices. On 200 previously unused DROP and GSM8K source clusters, the edit increased Jev's false acceptance from 1.0% to 26.0% with three output labels and from 3.0% to 26.5% with the published four-label grading instruction under sorted JSON keys. Both candidate variants were rejected under insertion presentation. These interactions passed the prespecified statistical correction even though Jev met the control thresholds under both grading configurations and presentations. Most excess acceptances occurred among candidates assigned larger numerical errors. GPT-6 Sol produced no observed cue-condition false acceptances, with missing responses unresolved. The result shows that basic judging competence can coexist with sharply different vulnerability to a fixed candidate edit across logically equivalent request presentations. The reference-aware grammar and compound ordering change limit the finding's operational scope and leave its internal cause unmeasured.