幻影填充:当表单需要答案时,语言模型会编造一个
PhantomFill: When the Form Demands an Answer, Language Models Invent One
浏览论文内容
中文总结 AI 辅助
研究语言模型在表单填充时的表现,发现表单格式会致其产生幻觉,通过实验展示不同模型在不同格式要求下的编造情况,发布幻影填充基准,测试一行模式代码修复方法,揭示模型在格式压力下诚实性受影响的问题。
中文摘要 AI 辅助
生产中的语言模型并非撰写散文,而是填充表单,如JSON字段、函数参数、提取模板等。我们发现表单本身会导致幻觉。对13个模型就相同输入问相同问题,仅改变答案格式。输入内容设计成无法回答的问题,如一个有12400个赞但无可见回复的热门帖子、一个通话未被转录的支持工单。在自由文本中,GPT - 5.5大多诚实地回答无回复数据。但给定情感的必填JSON字段时,它40次中有40次编造答案。13个模型中有10个因必填字段导致100%编造。明确的“证据不足”选项仅对前沿模型有效,部分模型会忽视直接指令。诚实受格式压力影响,我们发布了幻影填充基准及相关指标,测试的修复方法仅一行模式代码,而失败情况随处可见。
英文摘要
Language models in production do not write prose. They fill forms: JSON fields, function arguments, extraction templates. We show that the form itself causes hallucination. We ask thirteen models the same question about the same input and change only the answer format. The inputs are built so the question cannot be answered: a viral post showing 12,400 likes but no visible replies, a support ticket whose call was never transcribed. In free text, GPT-5.5 says there is no reply data 98% of the time. Given a required JSON field for sentiment, the same model invents an answer 40 times out of 40. It fabricates the mood of crowds it never saw and quotes customers it never heard. Required fields drive fabrication to 100% in ten of thirteen models. An explicit "insufficient evidence" option rescues only the frontier: all nine open-weight models ignore it. Under grammar-constrained decoding, where the escape token is guaranteed reachable by the sampler, five open models spend it zero times out of 203 trials on the three fields that carry the fabrication, and twelve times on the one field where escaping concedes nothing. They can emit the word. They decline to spend it where it costs them an answer. A direct instruction, do not infer sentiment, is overridden by the schema in four of six models. Resistance does not come with scale: within a single model family, the smallest model refuses, the mid-sized model fabricates, the largest refuses again. Honesty under format pressure is a training outcome that no one is measuring. Fabrication hides where hedging is impossible: in required enums and minimum-count arrays, fields where no disclaimer fits. We release PhantomFill, a benchmark with deterministic scoring and two reportable numbers: the Coerced Fabrication Rate and the Escape Utilization Rate. The fix we test is one line of schema. The failure we measure is everywhere.