发表机构
Harvard T.H. Chan School of Public Health(哈佛大学陈曾熙公共卫生学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过最小集设计测试LLM回答州政策问题时,错误答案是否源于跨辖区替代,发现归因脆弱,需完整同州参考集验证,并发布协议与数据。
AI 中文摘要
当大型语言模型(LLM)回答一个特定州的政策问题时,它可能是在产生幻觉,也可能是在返回另一个州真实存在的数值。我们采用最小集设计进行测试:问题措辞固定,仅改变辖区,涵盖美国50个州和哥伦比亚特区(共51个辖区)以及三个精确定义的医疗补助收入资格数量。黄金值来自官方数据手册,并在101个已核对的单元格中有102个与独立来源一致。在预注册协议下,Claude Sonnet 5.5和GPT-5.6 Sol在153个项目中分别有10个和25个项目可重复地给出另一个州的当前值,且两次独立重复结果相同。然而,归因是脆弱的。将任何等于另一个州值的错误答案归为跨辖区替代,比检查所问州自身记录中的每个数字能产生3-5倍更多的可重复替代,因为许多表面上的跨州答案实际上是所问州自身在不同约定或更早年份下的值。关于跨辖区错误的断言需要完整的同州参考集。我们将发布协议、黄金表和所有模型输出。
英文摘要
When an LLM answers a state-specific policy question wrongly, it may be hallucinating, or it may be returning a real value that holds in another state. We test this with a minimal-set design: the question wording is fixed and only the jurisdiction varies, across the 50 U.S. states and the District of Columbia (51 jurisdictions) and three exactly defined Medicaid income-eligibility quantities. Gold values come from an official data book and agree with an independent source in 101 of 102 checked cells. Under a pre-registered protocol, Claude Sonnet 5.5 and GPT-5.6 Sol reproducibly give another state's current value, identical across two independent repeats, for 10 and 25 of 153 items. Attribution is fragile, however. Crediting any wrong answer that equals another state's value yields 3-5x more reproducible substitutions than checking every number in the asked state's own records, because many apparent cross-state answers are the asked state's own values under another convention or from an earlier year. Claims about cross-jurisdiction error need a complete same-state reference set. We will release the protocol, gold table, and all model outputs.
Comments6 pages, 3 figures, 1 table