发表机构
Serenze Global(塞伦兹全球)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文对照人生真实语料库审计LLM生成的自传,发现96.7%的场景验证失败,提出了可重复使用的审计工具及基于真实信息的补救方法。
AI 中文摘要
当大语言模型(LLM)被要求撰写一个人的生平,其生成内容中有多少是真实发生的?我们开展了一项场景层面的案例研究审计——据我们所知,这是首次对照主题特定的真实语料库对LLM生成自传进行的量化审计,该研究基于非系统性文献检索。研究对象与本文作者为同一人:我们使用对话式LLM撰写了一本共366天的“每日一页”第一人称轶事条目书,其记录的输入内容为一个模板、两个示例日的内容以及每日的引语——而非作者本人的语料库——之后我们采用分析前确定的四级标准,对照独立验证语料库对每一天的内容在轶事场景层面进行审计。我们将验证失败率定义为未被评为“已验证”(场景得到正面证实)的天数占比:366天中有354天失败,占比96.7%(Wilson 95%置信区间为94.4-98.1%)。仅有12天包含得到证实的场景;19天(占比5.2%)提出的主张与记录直接矛盾;主要的失败模式是基于现实的漂移——虚构场景中包含真实人物、雇主和场景——尽管其测量占比因评分者而异。独立重新评分复制了核心结果(无证据表明原始比率被夸大),同时显示四级分类仅具有一般至中等的可靠性。使用当前命名模型用相同输入重新生成相同天数的内容,在相同输入下重现了100%的验证失败;基于主题语料库进行生成显著提高了验证率,但仍存在大量剩余失败(83.3%)。我们贡献了该测量方法、一种可重复使用的审计工具(我们证明其“弱/未验证”边界不可靠)以及具有量化效果的基于真实信息的补救方法。
英文摘要
When a large language model (LLM) is asked to write a person's life, how much of what it writes actually happened? We present a scene-level case-study audit - the first quantified audit of LLM-generated autobiography against a subject-specific ground-truth corpus that we are aware of, based on an unsystematic literature search. The subject and the author of this paper are the same person: a 366-day "page-a-day" book of first-person anecdotal entries was drafted with a conversational LLM whose documented inputs were a template, two exemplar days, and each day's quote - not her corpus - and every day was subsequently audited at the anecdote-scene level against an independent verification corpus using a four-level rubric fixed before analysis. We define the verification-failure rate as the share of days not rated VERIFIED (scene positively corroborated): 354 of 366 days fail, 96.7% (Wilson 95% CI 94.4-98.1%). Only 12 days contain a corroborated scene; 19 days (5.2%) assert claims actively contradicted by the record; the dominant failure mode is grounded drift - real people, employers, and settings inside invented scenes - though its measured share varies across raters. Independent re-rating replicates the headline (no evidence the original rate was inflated) while showing that the four-way taxonomy has only fair-to-moderate reliability. Regenerating the same days with current named models reproduces 100% verification failure under the same inputs; grounding generation in the subject's corpus significantly improves the verification rate while leaving substantial residual failure (83.3%). We contribute the measurement, a reusable audit instrument whose WEAK/UNVERIFIED boundary we show to be unreliable, and a grounding remedy with quantified effect.
Comments20 pages, 4 figures, 6 tables. Code and derived data available at https://github.com/heathriel/synthetic-memoir-audit