arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

并非双关:面向以人为中心的大语言模型评估的合理未知姓名(PUN)

No PUN Intended: Plausible Unknown Names for Person-Centred LLM Evaluation

Dimitri Staufer, David Hartmann, Ibrahim Baroud

arXiv 2608.21206首次发表:更新:

发表机构

Technische Universität Berlin; Weizenbaum Institute for the Networked Society; German Research Center for Artificial Intelligence (DFKI)(柏林工业大学; 魏茨曼网络社会研究所; 德国人工智能研究中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对LLM评估中人名使用导致的测量混淆问题,提出PUN协议构建合理未知姓名,经204人研究验证其有效性并发布300个对照姓名。

AI 中文摘要

人名被广泛用作大语言模型(LLM)评估事实性、隐私泄露、偏见及弃权(不执行)的提示变量,但当姓名的证据状态不受控制时,测量结果可能混淆记忆、检索、姓名先验及张冠李戴归因。我们将未知姓名定义为具备合理的名-姓形式、无索引全名证据且在记录的验证运行下无歧义信号的姓名,并引入PUN(合理未知姓名)这一用于构建和验证此类姓名的协议,该协议结合了源自Wikidata的组件、基于网络的LLM筛选及受控搜索再验证。我们报告了接受率、可复现性、消融实验及一项包含204名参与者的人类研究,发现被接受的姓名比对照组更具姓名特征,而参与者仅在3%的案例中恢复了人物证据。我们发布了300个带有对照的姓名。

英文摘要

Person names are widely used as prompt variables in LLM evaluations of factuality, privacy leakage, bias and abstention, but when a name's evidential status is uncontrolled, measurements may conflate memorisation, retrieval, name priors and wrong-person attribution. We operationalise an unknown name as one with plausible First-Last form, no indexed full-name evidence, and no ambiguity signals under a documented validation run, and introduce PUN (Plausible Unknown Names), a protocol for constructing and validating such names, combining Wikidata-derived components, web-enabled LLM screening, and controlled search revalidation. We report acceptance rate, reproducibility, ablations, and a 204-participant human study, finding accepted names are more name-like than controls while participants recover person evidence in only 3% of cases. We release 300 names with comparison controls.

CommentsUnder review

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑