发表机构
University of Pennsylvania; Haverford College(宾夕法尼亚大学; 哈弗福德学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过纵向审计实验证明,单次查询不足以可靠评估大语言模型的拒绝行为,需15至25次重复查询才能准确量化,并揭示了决策边界区域的存在。
AI 中文摘要
我们提出了初步的经验证据,表明单次观测查询不足以评估大语言模型的拒绝行为。利用一个纵向审计系统,我们在四个日期对GPT-4.1就两个社会敏感话题、20个维基百科来源各发出了100次相同的提示。拒绝结果与稳定的伯努利过程一致,但20%的来源落在决策边界区域内,在该区域内单次查询在很大程度上无法提供信息。可靠地量化拒绝行为需要15至25次重复查询,远高于现有评估中常见的单次观测标准。
英文摘要
We present preliminary empirical evidence that single-observation queries are insufficient for evaluations of LLM refusal behaviors. Using a longitudinal auditing system, we issued identical prompts 100 times each across four dates to GPT-4.1 for two socially salient topics across 20 Wikipedia sources. Refusal outcomes were consistent with a stable Bernoulli process, yet 20\% of sources fell within a decision-boundary region where a single query is largely uninformative. Reliable quantification of refusals required between 15 and 25 repeated queries, well above the single-observation standard common in existing evaluations.