arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大型语言模型拒绝行为研究中的随机性考量

Accounting for Stochasticity in Studies of Large Language Model Refusal

Emma Lurie, Stephanie T. Wang, Sorelle A. Friedler, Danaé Metaxa

arXiv 2609.33743首次发表:更新:

发表机构

University of Pennsylvania; Haverford College(宾夕法尼亚大学; 哈弗福德学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过纵向审计实验证明,单次查询不足以可靠评估大语言模型的拒绝行为,需15至25次重复查询才能准确量化,并揭示了决策边界区域的存在。

AI 中文摘要

我们提出了初步的经验证据,表明单次观测查询不足以评估大语言模型的拒绝行为。利用一个纵向审计系统,我们在四个日期对GPT-4.1就两个社会敏感话题、20个维基百科来源各发出了100次相同的提示。拒绝结果与稳定的伯努利过程一致,但20%的来源落在决策边界区域内,在该区域内单次查询在很大程度上无法提供信息。可靠地量化拒绝行为需要15至25次重复查询,远高于现有评估中常见的单次观测标准。

英文摘要

We present preliminary empirical evidence that single-observation queries are insufficient for evaluations of LLM refusal behaviors. Using a longitudinal auditing system, we issued identical prompts 100 times each across four dates to GPT-4.1 for two socially salient topics across 20 Wikipedia sources. Refusal outcomes were consistent with a stable Bernoulli process, yet 20\% of sources fell within a decision-boundary region where a single query is largely uninformative. Reliable quantification of refusals required between 15 and 25 repeated queries, well above the single-observation standard common in existing evaluations.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑