发表机构
Estonian Entrepreneurship University of Applied Sciences (EUAS)(爱沙尼亚应用科学创业大学(EUAS))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出骰子掷点法作为LLM品牌推荐重复查询审计的标准化协议,通过分解响应方差等方法,得出三类迭代指导,经外部验证可有效支撑相关审计工作。
AI 中文摘要
背景:研究人员日益使用重复相同提示来审计大型语言模型(LLM)品牌推荐中的随机变异,但目前尚无标准化协议用于设置迭代次数、选择稳定性指标或建立可靠性阈值。目的:我们将骰子掷点法形式化为一种可重复使用的协议,用于LLM品牌推荐的重复查询审计,其基于温度缩放的核采样生成模型。方法:总响应方差被分解为采样、提示措辞、运行间以及模型版本成分。该方法的核心包括:将迭代次数作为重复测量的负二项混合模型、作为分布自由效应量的Cliff's delta、保留依赖性的自助法、基于模拟的功效、概化理论分解、以及固定快照的漂移诊断。我们重新分析了五项品牌推荐审计研究,涉及约190000个观测值、270多个品牌、6种语言,迭代次数为5至40次。结果:从D研究中得出三类迭代指导:探索性(n=5,G=0.58)、验证性(n=10,G=0.74)和严格性(n=15,G=0.81),与效应量和概化目标相关。四类指标族(计数、集合、嵌入、公平性调整的PASOR)具有互补性,支持采用紧凑指标组而非单一指标。对三个独立语料库(Motoki等,100轮;Rozado,24个模型;llm-stability)的预注册外部验证,在39个单元格中重现了D研究的可靠性预测,无失败,且n=5的功效值精确到两位小数;固定迭代层级无法迁移,支持先试点再解决的思路。结论:该协议为LLM品牌推荐的重复查询审计提供了基于统计原理的基础,适用于真实自回归生成的条件非高斯结构。
英文摘要
Background: Researchers increasingly use repeated identical prompts to audit stochastic variation in large language model (LLM) brand recommendations, yet no standardized protocol exists for setting iteration counts, selecting stability metrics, or establishing reliability thresholds. Objective: We formalize the Dice Roll Method as a reusable protocol for repeated-query auditing of LLM brand recommendations, grounded in a generative model of temperature-scaled nucleus sampling. Methods: Total response variance is decomposed into sampling, prompt-phrasing, run-to-run, and model-version components. The stack: a negative-binomial mixed model with iterations as repeated measures; Cliff's delta as the distribution-free effect size; dependence-preserving bootstrap; simulation-based power; a generalizability-theory decomposition; drift diagnostics on pinned snapshots. We reanalyse five brand-recommendation auditing studies: approximately 190,000 observations, 270+ brands, 6 languages, iteration counts 5 to 40. Results: Three tiers of iteration guidance emerge from the D-study: exploratory (n = 5, G = 0.58), confirmatory (n = 10, G = 0.74), and rigorous (n = 15, G = 0.81), tied to effect-size and generalizability targets. The four metric families (count, set, embedding, fairness-adjusted PASOR) are complementary, motivating a compact metric battery over single indicators. A pre-registered external validation on three independent corpora (Motoki et al., 100-round; Rozado, 24 models; llm-stability) reproduces the D-study reliability prediction in 37 of 39 cells with no failures and the n = 5 power value to two decimals; the fixed tiers do not transfer, supporting a pilot-then-solve reading. Conclusion: The protocol gives repeated-query auditing of LLM brand recommendations a statistically principled footing under the conditional, non-Gaussian structure of real autoregressive generation.
Comments30 pages, 2 figures, 19 tables. Substantially revised; supersedes the Research Square preprint 10.21203/rs.3.rs-8883056/v1. Includes a pre-registered external validation on three independent corpora (Motoki et al., Rozado, llm-stability)