arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

噪声从何而来?大语言模型品牌答案中非确定性的方差成分分解

Where Does the Noise Come From? A Variance-Components Decomposition of Non-Determinism in LLM Brand Answers

Dmitrij Żatuchin

arXiv 2607.13304首次发表:更新:

AI 中文总结

研究大语言模型品牌答案中非确定性的来源,通过交叉随机效应分解将响应级品牌结果总方差分为提示内重采样等四个来源,应用于多语言多模型语料库,发现查询语言是最大方差来源,跨语言和模型能提升可靠性而非重复提示。

AI 中文摘要

测量大语言模型(LLMs)是否推荐某个品牌的团队面临可重复性问题:同一个问题问两次,答案会变动。实践中通常对每个提示进行几次(一般五次)重采样并求平均值,将提示内重采样视为噪声来源。但一个测得的品牌分数变动至少有四个可分离原因:提示内重采样、提示释义、模型身份和查询语言。我们指定了一种交叉随机效应(泛化理论)分解,将响应级品牌结果的总方差划分为这四个来源,并将这些成分嵌入到一个决策研究分配中,该分配返回为达到目标可靠性需要购买多少重复项、释义、模型和语言。我们将其应用于一个完全交叉的语料库,该语料库包含12933个关于20个中东欧品牌、8种语言和3个模型(参数模式下的GPT - 5.2和Gemini 3 Flash、基于地面检索的Perplexity)的LLM响应,有一个由1435个单元格组成的稳定性子集,每个单元格大约重采样五次。结果是每个响应的多语言情感极性。查询语言是最大的系统方面(一个响应方差的26.5%),而品牌身份仅为1.5%(组内相关系数ICC为0.0146),所以单个人工智能答案几乎没有品牌区分信号。一旦单元格项分离出纯重采样,重采样占方差的34.8%,品牌与上下文交互占29.6%;品牌与语言占8.6%(双语惩罚),而品牌与模型以及品牌与提示接近零。每单位查询预算,添加语言和模型比添加重复项能更大程度地降低相对误差方差:超过第五次重复时,重复项只能将其降低0.0003。品牌排名可靠性仍然很低,单个答案接近0.01,完全交叉设计时约为0.36,所以可靠性是通过跨语言和模型分布来实现的,而不是通过重复一个提示。

英文摘要

Teams measuring whether large language models (LLMs) recommend a brand face a reproducibility problem: ask the same question twice and the answer moves. Practice resamples each prompt a few times (commonly five) and averages, treating within-prompt resampling as the source of the noise. But a measured brand score moves for at least four separable reasons: within-prompt resampling, prompt paraphrase, model identity, and query language. We specify a crossed random-effects (generalizability-theory) decomposition that partitions the total variance of a response-level brand outcome into these four sources, and embed the components in a decision-study allocation that returns how many repeats, paraphrases, models, and languages to buy for a target reliability. We apply it to a fully crossed corpus of 12,933 LLM responses on 20 Central and Eastern European brands, 8 languages, and 3 models (GPT-5.2 and Gemini 3 Flash in parametric mode, Perplexity in grounded retrieval), with a stability subset of 1,435 cells resampled about five times. The outcome is per-response multilingual sentiment polarity. Query language is the largest systematic facet (26.5% of the variance of one response) against 1.5% for brand identity (ICC 0.0146), so a single AI answer carries almost no brand-discriminating signal. Once a cell term isolates pure resampling, resampling is 34.8% of variance and the brand-in-context interaction 29.6%; brand-by-language is 8.6% (a bilingual penalty) while brand-by-model and brand-by-prompt are near zero. Per unit of query budget, adding languages and models reduces relative-error variance far more than adding repeats: a repeat past the fifth reduces it by only 0.0003. Brand-ranking reliability stays low, near 0.01 for a single answer and about 0.36 at the full crossed design, so reliability is bought by spreading across languages and models, not by repeating one prompt.

Comments18 pages, 6 tables, 3 code listings

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑