发表机构
Dankook University(韩国檀国大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究大语言模型中问题顺序效应,将QQ等式发展为审计标准,理论上表征满足该等式的机制类别,方法上开发审计管道,实验发现强制二元下一个token对数概率不适用于分布级QQ审计,建议进行饱和诊断。
AI 中文摘要
人类调查受访者表现出满足QQ(量子问题)等式的问题顺序效应,这是射影量子问题顺序模型的先验、无参数预测。我们将QQ等式发展为自回归大语言模型顺序判断的审计标准。理论上,我们表征了哪些机制类别能稳健满足它:边际独立核满足QQ当且仅当所有四个不匹配转换率一致;极性和位置相关的重复族由精确的交叉对称条件和违反的封闭形式表征;满足QQ的行为在顺序匹配混合下是封闭的;秩2的默认语境标准转化为审计坐标为$|\qQQ|\le\OSS$。方法上,我们开发了适用于任何暴露下一个token对数概率的模型的预指定、记录审计管道。实验上,在两个框架下对一个开放权重指令调整模型的首次信号试点中,所有预指定健康门通过,但分别有17/18和7/8的项目对饱和,没有项目被认证为残留语境。因此,在测试模型和提示条件下,强制二元下一个token对数概率不足以进行分布级QQ审计;我们建议当下一个token分布被视为调查响应分布时进行预指定饱和诊断。
英文摘要
Question-order effects in human survey data have been reported to approximately satisfy the QQ (quantum question) equality, a parameter-free prediction of the standard projective quantum question-order model. We develop this equality into an audit framework for sequential binary judgments of autoregressive large language models (LLMs). Theoretically, we characterize mechanism families that satisfy QQ robustly, show that classical repetition can reproduce the equality exactly, and combine QQ with the rank-2 Contextuality-by-Default criterion through $|q_{QQ}| \le \mathrm{OSS}$. This separates order sensitivity, QQ imbalance, and residual contextuality rather than treating them as interchangeable signatures. Methodologically, we introduce a committed multi-turn forced-branch protocol that reconstructs order-conditioned joint distributions from next-token log-probabilities under counterbalanced label mappings and pre-specified health gates. A first-signal pilot on an open-weight instruction-tuned model reveals the central measurement problem. Although all pre-specified health gates passed, the binary-conditioned distributions were near-deterministic for 17 of 18 item pairs under the direct-evaluation framing and 7 of 8 under the persona framing. Label assignment materially changed several mapping-specific QQ verdicts, and no item was certified as residually contextual. Thus, under the tested conditions, the observed QQ outcomes did not uniquely identify a response mechanism in the presence of a saturated and label-sensitive measurement interface. The main implication is methodological: next-token probabilities should not be interpreted as survey-response distributions without first establishing adequate dispersion. We therefore argue that saturation screening and label counterbalancing should precede structural interpretation in distribution-level audits of LLM judgments.
Commentsv2: major revision. Five restructured findings separating order sensitivity, QQ imbalance, and residual contextuality; two-layer discriminant table; pipeline figure and per-item joint table; envelope pooling and certified Gamma bounds specified; retrospective batch-1 G3 re-validation (all verdicts preserved). No new pilot measurements