arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.10503cs.CL

每个词元都重要:用于测量大语言模型态度与偏差的精确李克特量表分布

Every Token Counts: Exact Likert-Scale Distributions for Measuring LLM Attitudes and Biases

Davood Wadi, Mohsen Ghodrat, Matthew Philp

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出结合人类心理测量学与LLMs机制的精确框架,通过因子实验、词元级PMFs及对应分析方法,分离LLMs的系统性原产国偏差。

中文摘要 AI 辅助

随着大语言模型(LLMs)越来越多地被部署为自主智能体,准确评估其潜在价值观与偏差至关重要。自然语言处理领域通常使用大型非结构化基准来评估模型,这些数据集虽对通用能力有效,但从根本上混淆了因果机制:即使检测到总体偏差,非结构化评估也无法区分其源于基线特性、情境混淆因素还是复杂交互。为解决此问题,我们提出一种用于LLMs受控行为评估的分析精确框架,将人类心理测量学与LLMs机制相结合,弥补设计、测量与分析中的缺口:首先,用完全交叉的因子实验替代非结构化提示,以系统分离因果主效应与交互效应;其次,通过直接操作精确的词元级概率质量函数(PMFs)消除蒙特卡洛文本采样噪声;最后,推导多变量序数共识度量与分布方差分析来分析这些PMFs。我们以五个LLMs的消费者民族中心主义案例研究验证该框架,证明其可分离出总体基准会掩盖的系统性原产国偏差。

英文摘要

As Large Language Models (LLMs) are increasingly deployed as autonomous agents, accurately evaluating their latent values and biases is critical. The NLP community typically evaluates models using large, unstructured benchmarks. While effective for general capabilities, these datasets fundamentally conflate causal mechanisms: even when an aggregate bias is detected, unstructured evaluations cannot disentangle whether it stems from baseline traits, contextual confounders, or complex interactions. To address this, we introduce an analytically exact framework for the controlled behavioral evaluation of LLMs. We bridge human psychometrics with LLM mechanics by resolving gaps in design, measurement, and analysis. First, we replace unstructured prompting with fully crossed factorial experiments to systematically isolate causal main and interaction effects. Second, we eliminate Monte Carlo text sampling noise by operating directly on exact, token-level Probability Mass Functions (PMFs). Third, we derive a multivariate ordinal consensus metric and a distributional ANOVA to process these PMFs analytically. We validate our framework with a case study on consumer ethnocentrism across five LLMs, demonstrating how our approach isolates systemic country-of-origin biases that aggregate benchmarks otherwise obscure.

发表机构

  • McGill University(麦吉尔大学)
  • University Canada West(加拿大西部大学)
  • Toronto Metropolitan University(多伦多都会大学)

机构由 AI 辅助整理,请以论文原文为准。

↑