AI 中文总结
提出 Rater Ising-Potts 模型,利用 LLM 嵌入权重和两两协议指标评估多类别评分可靠性,无需预设阈值,在 AERA 数据集上验证了与人类评分的高度一致性。
AI 中文摘要
Ising 模型被扩展为适用于多项数据的 Potts 模型。我们引入了一种 Rater Ising-Potts 模型,该模型利用评分者两两之间的协议指标和类别标签,权重由 LLM 嵌入派生。该模型不预设有序类别阈值或等距评分,而是直接关注评分者之间的两两一致性,并分配类别特定的正权重,使其特别适用于评分者依据评分指南评估回答时的多类别评分可靠性。我们在多种构答任务上展示了该模型的有效性,包括 AERA 数据集中平衡的简答题项和更具挑战性的不平衡论文提示。在这些设置中,模型与人类评分达到高度一致,绝大多数误分类发生在相邻评分等级之间,证实了其在不施加刚性假设的情况下保留评分量规序数结构的能力。我们引入了一种实用的相似性归一化和可选的幂变换,作为可调预处理步骤,以增强语义区分度,并可适应不同数据集。这些发现表明,LLM 派生的语义相似性,结合这种简约的 Potts 型公式和灵活的相似性缩放,为教育评估背景下的可靠性审计提供了一个稳健且可解释的框架。还讨论了扩展到多个评分者和分层评分过程的情况。
英文摘要
The Potts model extends the Ising model to multinomial data. We introduce a Rater Ising-Potts model that uses agreement indicators between pairs of ratings and category labels, with weights derived from LLM embeddings. The model does not presuppose ordered category thresholds or equidistant scoring; instead, it focuses on pairwise agreement among ratings and assigns category-specific positive weights, making it suited for multi-category scoring reliability. We evaluate the model on three constructed-response datasets spanning a corpus of K=14,466 short answers on a three-level rubric and two AERA essay prompts of roughly 1,200-1,400 responses on four-point rubrics. We compare three strategies for sharpening the similarity signal: top-K pruning, min-max normalization with a power transformation, and ColBERT late-interaction similarities. Top-K pruning, which replaces the dense similarity graph with a sparse local network of strongest semantic neighbors, consistently yields the highest accuracy and Cohen's kappa, and the selected neighborhoods are always a small fraction of the corpus. Power tuning consistently ranks second, while ColBERT is competitive on longer essay prompts and adds little on short answers. Across all settings, most misclassifications occur between adjacent score levels, confirming that the model preserves the ordinal structure of scoring rubrics without imposing rigid assumptions. These findings suggest that LLM-derived similarities, combined with a parsimonious Potts formulation and a sparse local graph, offer a robust and interpretable framework for reliability auditing in educational assessment. We discuss extensions to multiple raters and hierarchical rating designs.