AI 中文总结
研究针对自动作文评分中跨响应条件非均匀可靠性评估问题,引入含三个部分的条件可推广性框架,通过比较分析与经验扫描、依熵分层等,以限时L2写作自动评分验证,提供了评估非均匀可靠性的可移植工作流程。
AI 中文摘要
聚合可靠性估计可能掩盖响应条件下测量设计负担的异质性,单一的G或D研究可能错误表征特定层次设计的充分性。本研究引入了一个具有三个组成部分的条件可推广性框架。首先,将自动评分配置视为可接受的测量条件全域。其次,比较分析D研究投影与经验配置扫描。第三,证据以熵定义的响应层次为条件。通过限时L2写作的自动作文评分证明,总体设计可靠,按熵层次重新估计,可靠性保持较高但适度下降,该框架提供了评估非均匀可靠性的可移植工作流程。
英文摘要
Aggregate reliability estimates can obscure heterogeneity in measurement-design burden across response conditions, so a single G- or D-study may mischaracterize a design's adequacy for particular strata. This study introduces a conditional generalizability framework with three components. First, automated scoring configurations -- the encoder architectures and scoring-head families admissible within a fixed pipeline -- are treated as a universe of admissible measurement conditions rather than incidental modeling choices. Second, analytical D-study projections are compared with empirical configuration sweeps over a finite scoring pool, yielding two estimands of design adequacy whose agreement or divergence diagnoses the realized configuration universe. Third, evidence is conditioned on entropy-defined response strata, treating entropy as an operational stratification variable, not a construct claim about writing quality. Whereas recent generalizability-theory extensions address AI-generated item variants on the response side, this framework addresses the analogous scoring-side problem: AI-mediated scoring configurations. Demonstrated with automated essay scoring of timed L2 writing, the realized design was dependable in aggregate (Phi approx 0.76). Re-estimated within entropy strata, dependability stayed high but declined modestly and robustly (Phi = 0.88, 0.87, 0.84) -- a gradient implying different decision-study requirements, the highest-entropy stratum requiring the most crossed conditions. The framework offers a portable workflow for evaluating nonuniform dependability.