发表机构
Aflo Technologies Inc.(Aflo科技公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究构建了可解释的特征+LLM混合L2口语评估模型,其评分超越81%的人类评分者,还证实停顿编码不影响LLM流利度评分,且通过多维度方法验证了模型的公平性与可靠性。
AI 中文摘要
第二语言(L2)英语学习者很少有机会与伙伴练习口语,而口语也是最易引发焦虑的语言技能,这些缺口推动了自动口语练习与评分市场的快速增长。但自动评分只有在准确、可解释、公平且以合适的人类基准为参照时才值得信赖。我们构建了一个用于自发L2对话的可解释特征+LLM混合模型,该模型从未拟合人类标签,而是基于ICNALE全球评分档案进行评估:该档案包含140篇口语,由约80名经过训练的评分者按10项分析标准评分。我们对130篇具备可用音频的L2口语进行评分,其中确定性De-Jong语音时序组合达到rho=0.764,与单个文本LLM的流利度判断融合后,针对共识金标准的斯皮尔曼rho达到0.818,优于80名个体训练评分者中的81%:高于中位评分者(rho=0.73)、接近最优评分者,且达到可靠性校正最大值的约83%(kappa_max=0.99)。该融合模型较单独的组合模型提升了0.054(配对自助法95%置信区间[0.017, 0.108],排除0);LLM提供了粗略的流利度排名,由连续组合模型进行细化。我们还报告了关于停顿编码的受控零结果,在该样本量下效应被限制在约±0.1 rho以内:固定LLM和学习者词汇,仅改变提示中停顿的书写方式,内嵌停顿位置未优于聚合停顿统计量(-0.069,置信区间[-0.15, +0.08]),且基于子句中间的接地标准未带来可靠增益。流利度信号来自实测的语音时序特征,而非LLM的停顿书写方式。我们通过两种一致的学习者隔离方法、配对自助法置信区间、独白阴性对照、经典测量的逐特征复现以及按母语(L1)划分的公平性审计,为所有主张提供支撑。
英文摘要
Second-language (L2) English learners can rarely rehearse speaking with a partner. Speaking is also the most anxiety-laden skill. These gaps drive a fast-growing market for automated speaking practice and scoring. But an automated score is trustworthy only if it is accurate, interpretable, fair, and benchmarked against the right human bar. We build an interpretable feature-plus-LLM hybrid for spontaneous L2 dialogue. We evaluate it without ever fitting to the human labels, against the ICNALE Global Rating Archive: 140 speeches rated by ~80 trained raters on 10 analytic criteria. We score the 130 L2 speeches with usable audio. A deterministic De-Jong speech-timing composite reaches rho=0.764. Blended with a single text-LLM fluency judgment, it reaches Spearman rho=0.818 against the consensus gold. This agrees with the consensus better than 81% of the 80 individual trained raters: above the median rater (rho=0.73) and near the best, and at ~83% of the reliability-corrected maximum (kappa_max=0.99). The blend improves on the composite alone by +0.054 (paired-bootstrap 95% CI [0.017, 0.108], excludes 0); the LLM adds a coarse fluency ranking that the continuous composite refines. We also report a controlled null on pause encoding, bounded to effects below about +/-0.1 rho at this sample size. Holding the LLM and learner words fixed and varying only how pauses are written into the prompt, inline pause locations do not beat aggregate pause statistics (-0.069, CI [-0.15, +0.08]), and a grounded mid-clause criterion gives no reliable gain. The fluency signal comes from the measured speech-timing features, not from how pauses are written for the LLM. We back every claim with two agreeing learner-isolation methods, paired-bootstrap CIs, a monologue negative control, per-feature reproduction of classical measurements, and a per-L1 fairness audit.
Comments17 pages, 3 figures, 5 tables