arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.06300cs.AI

基于概念激活向量的L2口语评估系统的偏差分析

Bias Analysis of L2 Speaking Assessment Systems Using Concept Activation Vectors

Arya Labroo, Mengjie Qian, Kate Knill

首次发表
浏览论文内容

中文总结 AI 辅助

该研究将概念激活向量(CAV)分析扩展到BERT和Whisper两个L2口语评估系统,发现概念可恢复性及对概念的敏感性依赖于模型架构,稀疏自编码器(SAEs)可提升概念线性恢复性但会减弱激活空间敏感性。

中文摘要 AI 辅助

自动口语评估系统正越来越多地被部署在高风险场景中,用于给第二语言(L2)学习者的口语测试打分,因此至关重要的是要证明它们的分数取决于口语能力,而非第一语言(L1)或年龄等不相关的说话者属性。基于Transformer的基础模型提高了这些L2口语评分器的准确性,但它们的黑箱表示使公平性和可解释性分析变得更加困难。在之前使用概念激活向量(CAVs)检测基于特征的评分器中对不需要的属性(“概念”)的偏差的研究基础上,我们将基于CAV的分析扩展到两个神经口语评估系统:基于文本的BERT评分器和基于Whisper的语音-文本多模态评分器。CAVs将人类可解释的概念表示为模型激活空间中的方向,使我们能够区分概念是否被编码在模型的内部表示中,以及概念是否影响预测分数,后者使用基于梯度的敏感性指标进行量化。由于CAV依赖于线性可分性,而这在复杂的神经嵌入空间中不太可能出现,我们还研究了稀疏自编码器(SAEs)是否通过在稀疏潜在空间中学习CAV并将其映射回激活空间,从而提供更清晰的概念方向。我们的分析表明,概念可恢复性强烈取决于被探测的表示和架构,而非仅取决于概念本身。对概念的敏感性也具有架构依赖性。SAEs使概念更易线性恢复,但会减弱原始激活空间的敏感性,尤其是在低维层中。这些发现强调,在审核口语评估系统的偏差时,需要区分概念可恢复性与概念影响。

英文摘要

Automatic speaking assessment systems are increasingly deployed in high-stakes settings to mark second language (L2) learners' speaking tests, making it critical to show that their scores depend on speaking proficiency rather than irrelevant speaker attributes such as first language (L1) or age. Transformer-based foundation models have improved the accuracy of these L2 speaking graders, but their black-box representations make fairness and interpretability analysis more difficult. Building on prior work that used Concept Activation Vectors (CAVs) to detect bias towards unwanted attributes (`concepts') in feature-based graders, we extend CAV-based analysis to two neural speaking assessment systems: a text-based BERT grader and a speech-and-text multimodal grader based on Whisper. CAVs represent human-interpretable concepts as directions in a model's activation space, allowing us to distinguish between whether a concept is encoded in a model's internal representations and whether it influences the predicted score, the latter quantified using a gradient-based sensitivity metric. Since CAVs rely on linear separability, which is less likely in complex neural embedding spaces, we also investigate whether sparse autoencoders (SAEs) provide cleaner concept directions by learning CAVs in a sparse latent space and mapping them back to activation space. Our analysis shows that concept recoverability depends strongly on the representation and architecture being probed, rather than on the concept alone. Sensitivity to concepts is also architecture-dependent. SAEs make concepts more linearly recoverable, but attenuate the original activation-space sensitivity, especially in low-dimensional layers. These findings highlight the need to distinguish concept recoverability from concept influence when auditing bias in speaking assessment systems.

↑