arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

当目标领域发生变化时:高风险英语评估中人工智能介导的构念漂移

When the Target Domain Changes: AI-Mediated Construct Drift in High-Stakes English Language AssessmenW

Yi Gui

arXiv 2607.11213首次发表:更新:

发表机构

Measurement Incorporated(测量公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究高风险英语测试中人工智能介导的构念漂移问题,提出有界人工智能介导这一有效性设计原则,指出分数解释在用于支持人工智能介导的学术交流主张时应调整补充。

AI 中文摘要

高风险英语水平测试将标准化的、无辅助的表现作为学术英语水平分数解释的证据。这种解释仍然有意义,但随着目标语言使用领域越来越多地涉及生成式人工智能,从无辅助测试表现推断学术交流准备情况变得不那么明显。这个概念有效性论证将人工智能重新构建为高风险语言测试中的分数解释问题,而不仅仅是评分、反馈、安全或不当行为的操作问题。本文通过三个不均衡的层次综合当前文献表明,大多数工作将人工智能视为评估基础设施,而很少从理论上探讨其对构念有效性和推断依据的影响。它将人工智能介导的构念漂移定义为当目标领域所需的交际能力通过人工智能介导发生变化而测试构念仍锚定在无辅助表现模型时出现的不一致。它提出有界人工智能介导作为一种以有效性为导向的设计原则:一种标准化条件,即所有考生都可以使用相同的由机构控制的人工智能助手,具有预定义的辅助边界、记录的交互以及区分理解支持和答案生成的任务。本文认为,当用于支持关于人工智能介导的学术交流的主张时,分数解释应该缩小范围并加以补充。

英文摘要

High-stakes English proficiency tests treat standardized, unaided performance as evidence for score interpretations about academic English proficiency. This interpretation remains meaningful, but as target language use domains increasingly involve generative AI, the extrapolation from unaided test performance to academic communicative readiness becomes less self-evident. This conceptual validity argument reframes AI as a score-interpretation problem in high-stakes language testing, not only an operational issue of scoring, feedback, security, or misconduct. Synthesizing current literature in three uneven layers, the paper shows that most work treats AI as assessment infrastructure, while far less theorizes its implications for construct validity and extrapolation warrants. It defines AI-mediated construct drift as the misalignment that arises when communicative abilities required in the target domain change through AI mediation while test constructs remain anchored to an unaided-performance model. It proposes bounded AI mediation as a validity-oriented design principle: a standardized condition in which all test takers access the same institutionally controlled AI assistant, with predefined assistance boundaries, logged interactions, and tasks that distinguish comprehension support from answer generation. The paper argues that score interpretations should be narrowed and supplemented when used to support claims about AI-mediated academic communication.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑