arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.27939cs.CLcs.LG

从情感分类到可操作且负责任的反馈:2015-2026年学生教学评价中NLP的应用范围综述与证据图谱

From Sentiment Classification to Actionable and Responsible Feedback: A Scoping Review and Evidence Map of NLP in Student Evaluation of Teaching, 2015-2026

Jeff Eicher, Rafael da Silva

首次发表
浏览论文内容

中文总结 AI 辅助

本综述通过证据图谱揭示学生教学评价NLP研究存在行动性断层:从可演示输出(61.3%)到预期用户评估(11.6%)大幅下降,情感分析为主流任务,公平性指标罕见,强调技术多样性未带来教育价值提升。

中文摘要 AI 辅助

自然语言处理(NLP)应用于开放式教学评价评论(学生教学评价,SET)已追踪了该领域的技术演变——从词典和传统分类器到Transformer和大语言模型(LLM)——但尚不明确这种技术多样化是否伴随着教育价值和证据稳健性的相应提升。本范围综述(PRISMA-ScR)沿技术轴(研究问题1)和四个价值维度(研究问题2-5)绘制了421项研究(2015-2026年,2026年为部分数据)。采用双重互盲的LLM筛选,辅以抽样人工裁决,编码了七个提取领域,并在综合阶段进行了针对性的编码边界审查。联合图谱中最显著的量化差距是行动性断层:已展示输出或更强(A2+:258/421;61.3%)与预期用户评估或更强(A3+:49/421;11.6%)之间的差距,下降了49.7个百分点。情感分析仍是最常见的任务(300/421);诊断和生成深度占相当少数(D4-D5:占已解决案例的28.2%);正式公平性指标罕见(1.9%)。研究结果具有描述性,不支持关于进展的因果主张:技术共存和不均匀报告是图谱的一部分,但A2+到A3+的断层是贡献,而非质量阶梯。

英文摘要

Natural language processing (NLP) applied to open-ended teaching-evaluation comments (Student Evaluation of Teaching, SET) has tracked the field's technical evolution--from lexicons and conventional classifiers to transformers and large language models (LLMs)--but it is not evident that this technical diversification has been accompanied by corresponding gains in educational value and robustness of the evidence. This scoping review (PRISMA-ScR) maps 421 studies (2015-2026, 2026 partial) along a technical axis (RQ1) and four value dimensions (RQ2-RQ5). Dual mutually blinded LLM screening with sampled human adjudication coded seven extraction domains, with targeted codebook-boundary review at synthesis. The joint map's sharpest quantified gap is the actionability discontinuity: demonstrated output or stronger (A2+: 258/421; 61.3%) versus intended-user evaluation or stronger (A3+: 49/421; 11.6%), a 49.7 percentage-point drop. Sentiment analysis remains the modal task (300/421); diagnostic and generative depth is a substantial minority (D4-D5: 28.2% of resolved cases); a formal fairness metric is rare (1.9%). The findings are descriptive and do not support causal claims of progress: technological coexistence and uneven reporting are part of the map, but the A2+ to A3+ cliff is the contribution, not a quality ladder.

发表机构

  • Eastern University(东部大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑