arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于经过验证的教学反馈分类协议的耐久性和跨语言转移基准

A Durability and Cross-Language Transfer Benchmark for a Validated Teaching-Feedback Classification Protocol

Esteban U. Vega Barajas

arXiv 2607.11873首次发表:更新:

发表机构

Universidad de Guadalajara(瓜达拉哈拉大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究教学反馈分类协议的复用性问题,通过在原始西班牙数据上跨三代表示方法重新运行,并转移到英语,发现该协议具有耐久性,模型选择是部署决策而非方法特性。

AI 中文摘要

机构收集的开放式教学评估反馈数量远超其阅读量。先前研究引入了一种经过验证的协议,通过主题类别和情感对这类评论进行分类,该协议基于文档化的注释指南、注释者内部可靠性测量、分层交叉验证以及对西班牙机构语料库采用冻结编码器设计进行的保留评估构建。有两个问题限制了其复用性:随着表示方法的发展,固定于2019年时代冻结嵌入的协议是否仍具竞争力,以及它是否能转移到第二语言。我们在原始西班牙数据上跨三代表示方法重新运行该协议,并将其情感任务转移到英语。我们发现该协议具有耐久性,模型选择是部署决策而非方法特性。

英文摘要

Institutions collect far more open-ended teaching-evaluation feedback than they read. A prior study introduced a validated protocol for classifying such comments by thematic category and sentiment, built from a documented annotation guide, an intra-annotator reliability measurement, stratified cross-validation, and a held-out evaluation on a Spanish institutional corpus with a frozen-encoder design. Two questions limit its reuse: whether a protocol fixed to 2019-era frozen embeddings stays competitive as representation methods advance, and whether it transfers to a second language. We re-run it on the original Spanish data across three representation generations, sparse lexical features, frozen transformer embeddings, and prompted large language models, and transfer its sentiment task to English with a balanced 45,000-comment corpus checked against an aspect-labeled education dataset. Treating paired comparisons as descriptive, we find the protocol durable: a 2026 frontier model posts the highest thematic F1 on the hardest Spanish task, yet shows no sentiment advantage over a cheap model and no descriptive separation from it on English, so model choice is a deployment decision, not a property of the method.

Comments12 pages, 2 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑