发表机构
SCALE Initiative, Stanford University(斯坦福大学SCALE计划)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对LLM注释不透明的问题,提出EduBehaviors框架,通过测量可观察行为并学习分类器,在TalkMoves数据集上达到宏F1 0.673,并提供工具包以支持可审计的教育对话编码。
AI 中文摘要
大型语言模型使得与感兴趣构念相对应的教学注释能够快速部署,从而为在对话数据集上生成分类提供自然语言接口。然而,由于LLM推理的不透明性,我们无法获得可验证的、机制性的洞察,以了解模型为何为某个话语选择某个标签。我们引入了EduBehaviors框架,这是一种可解释、可扩展的教育数据注释方法,它利用LLM来测量与许多感兴趣构念相关的重复可观察行为,然后基于这些可观察行为学习一个针对该构念的分类器。我们在TalkMoves数据集上评估该框架,预测教师TalkMoves标签。我们最佳配置的宏F1得分为0.673,Cohen's kappa为0.688,证明与直接提示方法相比具有竞争力。此外,我们发布了EduBehaviors工具包,包含两个工具,使研究人员能够在其自身数据上实施EduBehaviors框架。
英文摘要
Large language models have allowed the rapid deployment of pedagogical annotations corresponding to constructs of interest, allowing a natural language interface for generating classifications on a conversational dataset. However due to the opaque nature of LLM reasoning, we have no verifiable, mechanistic insight into why a model chose a label for an utterance. We introduce the EduBehaviors framework, an interpretable, scalable approach to annotating educational data that uses LLMs to measure repeated observable behaviors relevant to many constructs of interest and then learns a classifier for the construct based on these observable behaviors. We evaluate the framework on the TalkMoves dataset, predicting the Teacher TalkMoves labels. Our best configuration results in a macro-F1 of 0.673 and 0.688 Cohen's kappa, proving competitive with direct prompting approaches. In addition, we release EduBehaviors Toolkit, two tools allowing researchers to operationalize the EduBehaviors framework in their own data.