arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

EduBehaviors:基于断言的模式,用于教育对话的可审计编码

EduBehaviors: Assertion-based Schemas for Auditable Coding of Educational Dialogues

Julian Bernado, Ana Trindade Ribeiro, Xander Beberman, Susanna Loeb

arXiv 2609.27043首次发表:更新:

发表机构

SCALE Initiative, Stanford University(斯坦福大学SCALE计划)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对LLM注释不透明的问题,提出EduBehaviors框架,通过测量可观察行为并学习分类器,在TalkMoves数据集上达到宏F1 0.673,并提供工具包以支持可审计的教育对话编码。

AI 中文摘要

大型语言模型使得与感兴趣构念相对应的教学注释能够快速部署,从而为在对话数据集上生成分类提供自然语言接口。然而,由于LLM推理的不透明性,我们无法获得可验证的、机制性的洞察,以了解模型为何为某个话语选择某个标签。我们引入了EduBehaviors框架,这是一种可解释、可扩展的教育数据注释方法,它利用LLM来测量与许多感兴趣构念相关的重复可观察行为,然后基于这些可观察行为学习一个针对该构念的分类器。我们在TalkMoves数据集上评估该框架,预测教师TalkMoves标签。我们最佳配置的宏F1得分为0.673,Cohen's kappa为0.688,证明与直接提示方法相比具有竞争力。此外,我们发布了EduBehaviors工具包,包含两个工具,使研究人员能够在其自身数据上实施EduBehaviors框架。

英文摘要

Large language models have allowed the rapid deployment of pedagogical annotations corresponding to constructs of interest, allowing a natural language interface for generating classifications on a conversational dataset. However due to the opaque nature of LLM reasoning, we have no verifiable, mechanistic insight into why a model chose a label for an utterance. We introduce the EduBehaviors framework, an interpretable, scalable approach to annotating educational data that uses LLMs to measure repeated observable behaviors relevant to many constructs of interest and then learns a classifier for the construct based on these observable behaviors. We evaluate the framework on the TalkMoves dataset, predicting the Teacher TalkMoves labels. Our best configuration results in a macro-F1 of 0.673 and 0.688 Cohen's kappa, proving competitive with direct prompting approaches. In addition, we release EduBehaviors Toolkit, two tools allowing researchers to operationalize the EduBehaviors framework in their own data.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑