校准可信度:为教育领域评估大型语言模型(LLM)协同设计指标与可视化方案
Calibrating Trustworthiness: Co-Designing Metrics and Visualizations for Evaluating LLMs in Education
浏览论文内容
中文总结 AI 辅助
本研究针对教育领域LLM响应的教学适配性评估缺口,通过与学习工程师协同设计可信度指标与可视化工具,提升了评估信度,为相关工具提出设计准则。
中文摘要 AI 辅助
大型语言模型(LLM)正在重塑教育技术,但针对其响应的教学适配性评估仍未得到充分探索,目前高度依赖构建该技术的学习工程师的专业知识。为弥合这一差距,本研究将可信度作为结构化评估视角,利用现有LLM可信度衡量指标系统识别潜在的教学干扰。通过与开发LLM驱动数字教材的学习工程师开展纵向协同设计流程,本研究完成了三项工作:(1)共同构建了包含20项具体指标的5项可信度指标,这些指标专为教学用途定制;(2)设计了可将可信度违规情况映射至LLM响应的可视化方案;(3)评估了这些工具如何帮助学习工程师对LLM响应进行A/B测试比较。将可信度明确化后,评分者间信度有所提升,同时帮助学习工程师解决了目标冲突并做出更一致的判断。本研究探讨了可信度作为教育领域LLM评估视角的新兴优势,并为未来评估工具提出了新的设计准则,以构建与教学适配的LLM驱动学习工具。
英文摘要
LLMs are reshaping educational technology, yet evaluating their responses for pedagogical alignment remains underexplored, relying heavily on the expertise of learning engineers building the technology. To bridge this gap, we explore trustworthiness as a structured lens for evaluation, leveraging existing measures of LLM trustworthiness to systematically identify potential pedagogical disruptions. Through a longitudinal co-design process with learning engineers developing an LLM-powered digital textbook, we: (1) co-constructed five trustworthiness metrics comprising 20 measures tailored to pedagogical use; (2) designed visualizations that map trustworthiness violations onto LLM responses; and (3) evaluated how these tools help learning engineers make A/B comparisons of LLM responses. Making trustworthiness explicit increased inter-rater reliability while helping learning engineers resolve conflicting objectives and produce more consistent judgments. We discuss the emergent benefits of trustworthiness as a lens for evaluating LLMs in education and propose new design guidelines for future evaluation tools that enable pedagogically-aligned, LLM-powered learning tools.