arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.25790cs.CY

自动化建设性评估:利用大型语言模型实现实践能力的可扩展与重复评估

Automating Constructive Assessment with Large Language Models: Toward Scalable and Repeated Evaluation of Practical Competence

Satoshi Takahashi, Atsushi Yoshikawa, Megumi Kose, Kenichi Suzuki, Chieko Inoue, Yumi Watanabe, Mari Sawada

首次发表
浏览论文内容

中文总结 AI 辅助

本研究利用大型语言模型自动化层级诊断推理评估,通过提示设计生成案例问题、评分和反馈,实现高一致性、低成本且可重复的建设性评估。

中文摘要 AI 辅助

本研究旨在通过开发和测试使用大型语言模型的评估过程,自动化层级诊断推理(HDR),这是一种评估实践判断技能的建设性方法。HDR是一项描述性任务,通过要求学生识别并解释基于案例研究的问题中的错误来衡量高阶认知技能。然而,开发和评估HDR需要专业知识和大量精力。因此,我们提出并通过实证验证了自动(1)生成包含符合教育意图的错误的案例问题,(2)对描述性答案进行评分,以及(3)基于错误答案生成结构化反馈,这些仅通过提示设计实现,无需微调。生成问题的内部一致性和结构效度得到了实验分数分布和Cronbach's alpha(0.78)的支持。在某些条件下,自动评分与人工评分的一致性达到100%。反馈被认为与人类教师的反馈一样具有说服力和实用性,展示了一个以可重复性、即时性和低成本实施基于HDR的建设性评估的实用框架。大型语言模型的灵活性还将支持在保持结构的同时进行重复和纵向评估,显示出在教育环境中广泛应用的潜力。

英文摘要

This study aimed to automate hierarchical diagnostic reasoning (HDR), a constructive method for evaluating practical judgment skills, by developing and testing an evaluation process using a large language model. HDR is a descriptive task that measures higher-order cognitive skills by requiring students to identify and explain errors in case-study-based problems. However, it requires expertise and effort to develop and evaluate. Hence, we proposed and empirically validated the automatic (1) generation of case problems containing errors aligned with educational intentions, (2) scoring of descriptive answers, and (3) generation of structured feedback based on incorrect answers, achieved solely through prompt design without fine-tuning. The internal consistency and construct validity of the generated problems were supported by the experimental score distribution and Cronbach's alpha (0.78). The agreement between automated and human ratings reached 100% under some conditions. The feedback was rated as being as convincing and useful as that from human instructors, demonstrating a practical framework for implementing HDR-based constructive assessment with reproducibility, immediacy, and low cost. The flexibility of large language models will also enable repeated and longitudinal assessments while maintaining structure, showing broad potential for application in educational settings.

发表机构

  • Nagoya University(名古屋大学)
  • GLOBIS University(GLOBIS大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑