GPT-4o-mini与教师平均评分在音乐分析回答自动评分中的比较验证:单次部署、可重复性及策略特异性偏差
Comparative Validation of GPT-4o-mini and Teacher Mean Scores for Automated Scoring of Music Analysis Responses: Single-Pass Deployment, Repeatability, and Strategy-Specific Bias
浏览论文内容
中文总结 AI 辅助
本研究以教师平均评分为基准,评估GPT-4o-mini对音乐分析回答的自动评分能力,发现Fs+CoT策略与教师评分一致性最强,不同策略评分特征不同,实际应用需针对性校准与人工监督。
中文摘要 AI 辅助
对开放式音乐分析回答进行评分耗时且需要对和声知识和形式理解进行细致判断。本研究以教师平均评分作为基准,评估GPT-4o-mini在基于 rubric(评分细则)的音乐分析短文评分中的有效性和可重复性。研究使用了包含300份大学水平学生回答的数据集,由教师从和声、曲式、推理、术语四个维度进行评分。GPT-4o-mini采用三种提示策略对相同回答进行评分:带思维链推理的少样本提示(Fs+CoT)、检索增强生成(RAG)、基于每次生成5个内部结果的自一致性(SC)。每种策略都进行了三次重复,期间模型、提示、评分细则和回答均保持不变。单次评分代表实际操作评分条件,而三次运行的中位数聚合用于检验鲁棒性。采用相关性、组内相关系数、Krippendorff's alpha、二次加权kappa和评分误差指标来评估与教师平均评分的一致性。Fs+CoT在单次评分和中位数聚合中均表现出与教师平均评分最强的一致性;RAG表现出系统性的过度评分;SC产生高度可重复的评分,但个体层面的一致性较弱。维度层面分析显示,评分性能随评分细则组件变化,术语维度的一致性通常弱于推理维度。这些发现表明,GPT-4o-mini可为复杂音乐分析回答生成稳定评分,但不同提示策略会产生不同的评分特征,因此实际应用需要针对策略进行校准、维度层面验证以及持续的人工监督。
英文摘要
Scoring open-ended music analysis responses is time-consuming and requires nuanced judgments of harmonic knowledge and formal understanding. This study evaluates the validity and repeatability of GPT-4o-mini for rubric-based scoring of music analysis essays, using teacher mean scores as the benchmark. A dataset of 300 university-level student responses was scored by teachers on four dimensions: Harmony, Form, Reasoning, and Terminology. GPT-4o-mini scored the same responses using three prompting strategies: few-shot prompting with chain-of-thought reasoning (Fs+CoT), retrieval-augmented generation (RAG), and self-consistency based on five internal generations per administration (SC). Each strategy was administered three times with the model, prompt, rubric, and response held constant. Single-pass scores represented an operational scoring condition, whereas median aggregation across three runs was used to examine robustness. Agreement with teacher mean scores was evaluated using correlation, intraclass correlation, Krippendorff's alpha, quadratic weighted kappa, and scoring error indices. Fs+CoT showed the strongest agreement with teacher mean scores in both single-pass scoring and median aggregation. RAG showed systematic over-scoring, whereas SC produced highly repeatable scores but weaker individual-level agreement. Dimension-level analyses showed that scoring performance varied across rubric components, with Terminology generally showing weaker agreement than Reasoning. These findings indicate that GPT-4o-mini can generate stable scores for complex music analysis responses, but prompting strategies produce distinct scoring profiles. Operational use therefore requires strategy-specific calibration, dimension-level validation, and continued human oversight.
发表机构
- Sejong University(世宗大学)
- Ewha Womans University(梨花女子大学)
机构由 AI 辅助整理,请以论文原文为准。