高效评分学习:面向低成本大语言模型作文评分的多臂老虎机驱动提示选择框架
Learning to Grade Efficiently: A Bandit-Driven Prompt-Selection Framework for Low-Cost LLM Essay Scoring
浏览论文内容
中文总结 AI 辅助
本研究针对自动作文评分的固定提示选择策略成本高的问题,提出多臂老虎机驱动的自适应提示选择框架,在保证准确率的同时减少78.4%的LLM调用,生成成本-可靠性学习曲线,为教育技术平台提供平衡成本与有效性的方案。
中文摘要 AI 辅助
大语言模型(LLMs)在自动作文评分(AES)中展现出强大能力,但现有方法通常采用固定提示选择策略,未考虑运营成本问题及最优配置的动态变化。本文提出一种新型成本感知方法,将每种提示类型视为多臂老虎机(MAB)控制器中的一个臂,从而在推理过程中自适应选择最优提示策略。我们在雅思写作任务2作文上开展实验,结果显示,该MAB框架在达到与穷尽网格搜索相当的评分准确率的同时,能减少78.4%的LLM调用次数以找到最佳评分方法。我们实现了四种不同的评分方案(多步与单步评估、带校准示例与不带校准示例),发现带示例的多步方法准确率最高。通过跟踪令牌使用量、延迟及一致性指标,我们生成了首个用于作文评分的成本-可靠性学习曲线,为需平衡运营成本与评估有效性的教育技术平台提供可操作的见解。本研究首次将在线控制机制应用于AES中的自适应提示策略选择,将提示选择从离线超参数优化问题转变为高效的在线学习任务。
英文摘要
Large Language Models (LLMs) demonstrate strong capabilities in automated essay scoring (AES), but contemporary approaches typically employ fixed prompt selection, failing to address operational cost concerns and evolving optimal configurations. We propose a novel cost-aware approach that treats each prompt type as an arm in a multi-armed bandit (MAB) controller, enabling adaptive selection of optimal prompting strategies during inference. Our experiments on IELTS Writing Task 2 essays show that the MAB framework achieves comparable scoring accuracy to exhaustive grid search while reducing LLM calls by 78.4\% to find the best grading approach. We implemented four distinct grading recipes (multi-step vs. single-step assessment, with vs. without calibration examples) and found that the multi-step approach with examples achieves the highest accuracy. By tracking token usage and latency alongside agreement metrics, we produce the first cost-reliability learning curves for essay scoring, providing actionable insights for educational technology platforms that must balance operational costs against assessment validity. This work represents the first application of online control mechanisms to adaptively select prompting strategies in AES, transforming prompt selection from an offline hyperparameter optimization problem into an efficient online learning task.
发表机构
- Carleton University(卡尔顿大学)
机构由 AI 辅助整理,请以论文原文为准。