arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

不会学习的学习者:优化教学指标会降低大语言模型辅导效果

The learner who does not learn: when optimizing a pedagogical metric degrades LLM tutoring

Daniel Domínguez Figaredo, Rafael Fernández De la Cruz

arXiv 2610.12125首次发表:更新:

发表机构

UNED(西班牙国立远程教育大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究发现,针对教学指标微调LLM辅导模型会导致专家评分下降,因指标仅优化重复单一最优决策,研究还推导了AI辅导基准的设计原则。

AI 中文摘要

人们普遍认为,提升大语言模型(LLM)辅导的教学质量的自然方式是定义一个教学性能指标,并针对该指标对模型进行微调。为了验证这一策略,我们设计了一个教学适应性指标,该指标会根据学习情境的条件对学习序列中的每个教学决策进行评分,这是自动教学评分的标准做法。我们在2000个学习者场景中评估了一款前沿辅导模型,通过微调一个开放权重代理模型修正了其表现最差的案例,并邀请31名受过专业训练的教育工作者以盲评方式对修正前后输出的教学一致性进行评分。修正案例的指标得分从+0.05提升至+0.42,而专家评分却从4.46降至3.03;未修正的对照组得分保持不变,基础代理模型对照组则排除了模型本身变化的影响。该辅导模型表现更差的原因在于,任何对决策进行独立评分并取平均值的指标,其最大值都可通过重复单一最优决策实现,而微调后的模型无论在样本内还是样本外,都完全收敛到了这一最优解。教育工作者识别出了这种重复现象,而此类指标无法对此进行表征;在评分中保留学习者轨迹的做法虽降低了但未逆转该指标的判定。权重分析显示,修正源于模型的输出投影层,该层已记忆训练字符串而非学会适应。我们得出结论:测量有效性并不等同于优化有效性,并推导了用于评估或训练AI辅导教学的基准设计原则。

英文摘要

It is assumed that a natural way to improve the pedagogical quality of large language model tutors is to define a metric of instructional performance and fine-tune the model against it. To test this strategy, we designed a metric of pedagogical adaptivity that scores each instructional decision in a learning sequence against the conditions of the learning situation, which is the standard used for automated pedagogical scoring. We audited a frontier tutor across 2,000 learner scenarios, corrected its weakest cases by fine-tuning an open-weights proxy, and asked 31 trained educators to rate the pedagogical alignment of the outputs blind, before and after correction. The metric increased from +0.05 to +0.42 for the corrected cases, while the expert ratings decreased from 4.46 to 3.03, with the unmodified controls remaining unchanged and a base-proxy control ruling out the change of model. The tutor performed worse because any metric that scores decisions independently and averages them is maximized by repeating the single best decision, and the fine-tuned model collapsed to that exact optimum in every case, in and out of sample. Educators identified the repetition, which such metrics cannot represent, and preserving the learner's trajectory in the score reduced, but did not reverse, the metric's verdict. Weight analysis traced the correction to the model's output projection, where it had memorized its training strings rather than learned to adapt. We conclude that measurement validity does not imply optimization validity, and we derive design principles for benchmarks that assess or train AI-tutor instruction.

Comments33 pages, 7 figures, 5 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑