arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2509.05346cs.AI

超越基准测试:面向个性化学习的大语言模型场景化评估

Beyond Benchmarking: Scenario-Based Evaluation of Large Language Models for Personalized Learning

  • The University of Queensland(昆士兰大学)
  • AI U Course(AI优课)

机构由 AI 辅助整理,请以论文原文为准。

Bo Yuan, Jiazi Hu

AI总结:

本研究提出场景化评估框架,以Gemini为外部评估器结合Bradley-Terry模型等,发现不同LLMs在个性化学习场景中教学行为存在差异,为AI增强教育提供有意义的模型行为信号。

AI中文摘要:

尽管大语言模型(LLMs)正越来越多地被用于支持个性化学习,但人们对其在真实学习场景中的教学行为差异仍知之甚少。现有的评估实践往往强调基准分数和整体模型排名,但这类方法通常无法深入洞察LLMs如何诊断学生的理解情况并生成个性化指导。本研究提出一种场景化评估框架,以细致考察LLMs在个性化学习支持中的行为。以课后辅导场景为例,将一名学生对一组数据结构问题的回答组成的数据集提供给多个LLMs,要求每个模型识别潜在的知识概念、推断学生的掌握情况并生成个性化改进指导。为支持一致、可复现且可扩展的比较,采用Gemini作为外部评估器,从诊断准确性、教学清晰度、可操作性、错误概念识别以及适配学生水平等多个与教学相关的维度开展评估。随后使用Bradley-Terry模型对得到的成对偏好进行拟合,以得出比较强度估计值,同时结合定性分析和语义可视化,进一步考察反馈结构、诊断深度以及建议针对性方面的差异。关键发现表明,不同LLMs在同一学习场景中表现出可区分的教学行为,且该场景化评估方法可提供具有教育意义的信号,用于理解AI增强教育中的模型行为。

英文摘要:

While large language models (LLMs) are increasingly being adopted to support personalized learning, there remains limited understanding of how their pedagogical behaviors differ in authentic learning scenarios. Existing evaluation practices often emphasize benchmark scores and overall model rankings, but such approaches usually provide limited insight into how LLMs diagnose student understanding and generate personalized guidance. This study proposes a scenario-based evaluation framework for closely examining LLM behavior in personalized learning support. Using a post-class tutoring setting as an illustrative example, a dataset comprising a student's responses to a set of data structures questions is provided to multiple LLMs. Each model is required to identify the underlying knowledge concepts, infer the student's mastery profile, and generate personalized guidance for improvement. To support consistent, reproducible and scalable comparison, Gemini is employed as an external evaluator across multiple pedagogically relevant dimensions, including diagnostic accuracy, instructional clarity, actionability, misconception identification, and appropriateness to the student's level. The resulting pairwise preferences are then fitted using the Bradley-Terry model to derive comparative strength estimates, while qualitative analysis and semantic visualization are used to further examine differences in feedback structure, diagnostic depth, and recommendation specificity. The key findings show that different LLMs exhibit distinguishable pedagogical behaviors within the same learning scenario and demonstrates how scenario-based evaluation can provide educationally meaningful signals for understanding model behavior in AI-enhanced education.

↑