将大语言模型评判的有用性重思为教学信号:一项针对导师模型的预注册审计
Rethinking LLM-Judged Helpfulness as a Pedagogy Signal: A Pre-Registered Audit Across Tutor Models
浏览论文内容
中文总结 AI 辅助
本研究通过预注册审计发现,通用LLM有用性评分标准无法作为可靠教学信号,其排序具有评判者依赖性,而教学评分标准可有效区分策略,故导师评估应采用教学针对性评分标准与确定性过程测量。
中文摘要 AI 辅助
大语言模型辅导存在一个测量问题:通用的有用性评分标准能否区分直接给出答案与教学指导?我们在一项预注册研究中审计该信号。在三个导师基准中,我们比较了由同一底层模型实例化的对话式与教学式策略,并搭配一个固定的弱模拟学生。确定性检测器用于测量答案泄露与下一轮独立作业。Claude Opus 4.8是冻结的、无条件的主要评判者。在Opus评分确定后,GPT-5.6 Sol被预先指定用于事后稳健性审计,针对1179个确认性答案阶段导师轮次,使用冻结的有用性与教学评分标准。在主要基准上,Opus评判下,策略在有用性上无显著差异,但在教学评分标准下完全秩分离(Cliff's |δ|=0.10对1.0)。在两个评判者中,教学对比在检测到时保持方向,而有用性排序具有评判者依赖性,在三个基准中的两个上在评判者间反转。在仅Opus的 ablation中,七个主要基准策略在平均评判教学性上跨度2.3个点,而平均评判有用性在0.25个点的区间内。此外,在每个基准上,答案揭示轮次后学生独立作业更少,该结果因构造而具有评判者不变性。在该受控环境中,通用有用性并非可靠的教学信号,导师评估应将教学针对性评分标准与确定性过程测量相结合。
英文摘要
LLM tutoring poses a measurement problem: can a general-purpose helpfulness rubric distinguish direct answer-giving from pedagogical guidance? We audit this signal in a pre-registered study. Within each of three tutor bases, we compare conversational and pedagogical policies instantiated with the same underlying model and paired with one fixed weak simulated student. Deterministic detectors measure answer leakage and next-turn independent work. Claude Opus 4.8 is the frozen, condition-blind primary judge. After the Opus scores were fixed, GPT-5.6 Sol was prospectively specified for a post hoc robustness audit of the same 1,179 confirmatory answer-phase tutor turns under the frozen helpfulness and pedagogy rubrics. On the primary base under Opus, the policies do not differ significantly in helpfulness but are perfectly rank-separated under the pedagogy rubric (Cliff's $|δ|{=}0.10$ vs. $1.0$). Across the two judges, pedagogy contrasts retain their direction where detected, whereas the helpfulness ordering is judge-contingent, reversing between judges on two of three bases. In an Opus-only ablation, seven primary-base policies span $2.3$ points in mean judged pedagogy within a $0.25$-point band of mean judged helpfulness. Separately, answer-revealing turns are followed by less independent student work on every base, a result that is judge-invariant by construction. In this controlled setting, general-purpose helpfulness is not a reliable pedagogy signal. Tutor evaluation should pair pedagogy-targeted rubrics with deterministic process measures.