训练大型语言模型助教以保留答案:一种用于调整苏格拉底式行为的监督架构与证据驱动方法
Teaching a Large Language Model Tutor to Withhold the Answer: A Supervisor Architecture and an Evidence-Driven Method for Tuning Socratic Behavior
浏览论文内容
中文总结 AI 辅助
本研究提出一种监督架构与证据驱动方法,训练LLM助教执行答案保留,通过自动评估调整苏格拉底式行为,使其在全部四个接受标准上达到完全合规。
中文摘要 AI 辅助
一名高效的大型语言模型(LLM)助教常常需要拒绝提供它能轻易生成的答案。在一项随机研究中,使用未受约束聊天机器人的学生在练习时得分更高,但在未使用该机器人的后续测试中得分更低;而同一模型的苏格拉底式受约束版本则保留了练习时的得分提升,并消除了后续的得分损失[4]。因此,可靠的答案保留是助教价值的核心,但一个有能力的模型在面对沮丧学生的催促时,仅通过提示无法可靠地执行保留操作。我们报告了一个已部署的助教系统,该系统将答案保留作为每一轮可机器验证的契约执行,以及一种针对证据调整该保留行为的方法。一个非LLM的策略核心仅读取可信的学习者状态,为八层帮助阶梯设置每一轮的上限;一个确定性检测器会剥离解决方案代码;还有一个独立的LLM评判员会对照契约检查每一条有风险的回复。我们使用无需人类受试者的自动评估来调整该行为:脚本化的学生角色会在实时流程中运行,并由更强的模型重新评分,我们记录每一次拒绝的陈述原因,以便按原因修复失败。这一过程揭示了一条可解释的“过度帮助阶梯”,从明显的解决方案泄露,到命名确切的错误,再到过度引用一般事实,每一次修复都会暴露下一个问题。该助教在全部四个接受标准上达到了完全合规。我们提供了“测量-诊断-修复”循环,作为任何必须拒绝自身具备的某一能力的LLM智能体的可复用方案。
英文摘要
An effective large language model (LLM) tutor must often decline to give an answer it could easily produce. In a randomized study, students who used an unguarded chatbot scored higher while practicing but lower on a later test taken without it, whereas a Socratically guarded version of the same model kept the practice gain and removed the later loss [4]. Reliable answer-withholding is therefore central to a tutor's value, yet a capable model pressed by a frustrated student does not withhold reliably on a prompt alone. We report a deployed tutoring system that enforces answer-withholding as a per-turn, machine-checkable contract, and a method for tuning that withholding against evidence. A non-LLM policy core, reading only trusted learner state, sets a per-turn ceiling on an eight-rung help ladder; a deterministic detector strips solution code; and a separate LLM judge checks each risky reply against the contract. We tune the behavior with an automated evaluation that uses no human subjects: scripted student personas are driven through the live pipeline and re-scored by a stronger model, and we record each rejection's stated reason so failures are fixed by cause. Doing so revealed an interpretable "over-help ladder," from blatant solution leaks, to naming the exact bug, to over-citing general facts, with each fix exposing the next. The tutor reached full compliance on all four acceptance criteria. We offer the measure, diagnose, and fix loop as a reusable recipe for any LLM agent that must refuse a capability it has.