arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.22553cs.CL

正确诊断,更好反馈:用于逻辑证明中忠实LLM辅导反馈的符号验证器

Correct Diagnosis, Better Feedback: A Symbolic-Verifier for Faithful LLM Tutoring Feedback in Logic Proofs

发表机构北卡罗来纳州立大学 · 普渡大学 · 肯尼索州立大学
查看机构详情
  • North Carolina State University(北卡罗来纳州立大学)
  • Purdue University(普渡大学)
  • Kennesaw State University(肯尼索州立大学)

机构由 AI 辅助整理,请以论文原文为准。

Tahreem Yasir, Arnav Mody, Xioayi Tian, Tiffany Barnes

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出基于符号验证器的架构,分离诊断与生成,在命题逻辑证明辅导中实现更高诊断正确性,并强调反馈忠实性需与正确性分开评估。

中文摘要 AI 辅助

有效的LLM辅导依赖于在生成反馈之前正确识别学生推理中的具体错误。我们在命题逻辑证明辅导中研究这一问题,其中学生行为可以对照形式推理规则进行检查。我们引入了一种基于验证器的架构,将诊断与语言生成分离。使用600个平衡的学生行为,我们比较了零样本LLM检测器、微调检测器和符号验证器。每个诊断由共享的推理和反馈代理处理,从而隔离初始诊断的影响。零样本检测器的宏F1为0.191;微调将其提升至0.709,但在结构相关的类别之间仍保留系统性错误。推理通常保留提供给它们的诊断,表明错误的诊断可以在流水线中被忠实地传播。反馈同样可以对其推理保持忠实、不泄露信息,并在教学上适当,同时针对错误的错误。基于验证器的反馈实现了最高的诊断正确性,专家评分与自动反馈评估大体不一致。这些发现表明,表面上的反馈质量可能掩盖上游诊断错误,并且忠实性必须与正确性分开评估。我们的代码公开可用。

英文摘要

Effective LLM tutoring depends on correctly identifying the specific error in a student's reasoning before generating feedback. We study this problem in propositional-logic proof tutoring, where student actions can be checked against formal inference rules. We introduce a verifier-grounded architecture that separates diagnosis from language generation. Using 600 balanced student actions, we compare a zero-shot LLM detector, a fine-tuned detector, and a symbolic verifier. Each diagnosis is processed by shared rationale and feedback agents, isolating the effect of the initial diagnosis. The zero-shot detector achieves a macro-F1 of 0.191; fine-tuning raises this to 0.709 but retains systematic errors between structurally related classes. Rationales generally preserve the diagnosis supplied to them, showing that an incorrect diagnosis can be faithfully propagated through the pipeline. Feedback can likewise remain faithful to its rationale, non-revealing, and pedagogically appropriate while addressing the wrong error. Verifier-grounded feedback achieves the highest diagnostic correctness, and expert ratings largely uneven with the automatic feedback evaluations. These findings show that apparent feedback quality can conceal upstream diagnostic errors and that faithfulness must be evaluated separately from correctness. Our code is publicly available

↑