发表机构
University of California, Berkeley(加州大学伯克利分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
LeanSide是一个形式化验证的协同推理系统,将学生自然语言证明自动形式化为Lean并反馈可理解信息,通过课堂部署分析系统属性对学习进展的影响,为人类-AI协同推理系统设计提供启示。
AI 中文摘要
大型语言模型越来越多地被用作演绎推理任务的协作者,但其输出可能产生幻觉或将用户偏离预期的推理路径。形式化证明助手提供机器检查的验证,但学习曲线陡峭,且要求比人类书面证明更细粒度的推理。我们探索了一种结合这些优势的界面,允许用户编写和修改自由形式的自然语言证明,同时一个经过验证的后端检查其推理并以用户的粒度返回反馈。我们在本科数学教育背景下研究这一界面,开发了LeanSide,一个形式化验证的协同推理系统,该系统将学生的推理自动形式化为Lean,并将验证器的输出非形式化为可理解的反馈。我们通过课堂部署进行了用户研究,并分析了哪些系统属性帮助学生取得进展,哪些导致他们陷入困境。我们利用这些发现为在人类-AI协同推理系统中使用形式化验证后端的设计启示。
英文摘要
Large language models are increasingly used as collaborators on deductive-reasoning tasks, but their outputs can hallucinate or pull users away from intended reasoning. Formal proof assistants provide machine-checked verification, but have a steep learning curve and require more granular reasoning than human written proofs. We explore an interface that combines these strengths, allowing users to write and revise free-form natural-language proofs while a verified backend checks their reasoning and returns feedback at the user's granularity. We study this interface in the context of undergraduate mathematics education by developing LeanSide, a formally verified co-reasoning system, which auto-formalizes student reasoning into Lean and informalizes verifier output into understandable feedback. We conducted user studies through classroom deployment and analyzed which system properties helped students make progress and which caused them to get stuck. We use these findings to derive design implications for using a formally verified backend in human-AI co-reasoning systems.
Comments16 pages, 13 figures