arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.21306cs.LGcs.AIcs.CLcs.CYcs.HC

人工智能助手过度协助

AI Assistants Overassist

发表机构斯坦福大学 · 加利福尼亚大学圣地亚哥分校 · 华盛顿大学
查看机构详情
  • Stanford University(斯坦福大学)
  • University of California, San Diego(加利福尼亚大学圣地亚哥分校)
  • University of Washington(华盛顿大学)

机构由 AI 辅助整理,请以论文原文为准。

Verona Teo, Raghav Jain, Tobias Gerstenberg, Max Kleiman-Weiner

首次发表
浏览论文内容

中文总结 AI 辅助

研究探讨大语言模型作导师时,其干预方式对学习的影响。介绍Int-Bench基准,模拟学生解题、教师决定干预。在多领域评估大语言模型教师干预情况,并与人类对比,发现其干预频繁且倾向提供完整方案,表明为短期成功而非深层学习优化。

中文摘要 AI 辅助

大语言模型(LLMs)越来越多地被用作导师和思维伙伴来帮助用户解决问题。虽然人工智能助手的指导可以辅助思考和促进学习,但这些益处取决于其帮助方式,过早或过于频繁的干预可能会阻碍真正的学习和认知参与。目前对人工智能系统在解决问题过程中如何做出干预决策了解甚少。本文介绍了Int-Bench,一个基于模拟的基准,用于评估学习过程中LLM的干预。Int-Bench模拟“学生”解决问题,“教师”监控学生推理并决定是否、何时以及如何干预。在代码调试、数学和脑筋急转弯三个领域评估LLM教师的干预频率、时机及其对即时任务成功和新问题泛化的影响。还将LLMs与人类比较,发现LLMs比人类干预更频繁、更早,且倾向于提供完整解决方案而非有针对性的提示。这些发现表明当前的LLM助手通常为短期成功而非支持深度学习和长期成功所需的推理过程进行优化。

英文摘要

Large language models (LLMs) are increasingly used as tutors and thought partners, helping users reason through problems. While guidance from AI assistants can scaffold thinking and foster learning, such benefits depend on how they help--for instance, intervening too early or too frequently may hinder true learning and cognitive engagement. Yet how AI systems navigate intervention decisions during problem-solving remains poorly understood. Here, we introduce Int-Bench, a simulation-based benchmark for evaluating LLM interventions during learning. Int-Bench simulates a "student" solving a problem while a "teacher" monitors the student's reasoning and decides whether, when, and how to intervene. Across three domains--code debugging, mathematics, and brain teasers--we evaluate LLM teachers on the frequency and timing of interventions, as well as their impact on both immediate task success and generalization to new problems. We also compare LLMs to humans, finding that LLMs intervene more frequently and earlier than humans. Moreover, in contrast to humans, they tend to provide complete solutions rather than targeted hints. These findings suggest that current LLM assistants often optimize for short-term success rather than supporting the reasoning processes needed for deeper learning and long-term success.

↑