探究人类与AI在多解问题上的差异
Investigating Human--AI Discrepancies via Multiple-Solution Problems
浏览论文内容
中文总结 AI 辅助
本研究通过270个多解推理谜题,发现AI模型间的解分布相似度远高于其与人类的相似度,且模型分布多样性更低,揭示人机问题解决过程的显著差异。
中文摘要 AI 辅助
前沿人工智能(AI)模型的基准测试基于其是否得出正确答案。然而,许多问题存在多个正确答案,不同的人或同一模型多次采样得到的重复尝试,会描绘出这些答案上的一个分布。在本工作中,我们探讨人类与模型的推理是否会导致在有效解上的不同分布。我们的测试平台包含五个谜题家族中的270个推理谜题。这些多解谜题每个有3到8个有效解,并且足够简单,使得人类和模型都能可靠地解决它们。由此产生的分布差异显著:模型彼此之间不同,但相互之间的相似度远高于它们与人类的相似度。此外,在每个谜题家族内,模型的分布比人类的分布多样性更低。我们跨谜题类别比较这些差异,并追踪它们对推理努力设置、提示以及保持谜题解不变的扰动的响应。综合来看,这些结果指出了人类与AI问题解决过程之间的显著差异,以及它们在同等可辩护解中的选择。随着AI系统在社会中的逐步部署成为焦点,评估此类差异(超越一维准确性指标)变得越来越重要。数据和代码可在以下网址获取:https://this https URL
英文摘要
Frontier artificial intelligence (AI) models are benchmarked on whether they reach a correct answer. Yet many problems admit several correct answers and repeated attempts, by different people or by the same model resampled, trace out a distribution over them. In this work, we ask whether human and model reasoning lead to different distributions over valid solutions. Our testbed comprises 270 reasoning puzzles across five puzzle families. These multiple-solution puzzles each have 3 to 8 valid solutions and are simple enough that humans and models can solve them reliably. The resulting distributions differ markedly: models differ from one another, yet resemble each other far more than they resemble humans. Model distributions are, moreover, within every puzzle family, less diverse than human ones. We compare these discrepancies across puzzle categories, and trace how they respond to reasoning-effort settings, to prompting, and to perturbations of the puzzle that leave its solutions unchanged. Together, these results point at significant differences between human and AI problem-solving processes, and their choice among equally defensible solutions. As progressive deployment of AI systems in society comes into focus, evaluating such differences (beyond one-dimensional accuracy metrics) is increasingly important. Data and code are available at https://hai-discrepancies.github.io/
发表机构
- Stanford University(斯坦福大学)
机构由 AI 辅助整理,请以论文原文为准。