可用但未被认领:人机协同的实证研究
Available but Unclaimed: An Empirical Study of Human-AI Synergy
浏览论文内容
中文总结 AI 辅助
本研究通过大规模被试实验发现,人机协同中LLM能力提升仅部分转化为用户表现,且遵从度随能力变化,呼吁设计支持选择性弃权以保留独立推理。
中文摘要 AI 辅助
人们越来越多地依赖大型语言模型(LLM)进行推理,然而互补的能力并不能保证整体表现优于两个组成部分。在一项被试间研究中,参与者(N=535)在无辅助或分别使用GPT-5.6-Luna、Claude Opus 4.8、Gemini 3.6 Flash或Kimi K3的情况下,解决了包含矩阵推理、心理旋转、三段论和字母串类比共40道题的测试集。每次辅助试验都要求参与者咨询模型。每个模型在匹配的引出条件下,对每个题目单独作答100次。辅助与无辅助准确率之差随题目层面的LLM能力提高而增加。在不同任务中,遵从度有所变化,且在任务内部随能力提高而增加。与无辅助信心相比,获得建议后的信心区分正确与错误答案的能力较弱。在参考比较中,LLM准确率提升的大约一半传递到了辅助准确率中。这一准确率增益中有多少到达参与者,在不同模型间存在差异。这些发现促使我们评估LLM与人类交互时的表现,并设计支持选择性遵从(弃权(不执行))的机制,以保留独立推理。
英文摘要
People increasingly reason with large language models (LLMs), yet complementary capabilities do not guarantee outperforming both components. In a between-subjects study, participants (N=535) solved a 40-item battery of matrix reasoning, mental rotation, syllogisms, and letter-string analogies, unaided or with GPT-5.6-Luna, Claude Opus 4.8, Gemini 3.6 Flash, or Kimi K3. Each assisted trial required consultation with the model. Each model answered every item alone 100 times under matched elicitation. The assisted-unaided accuracy difference increased with item-level LLM competence. Deference varied across tasks and increased with competence within tasks. Post-advice confidence distinguished correct from incorrect answers less strongly than unaided confidence. In a reference comparison, about half the increase in LLM accuracy carried through to assisted accuracy. How much of that accuracy gain reached participants differed across the models. These findings motivate evaluating LLMs in interaction with humans and designing support for selective deference that preserves independent reasoning.
发表机构
- Aalto University(阿尔托大学)
- University of Innsbruck(因斯布鲁克大学)
- Humboldt-Universität zu Berlin(柏林洪堡大学)
- Ludwig-Maximilians-Universität München(慕尼黑大学)
机构由 AI 辅助整理,请以论文原文为准。