arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31095cs.HC

自信而非更明智:人机交互中的邓宁-克鲁格效应

Confident, Not Wiser: The Dunning-Kruger Effect in Human-AI Interaction

Daniela Fernandes, Michelle Rausch, Agnes Mercedes Kloft, Daniel Buschek, Robin Welsch

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过对比人类单独与人类+AI在推理任务上的表现,发现AI辅助提升成绩但未改善自我评估,且加剧了邓宁-克鲁格效应,强调了元认知增强与表现增强的区分及界面设计启示。

中文摘要 AI 辅助

AI辅助可以提高表现,但不会改善自我评估。我们报告了一项研究(N=366),比较了人类单独和人类+AI在推理任务上的表现,其中AI模型在同一项目上进行了基准测试。参与者估计了整体和分块表现,并对他们的答案进行了信心评级。人类+AI组获得了更高的分数,但自我估计与表现的相关性较弱。两组的平均过度估计相似,覆盖了个体错误。在不同任务中,人类+AI组中信心区分正确答案与错误答案的准确性较低,而任务内差异仍不确定。两组都发现了邓宁-克鲁格模式,人类+AI组观察到的对比更大。对分数噪声的控制减少了但未消除该模式,受控组差异仍无定论。一个扩展的计算模型描述了整体和分块估计。我们的发现区分了表现增强与元认知增强,并激励支持验证、传达任务特定AI模型性能以及帮助用户评估其联合工作质量而非产生答案的界面。

英文摘要

AI assistance can improve performance without improving self-assessment. We report a study (N=366) comparing Human alone and Human+AI performance on reasoning tasks, for which the AI model is benchmarked on the same items. Participants estimated global and block performance and rated confidence in their answers. Human+AI achieved higher scores, but self-estimates tracked performance weakly. Average overestimation was similar across groups, covering individual errors. Across tasks, confidence distinguished correct from incorrect answers less accurately in the Human+AI group, while within-task differences remained uncertain. The Dunning-Kruger pattern was found in both groups, with a larger observed contrast in Human+AI. Controls for score noise reduced but did not eliminate the pattern, with the controlled group difference remaining inconclusive. An extended computational account describes global and block estimates. Our findings distinguish performance augmentation from metacognitive augmentation and motivate interfaces that support verification, communicate task-specific AI model performance, and help users evaluate the quality of their joint work rather than produce answers.

发表机构

  • Aalto University(阿尔托大学)
  • University of Bayreuth(拜罗伊特大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑