arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越“ChatGPT 也会犯错”:设计干预措施以支持 AI 辅助工作中的元认知监控

Beyond "ChatGPT Can Make Mistakes": Designing Interventions to Support Metacognitive Monitoring in AI-Assisted Work

Manuel A. D. Santos, Paul Thiesse, Steeven Villa, Daniela Fernandes, Albrecht Schmidt, Verena Distler, Robin Welsch

arXiv 2609.17065首次发表:更新:

发表机构

Aalto University; LMU Munich(阿尔托大学; 慕尼黑大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过实验比较多种干预措施,发现可靠性卡片和对比回复能有效提升用户在 AI 辅助工作中的元认知监控能力,且监控与任务绩效可分离设计。

AI 中文摘要

AI 辅助对用户提出了元认知需求,用户必须判断自身能力以及系统的能力。然而,设计者缺乏关于选择哪些干预措施、将其置于何处以及如何判断其是否有效的比较性证据。我们从 11 位专家处征集了 30 项干预措施,并结合先前工作,将其组织成一个设计空间,涵盖时间(干预措施何时起作用)、层级(评判谁的能力)和来源(谁提供监控线索)三个维度。一项被试间实验(N = 917;12 个规划与组织问题)将每任务可靠性卡片、对比回复、暂停点和问题后反思与基线 LLM 助手进行了比较。可靠性卡片和对比回复降低了估计误差和过度自信,并提高了总体置信度区分度。未发现任务绩效提升或平均项目内区分度的增益。我们贡献了一个共享词汇表、一个设计空间,以及表明所测量的监控和任务绩效是可分离的设计目标的证据。

英文摘要

AI assistance places a metacognitive demand on users, who must judge their own competence and the system's. Yet designers lack comparative evidence on which interventions to choose, where to place them, and how to tell whether they worked. We elicited 30 interventions from 11 experts and, with prior work, organized them into a design space of time (when an intervention acts), level (whose competence is judged), and source (who supplies the monitoring cue). A between-subjects experiment (N = 917; 12 planning-and-organizing problems) compared a per-task reliability card, contrasting replies, pause points, and post-problem reflection against a baseline LLM assistant. Reliability cards and contrasting replies reduced estimation error and overconfidence and increased aggregate confidence discrimination. No task-performance improvement or average within-item discrimination gain was established. We contribute a shared vocabulary, a design space, and evidence that measured monitoring and task performance are separable design targets.

Comments40 pages, 13 figures, including appendices

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑