发表机构
Oxford Centre for Impact Research(牛津影响研究中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过随机在线实验发现,在AI生成前加入结构化元认知引导(Cognistance)相比单次提示,可将企业决策文档质量提升32%,并显著减少离题输出,但需额外时间成本。
AI 中文摘要
生成式人工智能加速并大多改善了专业工作,但存在一种担忧:将答案的生产和评估都委托给AI的用户可能接受较弱的输出,并较少参与底层推理(认知放弃)。迄今为止提出的干预措施,如无辅助练习或减缓采用,均在工作任务之外。我们测试了一种不同的方法:一个交互式元认知脚手架层(Cognistance,牛津影响研究中心(OCIR)开发的原型,在AI生成交付物之前要求用户澄清背景、选择战略方向并解释其推理)。使用脚手架后,平均综合质量提高了32%,且每位评分者的方向一致。收益在权衡阐述和战略一致性方面最大,而在技术特异性方面没有收益。总体效应的很大一部分反映在弱提示的挽救上:离题(任务外)交付物从34%降至5%。在那些自己的提示已陈述数据本地化问题的参与者中,优势为21%。治疗组参与者报告了更高的参与度,平均多花费约2.4分钟(10.46分钟)。即时回忆得分更高,这初步表明更好的保留,但在这一有限实验中,对敏感性分析不稳健。自我质量评分在两种条件下均未跟踪评分质量。生成前的结构化引导以适度的时间成本提高了AI辅助战略文档的评分质量和任务相关性。延迟保留、错误检测以及在真实组织中的效果是下一阶段研究的优先事项。
英文摘要
Generative AI speeds, and mostly improves, professional work, but there is concern that users who delegate both the production and the evaluation of an answer may accept weak output and engage less with the underlying reasoning (cognitive surrender). Interventions proposed so far, such as unassisted practice or slowing adoption, sit outside the working task. We tested a different approach: an interactive metacognitive scaffolding layer (Cognistance, a prototype developed at the Oxford Centre for Impact Research (OCIR) that asks users to clarify context, choose a strategic direction and explain their reasoning before the AI generates a deliverable). Mean composite quality was 32% higher with the scaffold, with the same direction for every rater. Gains were largest for trade-off articulation and strategic coherence and absent for technical specificity. A large part of the aggregate effect reflected rescue of weak prompts: floor-scored (off-task) deliverables fell from 34% to 5%. Among participants whose own prompt already stated the data-localisation problem, the advantage was 21%. Treatment participants reported greater involvement and took about 2.4 minutes longer on average (10.46 minutes). Immediate recall scores were higher, which tentatively suggests better retention, but in this limited experiment, was not robust to sensitivity analyses. Self-ratings of quality did not track rated quality in either condition. Structured elicitation before generation improved the rated quality and task relevance of AI-assisted strategy documents at modest cost in time. Delayed retention, error detection and effects in live organisations are the priorities for the next stage of research.
Comments12 pages