AI 中文总结
本研究通过全文审计发现人机协作研究中声明分母漂移问题,即相同标签掩盖不同人类单位,并提出检查点以改进证据综合。
AI 中文摘要
人机协作(HAT)综述通常按顾问、队友或协调者等标签对研究进行分组。然而,相同的标签可能描述一个人采纳AI建议、几个人围绕AI进行协调,或一种分配权威与责任的工作流程。因此,将这些研究合并可能会改变声明背后的人类单位。我们考察全文证据如何改变声明背后的研究集合。我们从419条标题/摘要图谱中有目的地选取了86篇全文进行审计。我们发现,全文阅读改变了40条记录的核心成员资格:74个明显的核心候选中有36个移出,而12个边界候选中有4个移入。团队词汇并不能可靠地识别社会单位:27个人机二元组中有14个,23个多人同伴团队中有20个使用了团队或协作术语。86篇论文中只有20篇明确说明了谁可以看到AI输出。四次盲法语言模型运行一致标记了53个筛选案例和59个安排,但这些共识决策中有32%和34%与全文标签不同。这些结果将声明分母漂移确定为HAT研究中的一个综合问题。我们贡献了一个以人类安排为中心的全文审计,以及一个声明合并检查点,用于决定何时可以比较关于信任、协调、绩效、效率和问责制的证据。
英文摘要
Human-AI Teaming (HAT) reviews often group studies by labels such as advisor, teammate, or coordinator. Yet the same label can describe one person taking AI advice, several people coordinating around AI, or a workflow that distributes authority and responsibility. Pooling these studies can therefore change the human unit behind a claim. We examine how full-text evidence changes the set of studies behind a claim. We audited 86 full texts purposively selected from a 419-record title/abstract map. We find that full-text reading changed core membership for 40 records: 36 of 74 apparent core candidates moved out, while 4 of 12 boundary candidates moved in. Team vocabulary did not reliably identify the social unit: 14 of 27 human-AI dyads and 20 of 23 multi-human peer teams used team or collaboration terms. Only 20 of 86 papers specified who could see AI output. Four blinded language-model runs unanimously labeled 53 screening cases and 59 arrangements, yet 32% and 34% of those consensus decisions differed from the full-text labels. These results identify claim-denominator drift as a synthesis problem in HAT research. We contribute a full-text audit centered on human arrangements and a claim-pooling checkpoint for deciding when evidence about trust, coordination, performance, efficiency, and accountability can be compared.
Comments22 pages, 4 figures, 6 tables. Preprint