arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从任务成功到生产性成功:通过质量与成本评估人机协作

From Task Success to Productive Success: Evaluating Human-AI Collaboration by Quality and Cost

Saki Imai, Mert İnan, Malihe Alikhani

arXiv 2609.21117首次发表:更新:

发表机构

Northeastern University(东北大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出一个以生产力为导向的框架,通过质量与交互成本的比率评估人机协作,发现相同质量下成本差异可达70倍,并区分生产性成功与昂贵成功。

AI 中文摘要

AI生产力通常通过任务完成时间、经济价值或结果质量的改进来衡量。然而,这些衡量方式通常将协作视为一个黑箱,它们捕捉了产生了什么输出,但未捕捉产生该输出所需的交互成本。受经济学文献启发,我们引入了一个以生产力为导向的框架,将人机协作评估为相对于交互成本的结果质量。在跨越四个任务的两个数据集中,我们表明:(1)具有相同质量评级的会话在交互成本上可能相差高达70倍;(2)质量-成本关系因任务而异,有些任务奖励延长交互,而其他任务则偏好快速收敛;(3)主观用户评分不能可靠地替代生产力;(4)生产性会话的特点是智能体更早地进行探测,用户花费更少的精力修复交互。通过区分生产性成功与昂贵成功,我们的框架使交互成本可见,并展示了对话分析如何为AI系统的评估和设计提供信息。

英文摘要

AI productivity is often measured by task completion time, economic value, or improvements in outcome quality. However, these measures usually treat collaboration as a black box where they capture what output was produced, but not the interaction cost required to produce it. Motivated by economics literature, we introduce a productivity-oriented framework for evaluating human-AI collaboration as outcome quality relative to interaction cost. Across two datasets spanning four tasks, we show that: (1) sessions with identical quality ratings can differ by up to 70 times in interaction cost; (2) quality-cost relationships vary by task, with some tasks rewarding extended interaction and others favoring fast convergence; (3) subjective user ratings are not reliable substitutes for productivity; and (4) productive sessions are characterized by agents probing earlier and users spending less effort repairing the interaction. By distinguishing productive success from costly success, our framework makes interactional cost visible and shows how dialogue analysis can inform the evaluation and design of AI systems.

CommentsEMNLP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑