arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.05527cs.AIcs.HC

超越“AI 帮助人类”:智能体时代人机团队中面向决策的评估设计

Beyond "AI Helps Humans": Decision-Targeted Evaluation Design for Human-Agent Teams in the Agentic Era

Hamed Khosravi, Xiaoming Huo

首次发表
浏览论文内容

中文总结 AI 辅助

针对人机团队部署决策,提出TEAM-Design规则,在固定重放预算下优化任务重放分配,以准确判断人机工作流是否优于仅人类或仅智能体方案,并控制错误决策风险。

中文摘要 AI 辅助

无论编码智能体在工程师监督下工作,还是临床模型辅助放射科医生,部署问题都在于是否保留人机工作流,还是仅由人类或仅由智能体单独执行。只有当人机工作流优于这两种替代方案时,才值得保留。然而,一旦部署,这两种替代结果都无法观察到:恢复其中一种意味着在该替代方案下重放任务,而每次重放都会消耗专家时间或计算资源。在固定的重放预算下,设计问题因此变成:哪些任务应更可能接受仅人类重放,哪些应接受仅智能体重放。现有方法并未直接针对这一决策。智能体基准测试不选择要测量的缺失基线,基于方差的采样忽略了两个比较中哪一个更接近失败,而贝叶斯信息方法专注于学习模型参数而非做出部署决策。我们提出 TEAM-Design,一种为每个任务分配两个重放概率(每个基线一个)的规则。它在缺失基线结果难以根据任务已知信息预测且该比较更难建立时提高概率,在重放成本高昂时降低概率。我们证明该规则解决了这一预算设计问题,并且从记录的概率中随机抽取重放仍能控制错误宣称工作流优于两者的概率。我们重新分析了 6 个临床场景(其中没有一个人机工作流优于两种替代方案)和一个编码基准(其中有一个),然后在合成设计以及基于真实胸部 X 光片阅读器研究构建的半合成设计上评估 TEAM-Design。当两个比较中有一个明显比另一个更难确定时,TEAM-Design 效果最佳;当两者难度相似时,其表现可能不如基于方差的分配。

英文摘要

Wherever a coding agent works under engineer supervision, or a clinical model assists a radiologist, the deployment question is whether to keep the human-AI workflow or replace it with the human alone or the agent alone. The human-AI workflow is worth keeping only if it beats both of those alternatives. Yet once it is deployed, neither alternative outcome is observed: recovering one means replaying the task under that alternative, and every replay costs expert time or compute. Under a fixed replay budget, the design question is therefore which tasks should be more likely to receive a human-only replay, and which an agent-only replay. Existing methods do not directly target this decision. Agent benchmarks do not choose which missing baseline to measure, variance-based sampling ignores which of the two comparisons is closer to failing, and Bayesian information methods focus on learning model parameters instead of making the deployment decision. We propose TEAM-Design, a rule that gives every task two replay probabilities, one per baseline. It raises a probability where the missing baseline outcome is hard to predict from what is already known about the task and where that comparison is harder to establish, and lowers it where replay is expensive. We prove that the rule solves this budgeted design problem, and that drawing the replays at random from recorded probabilities still controls the chance of wrongly declaring that the workflow beats both. We reanalyze 6 clinical settings, where no human-AI workflow beats both alternatives, and a coding benchmark, where one does, then evaluate TEAM-Design on synthetic designs and on a semi-synthetic design built from a real chest X-ray reader study. TEAM-Design works best when one of the two comparisons is clearly harder to settle than the other, and can do worse than variance-based allocation when the two are similarly difficult.

发表机构

  • Georgia Institute of Technology(佐治亚理工学院)

机构由 AI 辅助整理,请以论文原文为准。

↑