arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI评估应与人类协同开展

AI Evaluation Should Work With Humans

Jan Kulveit, Gavin Leech, Tomáš Gavenčiak, Raymond Douglas

arXiv 2608.13577首次发表:更新:

发表机构

Charles University; University of Cambridge(查理大学; 剑桥大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该论文提出AI评估应从聚焦超人类自主性能转向评估人机团队性能,以培育作为人类能力补充的AI系统,获得更优社会成果。

AI 中文摘要

本立场论文指出,当前主流的AI评估范式聚焦于超人类自主性能,隐含目标是替代人类,这正引导AI发展走向错误方向。相反,AI界应转向评估人机团队的性能。我们认为,这种协作转向将培育出真正作为人类能力补充的AI系统,从而比现有流程带来更好的社会成果。

英文摘要

This position paper argues that the dominant paradigm of AI evaluation (which focuses on superhuman autonomous performance and so implicitly targets the goal of replacing humans) is guiding AI development in the wrong direction. Instead, the AI community should pivot to evaluating the performance of human--AI teams. We argue that this collaborative shift will foster AI systems that act as true complements to human capabilities and therefore lead to far better societal outcomes than will the current process.

CommentsAccepted to ICML 2026 Position Paper Track

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑