更快的人工智能,参差不齐的前沿:快速跨越、锯齿状前沿与人类判断力的重新定位
Faster AI, Uneven Frontier: Rapid Crossings, a Jagged Frontier, and the Repositioning of Human Judgment
AI总结:
研究人工智能在认知任务上快速跨越人类基线但前沿参差不齐的现象,指出人类在一些方面仍具优势,基准结果有偏差,人类将系统用作认知扩展,人机协作需重新定位人类角色,阐述了相应状况及意义。
AI中文摘要:
在2023年至2026年期间,前沿人工智能系统在越来越多有限的、明确规定的、可评估的认知任务上超越了有记录的人类专家基线,包括研究生水平的科学问题、竞赛数学、软件工程基准和结构化诊断推理,且此类系统能以50%可靠性完成的任务长度大约每七个月就会翻倍。这些跨越迅速且广泛,但前沿参差不齐:人类在长期可靠性、真正新颖的问题、校准的自我认知、样本高效学习和具身行动方面仍具有决定性优势,基准结果夸大了部署能力。同时,人类越来越将这些系统用作认知扩展。卸载文献预测了对独立技能的成本影响,早期实地证据与之相符,但关于先前技术的最大元分析证据则相反,生成式人工智能是否不同尚无定论。最后,关于人机协作的实验记录表明,简单组合往往不如较强的一方,这意味着人类的贡献必须重新定位到规范、验证和监督上,这种转变在实验中可见,但在实地劳动力市场数据中目前几乎不可见。本文阐述了由此产生的状况,即在参差不齐的前沿上快速跨越,人类角色必须重新设计而非捍卫,并阐述了其理论和实践意义。
英文摘要:
Between 2023 and 2026, frontier AI systems crossed documented human expert baselines on a growing set of bounded, well-specified, evaluable cognitive tasks, including graduate-level science questions, competition mathematics, software-engineering benchmarks, and structured diagnostic reasoning, while the length of tasks such systems can complete at 50% reliability doubled roughly every seven months. These crossings are rapid and broad, but the frontier is jagged: humans retain decisive advantages in long-horizon reliability, genuinely novel problems, calibrated self-knowledge, sample-efficient learning, and embodied action, and benchmark results overstate deployed capability for reasons that are themselves now documented, namely contamination, construct validity, vendor self-evaluation, and the gap between 50% reliability and the reliability that economic work requires. Concurrently, humans increasingly use these systems as cognitive extensions. The offloading literature predicts costs to unaided skill, and early field evidence is consistent with such costs, though the largest meta-analytic evidence on prior technologies points the other way, and the question of whether generative AI differs is open. Finally, the experimental record on human-AI collaboration shows that naive combination often underperforms the stronger partner, implying that the human contribution must be repositioned toward specification, verification, and oversight, a shift visible in experiments but, so far, barely visible in field labor-market data. This paper states the resulting position, rapid crossings on a jagged frontier with a human role that must be redesigned rather than defended, and draws out its theoretical and practical implications.