生成式人工智能悖论:“它能创造之物,未必能理解”
The Generative AI Paradox: "What It Can Create, It May Not Understand"
- University of Washington(华盛顿大学)
- Allen Institute for Artificial Intelligence(艾伦人工智能研究所)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出并验证生成式人工智能悖论假说:生成模型虽能超越人类生成专家级输出,但其生成能力并不依赖于理解能力,在语言和图像实验中理解表现仍弱于人类且更脆弱。
AI中文摘要:
近期生成式人工智能的浪潮引发了前所未有的全球关注,人们对可能达到超人水平的人工智能既感到兴奋又充满担忧:如今模型只需数秒就能产出足以挑战甚至超越人类专家能力的输出。与此同时,模型在理解方面仍会犯下一些基本错误,而这些错误即使是非专家人类也不应出现。这向我们呈现了一个明显的悖论:我们该如何调和看似超人的能力与少数人类才会犯的错误持续存在之间的矛盾?在这项工作中,我们提出,这种张力反映了当今生成式模型中智能的配置方式相对于人类智能的配置方式存在分歧。具体而言,我们提出并检验了生成式人工智能悖论假说:生成模型因被直接训练来复现专家级输出,从而获得了不依赖于——因此可以超越——其对同类输出的理解能力的生成能力。这与人类形成对比,对人类而言,基本理解几乎总是先于生成专家级输出的能力。我们通过受控实验检验这一假说,分析生成模型在语言和图像两种模态下的生成与理解。我们的结果表明,尽管模型在生成方面可以超越人类,但在理解能力的测量上始终不及人类,同时生成与理解表现之间的相关性更弱,并且对对抗性输入更加脆弱。我们的发现支持了以下假说:模型的生成能力可能并不依赖于理解能力,并呼吁在通过与人类智能类比来解释人工智能时保持谨慎。
英文摘要:
The recent wave of generative AI has sparked unprecedented global attention, with both excitement and concern over potentially superhuman levels of artificial intelligence: models now take only seconds to produce outputs that would challenge or exceed the capabilities even of expert humans. At the same time, models still show basic errors in understanding that would not be expected even in non-expert humans. This presents us with an apparent paradox: how do we reconcile seemingly superhuman capabilities with the persistence of errors that few humans would make? In this work, we posit that this tension reflects a divergence in the configuration of intelligence in today's generative models relative to intelligence in humans. Specifically, we propose and test the Generative AI Paradox hypothesis: generative models, having been trained directly to reproduce expert-like outputs, acquire generative capabilities that are not contingent upon -- and can therefore exceed -- their ability to understand those same types of outputs. This contrasts with humans, for whom basic understanding almost always precedes the ability to generate expert-level outputs. We test this hypothesis through controlled experiments analyzing generation vs. understanding in generative models, across both language and image modalities. Our results show that although models can outperform humans in generation, they consistently fall short of human capabilities in measures of understanding, as well as weaker correlation between generation and understanding performance, and more brittleness to adversarial inputs. Our findings support the hypothesis that models' generative capability may not be contingent upon understanding capability, and call for caution in interpreting artificial intelligence by analogy to human intelligence.