发表机构
Toulouse School of Economics; Carnegie Mellon University(图卢兹经济学院; 卡内基梅隆大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出贝叶斯智能理论,证明智能体行为可解释为贝叶斯更新当且仅当报告非完全矛盾,并刻画智能顺序及聚合粗略报告的困难。
AI 中文摘要
从可观察行为推断智能是人工智能中的一个基础性挑战。我们为语言模型等智能体发展了一套贝叶斯智能理论。每个提示都会引发一个可能不完美的内部实验;智能体通过贝叶斯规则更新一个全支撑先验,并忠实地报告其对问题可能答案的后验。重复实验从同一未观察实验的固定状态中抽取新的、独立的结果。我们证明,当且仅当智能体的报告并非完全矛盾(即,在所有提示的每个报告下,某些状态仍然可能)时,智能体的行为才承认这种解释。报告频率和正概率的大小不施加进一步限制。我们进一步提出并刻画了一种智能顺序的行为含义,该顺序使得两个智能体的行为与一个智能体拥有更信息丰富的实验相一致:应存在报告分布的耦合,使得信息更丰富的智能体的报告排除其对应方排除的每个答案。最后,我们展示了聚合来自智能体的粗略报告的困难:除非智能体报告关于世界完整状态的信念,否则最优聚合可以对未被排除的状态分配任意权重。这些结果为理解智能体行为何时是智能的提供了基础,并强调了拒绝贝叶斯理性的困难。
英文摘要
Inferring intelligence from observable behavior is a foundational challenge in artificial intelligence. We develop a theory of Bayesian intelligence for agents such as language models. Each prompt induces a possibly imperfect internal experiment; the agent updates a full-support prior by Bayes' rule and faithfully reports its posterior over the possible answers to the question. Repetitions draw fresh, independent outcomes from the same unobserved experiment at one fixed state. We show that the agent's behavior admits this explanation if and only if its reports are not fully contradictory, i.e., some state remains possible under every report across all prompts. Report frequencies and the sizes of positive probabilities impose no further restrictions. We further propose and characterize the behavioral implications of an intelligence order that makes the behavior of two agents consistent with one agent having access to a more informative experiment: there should exist a coupling of report distributions such that the more informative agent's report excludes every answer excluded by its counterpart. Finally, we show the difficulty of aggregating coarse reports from intelligent agents: unless the agent reports a belief about the complete state of the world, the optimal aggregation can assign arbitrary weights to states that have not been excluded. These results provide a basis for understanding when agents' behavior is intelligent and highlight the difficulty of rejecting Bayesian rationality.