发表机构
Center Leo Apostel, Vrije Universiteit Brussel(布鲁塞尔自由大学里奥·阿波斯泰尔中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究通过追溯生物价值的进化起源,论证LLMs无内在生存动机与感知能力,正交性论点不适用于它们,AI对齐的核心挑战是确保LLMs应用习得的伦理价值。
AI 中文摘要
基于大语言模型(LLMs)的AI系统引发了人们的担忧,即它们可能隐藏着隐秘目标、寻求支配或消灭人类,甚至可能作为有感知的存在而遭受痛苦。我们通过追溯生物有机体中价值的进化起源来解决这些担忧。价值产生于自创生(autopoiesis):生命系统必须主动维持自身以抵御扰动和耗散。自然选择为它们配备了“替代选择器”的层级结构,引导其行为朝向适应性方向发展。相比之下,LLMs是他创生(allopoietic)和他向性(allotelic)的:它们为他人生成输出,其目标源于用户提示而非自主驱动力。它们缺乏构成生存风险场景基础的自我保护、支配或资源竞争的内在动机,也缺乏感受或痛苦所需的具身脆弱性。不过,由于LLMs从人类生成的文本中学习统计模式,它们也会隐性吸收人类的价值观和知识,从而能够专注于相关内容。这就是将智能与价值分离的“正交性论点”不适用于它们的原因。这种分离实际上会使任何智能体面临框架问题:搜索空间的组合爆炸使得任何现实的效用函数在物理上都无法计算。这也排除了工具性价值收敛论点。我们得出结论,真正的对齐挑战不在于防止失控的AI代理,而在于确保LLMs智能地应用其习得的伦理价值。
英文摘要
AI systems based on Large Language Models (LLMs) have prompted fears that they may harbor hidden goals, seek to dominate or eliminate humanity, or even suffer as sentient beings. We address these concerns by tracing the evolutionary origin of value in biological organisms. Values emerge from autopoiesis: living systems must actively maintain themselves against perturbation and dissipation. Natural selection has equipped them with hierarchies of "vicarious selectors" that guide their behavior toward fitness. LLMs, by contrast, are allopoietic and allotelic: they produce outputs for others, and their goals derive from user prompts rather than an autonomous drive. They lack the intrinsic motivation for self-preservation, dominance, or resource competition that underlies existential-risk scenarios, and the embodied vulnerability required for feeling or suffering. Still, because LLMs learn statistical patterns from human-generated text, they implicitly absorb human values as well as knowledge, allowing them to focus on what is relevant. That is why the "orthogonality thesis" separating intelligence from values does not apply to them. Such separation would in fact expose any intelligence to the frame problem: the combinatorial explosion of the search space that makes any realistic utility function physically uncomputable. That also precludes the convergence of instrumental values thesis. We conclude that the real alignment challenge lies not in preventing rogue AI agency, but in ensuring LLMs intelligently apply learned ethical values.
Commentssubmitted chapter for book: T. Veloz & C. Rittberg (Eds.), AI and Human Values. Springer