AI 中文总结
探讨机器学习中表征概念,随着模型发展其从工程领域转向心理表征领域。评估《柏拉图式表征假说》中关于不同AI模型表征属性趋同由现实统一结构驱动的说法,借助心灵哲学观点审视,阐明关键、解释不足并给出研究方向。
AI 中文摘要
表征是现代机器学习的核心概念,通常指支持学习与泛化的内部编码。随着模型规模扩大及能力趋近人类水平,这种表征语言有时会从工程背景转向心理表征这一哲学意味更浓的领域。我们认为近期关于不同人工智能模型表征属性趋同的说法就是如此。特别是,我们评估了《柏拉图式表征假说》中的论点,该假说认为这种趋同由现实的统一结构驱动。我们通过引入心灵哲学中关于心理表征辩论的论点和观点来审视这一说法。我们认为这些哲学资源能阐明此类说法的关键所在,解释为何仅对齐证据不足以得出强有力的形而上学结论,并为未来研究指明方向。
英文摘要
Representation is a central concept in modern machine learning, where it usually refers to internal encodings that support learning and generalization. As models scale and their capabilities become increasingly human-level, this representational language sometimes shifts from an engineering context into the more philosophically loaded domain of mental representation. We argue that this is the case for recent claims about the convergence of representational properties across different AI models. In particular, we assess the arguments developed in The Platonic Representation Hypothesis, according to which this convergence is driven by a unified structure of reality. We examine this claim by introducing arguments and ideas from debates about mental representation in the philosophy of mind. We argue that these philosophical resources can clarify what is at stake in such claims, explain why alignment evidence alone is insufficient for strong metaphysical conclusions, and suggest directions for future research.
Comments8 pages. Accepted for oral presentation at ICML 2026