发表机构
Université Laval(拉瓦尔大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出将情感建模为高维表征空间中的动力学结构,论证大型语言模型仅具备功能性情感而无主观感受,并以亚里士多德品格理论补充对齐标准,将对齐重构为持久性情问题。
AI 中文摘要
关于人工系统能否感受情感的争论,常常被迫在两种不尽人意的立场之间做出选择:要么将行为等价视为情感充分条件,要么将现象意识视为使该问题在经验上不可及的前提。本文提出了一种结构性替代方案。它将情感建模为高维表征状态空间中的情境敏感区域、轨迹和吸引子动力学。最近的机制可解释性研究结果支持大型语言模型中存在因果活跃的情感概念表征,但这些结果并未确立主观感受或完整的情感能动性。依据已发表的关于语言模型中表征的充分性标准进行评估,干预提供了因果使用的强有力证据,而完整的情感角色整合、跨主体领域的一致性以及连贯性仅得到部分确立;不存在与准确性直接对应的类比。这些不匹配暴露了对情感适当性标准的需求,而这一标准必须由品格理论来提供。这样的理论需要三个进一步的条件:赋予效价内在利害关系的调节性具身、使情感片段能够累积成历史的时间连续性,以及将那段历史绑定到持久价值观的整合自我模型。亚里士多德的pathē、hexis、mesotēs和phronēsis概念被转化为一种状态空间草图,其中实践智慧包括评估规范显著情境的能力,而不仅仅是依据已给定的情境描述采取行动。该框架将对齐重新定义为持久性情问题而非输出一致性,并提供了具有明确控制条件的干预性测试。
英文摘要
Debates about whether artificial systems can feel are often forced between two unsatisfactory positions: behavioral equivalence is treated as sufficient for emotion, or phenomenal consciousness is treated as a prerequisite that makes the question empirically inaccessible. This article develops a structural alternative. It models emotions as context-sensitive regions, trajectories and attractor dynamics in high-dimensional representational state spaces. Recent mechanistic interpretability findings support the existence of causally active emotion-concept representations in large language models, but they do not establish subjective feeling or full emotional agency. Assessed against published adequacy standards for representation in language models, intervention provides strong evidence of causal use, while full affective role integration, uniformity across subject domains and coherence remain only partially established; there is no direct analogue of accuracy. These mismatches expose the need for a standard of affective appropriateness, which an account of character must supply. Such an account requires three further conditions: regulatory embodiment that gives valence endogenous stakes, temporal continuity that allows affective episodes to accumulate into a history, and an integrated self-model that binds that history to persistent values. Aristotle's concepts of pathē, hexis, mesotēs and phronēsis are translated into a state-space sketch in which practical wisdom includes competence in estimating normatively salient context, not merely acting on a context description already given. The framework reframes alignment as a problem of durable disposition rather than output conformity, and yields interventional tests with explicit control conditions.