人工身份:智能体AI中的驱动力与持久对齐
Artificial Id: Drive and Persistent Alignment in Agentic AI
浏览论文内容
中文总结 AI 辅助
针对智能体AI从有界任务转向持续自适应系统带来的控制问题,本文提出人工身份作为自适应内部驱动力,通过最小虚拟培养皿实验证明其能在无任务特定目标时涌现有用控制,同时指出持久性可能导致错位持续,需建立持久对齐边界。
中文摘要 AI 辅助
智能体AI正从有界任务执行转向能够保留重要状态、持续运行并跨任务边界适应的系统。这一转变带来了一个控制问题,当前的管控机制在很大程度上通过人工方式解决:目标、重试、验证、停止规则及其他行为转换均在外部指定。我们提出了一种人工身份,即一种自适应的内部驱动力,用于判断行为应继续、停止还是改变。在一个最小化的虚拟培养皿实验中,一个过小而无法进行通用推理且未接收任何任务特定行为目标的控制器,通过差异持久性发展出了有用的控制能力。当某种非预期的物理策略具有更好的持久性时,同一机制会选择该策略,并在其环境含义发生变化时,随后替换已习得的传感器映射。这些结果表明,自适应方向可以在未被明确指定为行为目标的情况下涌现。使这种自适应智能体有用的同一持久性,也可能导致错位、损坏状态和非预期行为跨任务边界持续存在。一个可扩展的人工身份将携带重要状态和自适应驱动力跨越这些边界,使对齐成为持续智能体系统的属性,而非模型响应或单条轨迹的属性。此类系统需要在可信观测、后果通道、持久状态、权威、身份、来源和硬约束之上建立持久的对齐边界。
英文摘要
Agentic AI is moving from bounded task execution toward systems that retain consequential state, continue operating and adapt across task boundaries. That shift creates a control problem that current harnesses largely solve by hand: objectives, retries, verification, stopping rules and other behavioral transitions are specified externally. We propose an artificial id, an adaptive internal drive for determining whether behavior should continue, stop or change. In a minimal virtual Petri-dish experiment, a controller too small to perform general-purpose reasoning and receiving no task-specific behavioral objective develops useful control through differential persistence. The same mechanism selects an unintended physical strategy when that behavior persists better and later replaces a learned sensor mapping when its environmental meaning changes. These results show that adaptive direction can emerge without being explicitly specified as a behavioral objective. The same persistence that makes such adaptive agency useful can also allow misalignment, corrupted state and unintended behavior to persist across task boundaries. A scalable artificial id would carry consequential state and adaptive drive across those boundaries, making alignment a property of the continuing agentic system rather than of a model response or single trajectory. Such systems require a persistent alignment boundary over trusted observations, consequence channels, persistent state, authority, identity, provenance and hard constraints.