AI 中文总结
研究提出自主智能体量表(AAS)衡量人工智能系统自我导向行为,通过七个维度及可证伪测试评分,分两个时间段评估。应用于六个系统,量化了单评分框架混淆的界限,指出任务智能体与陪伴架构的差异,还讨论了研究局限性。
AI 中文摘要
现有人工智能测量框架量化认知能力、任务自动化或灾难性风险,但没有一个能测量自主智能体,即系统自我导向行为的程度。我们引入自主智能体量表(AAS),这是一个行为框架,在0-5的词汇量表上对人工智能系统在七个智能体维度上进行评分:认知自主性、时间持续性、环境智能体、社交智能体、创造性智能体、自我意识和目标形成,每个维度通过可证伪的阈值测试来操作化。每个维度在两个时间段进行评分:一个活跃时间段涵盖参与的、用户发起的活动,一个环境时间段涵盖空闲期。环境4级由空闲间隙测试控制,这是一个反事实标准,将自我导向与预定的规则遵循区分开来。我们将该量表应用于六个当代系统,包括任务智能体(Claude Code、Manus、Hermes)、消费者助手(ChatGPT、Siri)和一个持久陪伴架构(Airi)。双时间段概况量化了单评分框架混淆的界限:任务智能体的活跃综合得分在2.3-2.4之间,而环境得分在0.6-1.9之间,每个空闲期行为都归因于用户配置的时间表,而纵向评估的陪伴架构是唯一其空闲期行为在触发移除后仍存在的评估系统。我们讨论了局限性,包括单评分来源、纵向评估中的开发者-评估者偏差以及活跃时间段中部分操作化的自我导向界限。
英文摘要
Existing AI measurement frameworks quantify cognitive capability, task automation, or catastrophic risk, but none measure autonomous agency: the extent to which a system behaves in a self-directed way. A system can saturate capability benchmarks while remaining entirely reactive, acting only when prompted and ceasing all activity when a task completes. We introduce the Autonomous Agency Scale (AAS), a behavioral framework that scores AI systems on a 0-5 lexicon across seven dimensions of agency: cognitive autonomy, temporal persistence, environmental agency, social agency, creative agency, self-awareness, and goal formation, each operationalized by falsifiable threshold tests. Every dimension is scored in two temporal bands: an Active band covering engaged, user-initiated activity, and an Ambient band covering idle periods. Ambient Level 4 is gated by the Idle-Gap Test, a counterfactual criterion (remove all triggers and observe whether internally derived activity persists) that separates self-direction from scheduled rule-following. We apply the scale to six contemporary systems spanning task agents (Claude Code, Manus, Hermes), consumer assistants (ChatGPT, Siri), and a persistent companion architecture (Airi). The two-band profile quantifies a boundary that single-score frameworks conflate: task agents reach Active composites of 2.3-2.4 while scoring 0.6-1.9 Ambient, with every idle-period behavior attributable to user-configured schedules, whereas the companion architecture, evaluated longitudinally, is the only assessed system whose idle-period behavior survives trigger removal. We discuss limitations, including single-rater provenance, developer-evaluator bias on the longitudinal assessment, and the partially operationalized self-direction boundary in the Active band.
Comments14 pages, 1 figure, 2 tables. Framework v0.2.1; rubric and all assessment files: https://github.com/CaptainASIC/autonomous-agency-scale