发表机构
Cheung Kong Graduate School of Business(长江商学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究人工智能代理在复杂系统中的行为评估问题,提出层归因诊断框架,该框架含基础计算层和行为调制层,阐明了替代有效性等三个结果,为评估代理行为提供了新方法。
AI 中文摘要
人工智能代理越来越多地在行为科学家所研究的临床、政治、科学和社会系统中运作。评估这些系统需要源级别诊断,因为相同的行为模式可能源于代理的表征基础,也可能源于塑造其表现的角色、目标、交互结构和治理规则。本文提出了一个人工智能代理行为的诊断框架:层归因。基础计算层通过架构、内存、感知、注意力和表征定义了哪些行为是可能的。行为调制层通过身份、资源、目标、社会交互、制度约束和治理来塑造这些能力的表达方式。该框架阐明了三个结果:替代有效性是模型 - 任务 - 层关系,人机差异提供诊断证据,治理需要在干预前进行源归因。因此,将人工智能代理视为行为主体需要评估方法,在决定如何解释、验证或治理行为之前,先确定行为的起源。
英文摘要
AI agents increasingly act within the same clinical, political, scientific, and social systems that behavioral scientists study. Evaluating these systems requires source-level diagnosis: the same behavioral pattern may arise from an agent representational substrate or from the roles, objectives, interaction structures, and governance rules that shape its expression. This Perspective proposes a diagnostic framework for AI agent behavior: layer attribution. The foundational computational layer defines what behaviors are possible through architecture, memory, perception, attention, and representation. The behavioral modulation layer shapes how those capacities are expressed through identity, resources, objectives, social interaction, institutional constraints, and governance. The framework clarifies three consequences: surrogate validity is a model-task-layer relation, human-AI divergence provides diagnostic evidence, and governance requires source attribution before intervention. Treating AI agents as behavioral actors therefore requires evaluation methods that determine where behavior originates before deciding how to explain, validate, or govern it.