交互就绪:构建和评估人类角色AI智能体的框架
Interaction Readiness: A Framework for Building and Evaluating AI Agents in Human Roles
浏览论文内容
中文总结 AI 辅助
该研究针对带角色属性AI智能体的评估缺口,提出交互就绪框架,通过分离内容与交互规范、四项智能体操作实现评估,证实内容与交互质量独立,并给出规范模板与审核流程。
中文摘要 AI 辅助
构建带角色属性的AI智能体的产品与工程团队面临评估缺口:智能体可产出准确、安全且流畅的内容,却仍无法满足其分配角色的行为要求。本文提出交互就绪(Interaction Readiness)作为针对该缺失性能层的规范与评估框架。该框架将内容规范(管控智能体知晓与表述的内容)与交互规范(定义智能体在角色主导交互中的行为方式)分离;交互规范要求团队在部署前明确角色目的、权限边界、常见场景、边界案例、修复行为及审核标准。我们通过四项智能体操作实现交互就绪:理解目的、校准权限、管理语气及修复故障。利用StudyChat(学生与AI辅导智能体交互的公开数据集),我们证实内容准确性与交互质量为独立维度:智能体可事实正确却无法胜任辅导角色,或交互合理却技术错误;最常见的失败是权限校准失误——智能体常知晓如何回答,却不知辅导角色是否允许、何时允许及如何回答。本文将这些发现转化为产品与工程团队可在部署前后应用的规范模板与审核流程。
英文摘要
Product and engineering teams building role-bearing AI agents face an evaluation gap: an agent can produce accurate, safe, and fluent content while still failing the behavioral requirements of its assigned role. This paper introduces Interaction Readiness as a framework for specifying and evaluating that missing layer of performance. The framework separates content specifications, which govern what an agent knows and says, from interaction specifications, which define how an agent should conduct itself in a role-governed exchange. Interaction specifications require teams to define role purpose, authority boundaries, recurring situations, boundary cases, repair behaviors, and audit criteria before deployment. We operationalize interaction readiness through four agent operations: understanding purpose, calibrating authority, managing tone, and repairing breakdowns. Using StudyChat, a public dataset of student interactions with an AI tutoring agent, we show that content accuracy and interaction quality are independent dimensions: an agent may be factually correct while failing as a tutor, or interactionally sound while technically wrong. The most persistent failure is authority miscalibration: the agent often knows how to answer, but not whether, when, or how the tutor role permits it to answer. The paper translates these findings into a specification template and audit procedures that product and engineering teams can apply before and after deployment
发表机构
- Columbia University(哥伦比亚大学)
机构由 AI 辅助整理,请以论文原文为准。