arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SPEAR:交互式人机对齐的五个原则

SPEAR: Five Principles for Interactive Human-Agent Alignment

Tao Long, Lydia B. Chilton

arXiv 2610.07204首次发表:更新:

发表机构

Columbia University(哥伦比亚大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出将人机对齐视为交互设计问题,并引入SPEAR五原则:规范、过程、评估、适应与重新校准,以指导智能体在长期社会情境中的持续对齐。

AI 中文摘要

近期的人工智能对齐工作通常将对齐定义为部署前的优化问题:收集人类反馈,学习偏好或原则,微调模型,然后部署对齐的系统。这种框架取得了重大进展,但它未能充分说明当AI系统在情境化、长期和社交环境中代表用户作为智能体行动时会发生什么。本立场论文将人机对齐重新定义为持续的交互设计问题。我们提出SPEAR,即交互式对齐的五个支柱:规范(人们如何表达意图并建立共同理解)、过程(智能体如何决定何时行动、询问、弃权(不执行)或暂停)、评估(人们如何判断智能体是否成功)、适应(智能体如何在重复使用中适应用户)以及重新校准(人们如何根据智能体调整其信任、期望和行为)。

英文摘要

Recent AI alignment work often frames alignment as a pre-deployment optimization problem: collect human feedback, learn preferences or principles, finetune the model, and deploy an aligned system. This framing has produced major progress, but it under-specifies what happens once AI systems act as agents on users' behalf in situated, long-term, and social contexts. This position paper reframes human-agent alignment as an ongoing interaction design problem. We propose SPEAR, five pillars of interactive alignment: Specification (how people express intent and establish shared understanding), Process (how agents decide when to act, ask, defer, or pause), Evaluation (how people judge whether agents succeeded), Adaptation (how agents adapt to users over repeated use), and Recalibration (how people adapt their trust, expectations, and behavior in response to agents).

Comments3 pages. Best Talk Award at the ACM Conference on Human-AI Complementarity and Alignment (HCOMP 2026)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑