定义AI智能体:标准、指标与基准的汇编
Defining AI Agents: A Compendium of Criteria, Metrics, and Benchmarks
浏览论文内容
中文总结 AI 辅助
针对AI智能体缺乏统一定义的问题,本文围绕五个维度开展综述,整合评估指标与基准,并推出“智能体汇编”资源,以促进可复现研究和系统比较。
中文摘要 AI 辅助
人工智能中的“智能体”一词缺乏标准定义,这使AI智能体研究的评估、比较和可复现性变得复杂。我们通过一项围绕智能体性的五个维度组织的调查来解决这一模糊性:环境交互、学习与适应、自主性、目标导向行为和时间连贯性。对于每个维度,我们考察了先前工作中该底层能力是如何被概念化的,并综合了用于评估该能力的指标、基准和评估框架。本综述提供了当前智能体评估格局的结构化描述,既强调了已确立的方法,也指出了评估仍然有限或不一致的领域。我们还引入了“智能体汇编”(Agent Compendium),这是一个面向公众的数字资源,用于组织和扩展本综述中确定的评估方法。调查与汇编共同为跨AI系统评估和比较智能体能力提供了一个通用结构,支持更具可复现性的研究、更清晰的沟通以及对人工智能体更系统的研究。
英文摘要
The term agent in artificial intelligence lacks a standard definition, complicating the evaluation, comparison, and reproducibility of AI agent research. We address this ambiguity through a survey organized around five dimensions of agenticness: environmental interaction, learning and adaptation, autonomy, goal-directed behavior, and temporal coherence. For each dimension, we examine how the underlying capability has been conceptualized across prior work and synthesize the metrics, benchmarks, and evaluation frameworks used to assess it. This review provides a structured account of the current landscape of agent evaluation, highlighting both established approaches and areas where evaluation remains limited or inconsistent. We additionally introduce the Agent Compendium, a public-facing digital resource that organizes and extends the evaluation methods identified through this review. Together, the survey and compendium provide a common structure for evaluating and comparing agent capabilities across AI systems, supporting more reproducible research, clearer communication, and more systematic study of artificial agents.
发表机构
- Duke University(杜克大学)
机构由 AI 辅助整理,请以论文原文为准。