arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.04168cs.AI

智能体认知深度:评估LLM智能体的操作标准

Agentic Cognitive Depth: Operational Criteria for Evaluating LLM Agents

Nijesh Upreti, Chris Sypherd, Vaishak Belle

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出智能体认知深度概念,通过五个操作标准评估LLM智能体轨迹级性能,并提供扰动程序与基准扩展方案。

中文摘要 AI 辅助

智能体大语言模型(LLM)系统通常实现为一个LLM在循环中,包含规划、记忆、工具和控制流。这种以应用为中心的视角将智能体LLM研究连接到可部署系统,并留下了如何超越端到端任务成功来评估此类系统的问题。基于这一视角,我们定义了智能体认知深度,作为跨越五个操作标准的轨迹级剖面。该剖面包含上下文敏感性($C$)、时间连续性($T$)、多模态协调($M$)、自适应交互($A$)和元认知监控($Mc$)。前四个标准衡量控制流、记忆、工具和规划在轨迹中的使用情况。第五个标准衡量系统是否监控和调节整个运行过程。对于每个标准,我们给出操作代理和扰动程序,然后将剖面连接到智能体的世界模型。我们提供了扩展GAIA、SWE-bench、WebArena和TRIP-Bench等基准所需的结构,以增加每个标准的诊断。符号验证器、结构化记忆、规划器耦合和工具约束提供了构建和测试这些能力的实用方法。

英文摘要

Agentic large language model (LLM) systems are commonly implemented as an LLM in a loop with Planning, Memory, Tools, and Control Flow. This application-focused view connects agentic LLM research with deployable systems and leaves open how such systems should be evaluated beyond end-to-end task success. Building on this view, we define agentic cognitive depth as a trajectory-level profile across five operational criteria. The profile contains context sensitivity ($C$), temporal continuity ($T$), multimodal coordination ($M$), adaptive interaction ($A$), and metacognitive monitoring ($Mc$). The first four criteria measure how well Control Flow, Memory, Tools, and Planning are used across a trajectory. The fifth measures whether the system monitors and regulates the full run. For each criterion, we give operational proxies and a perturbation procedure, then connect the profile to the agent's world model. We provide the structure needed to extend benchmarks such as GAIA, SWE-bench, WebArena, and TRIP-Bench with per-criterion diagnostics. Symbolic verifiers, structured memory, planner coupling, and tool constraints provide practical ways to build and test these capacities.

发表机构

  • The University of Edinburgh(爱丁堡大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑