发表机构
Salesforce(Salesforce)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
EDGE基于AgentGraph等框架,通过图遍历生成对话路径评估集,定义新指标量化智能体确定性,证实带受控节点转换的智能体确定性更优。
AI 中文摘要
随着智能体系统发展为复杂的多智能体编排工作流,对可衡量智能体行为一致性与确定性的系统化框架的需求日益迫切。本文提出一种基于AgentGraph的形式化评估方法,AgentGraph是一种规划器,通过领域特定语言(DSL)以可动态调整的有向图表示智能体推理。我们利用该结构形式主义,采用图遍历算法穷举枚举对话路径,形成覆盖智能体完整行为空间的综合评估集;随后系统重放这些可复现轨迹,对比观测到的输出与状态转换和预期DSL规范的差异。为量化可靠性,我们定义了新指标,用于测量精确重放及其语言变体的响应与轨迹确定性、结构依从性和语义一致性。我们的系统结果表明,使用AgentGraph、LangGraph等框架配置且具有显式结构节点转换的智能体,相较于未配置受控转换的智能体,表现出更优的确定性。
英文摘要
As agentic systems evolve into complex multi agent orchestration workflows, there is a growing and critical need for systematic frameworks that measures an agent's behavioral consistency and determinism. In this paper, we introduce a formal evaluation methodology that is grounded in AgentGraph, a planner powered by a domain specific language that represents agent reasoning through a dynamically adjustable directed graph. We leverage this structural formalism and utilize graph traversal algorithms that exhaustively enumerate conversational paths, forming a comprehensive evaluation set that captures the agent's complete behavioral space. We then systematically replay these reproducible trajectories to compare observed outputs and state transitions against the intended DSL specification. To quantify reliability, we define novel metrics that measure response and trajectory determinism, structural adherence and semantic consistency across both exact replays and their linguistic variants. Our system's results demonstrate that agents configured using frameworks like AgentGraph and LangGraph with explicitly structured node transitions show superior determinism over agents that are not configured with controlled transitions.