AI 中文总结
提出DynGraphAgentBench基准,用于在延迟反馈下评估动态图异常检测中智能体生命周期控制,包含多数据集、多检测器和多部署窗口,验证了控制器对延迟证据的不同适应策略。
AI 中文摘要
动态图异常检测需要随着图结构和类别分布漂移而反复做出决策,然而检测器基准通常在已知当前标签后对固定流程进行评分。我们提出了DynGraphAgentBench,一个在延迟反馈下用于智能体生命周期控制的可执行基准。它包含七个时间图数据集,涵盖节点级和边级异常任务,十一个可选择的检测器,以及每个数据集的八个按时间顺序的部署窗口。在每个窗口中,控制器只能看到时间因果聚合上下文、注册的模型卡以及自身已成熟的历史。它必须在当前窗口的训练或候选分数存在之前选择检测器。一个沙箱执行器在成熟数据上训练所选架构,对隐藏的部署窗口进行评分,并在一个窗口延迟后发布结果。一个确定性验证器检查决策时机、泄漏防护、合法操作、训练范围和持久化工件。我们使用平均精度和固定审查深度下的捕获率来衡量检测效用,并通过模型切换和计算量来表征适应性。来自两个主要控制器和一个无记忆参考在四个数据集上的完整八窗口轨迹,以及另外三个控制器在三个数据集上的轨迹,揭示了在无详尽当前窗口预言机的情况下对延迟证据的有用、昂贵和无效的反应。
英文摘要
Dynamic graph anomaly detection requires repeated decisions as graph structure and class prevalence drift, yet detector benchmarks usually score a fixed pipeline after current labels are known. We introduce DynGraphAgentBench, an executable benchmark for agentic lifecycle control under delayed feedback. It comprises seven temporal graph datasets with node- and edge-level anomaly tasks, eleven selectable detectors, and eight chronological deployment windows per dataset. In each window, a controller sees only time-causal aggregate context, registered model cards, and its own matured history. It must choose a detector before current-window training or candidate scores exist. A sandboxed executor trains the chosen architecture on mature data, scores a hidden deployment window, and releases the outcome after a one-window delay. A deterministic verifier checks decision timing, leakage guards, legal actions, training scope, and persisted artifacts. We measure detection utility with average precision and capture at fixed review depth, and characterize adaptation through model switches and compute. Complete eight-window trajectories from two primary controllers and a no-memory reference on four datasets, together with three additional controllers on three datasets, expose useful, costly, and ineffective reactions to delayed evidence without granting an exhaustive current-window oracle.
Comments16 pages, 2 figures, 5 tables