arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TelemetrySuffBench:智能体遥测是否足以用于故障起源诊断?

TelemetrySuffBench: Is Agent Telemetry Sufficient for Failure-Origin Diagnosis?

Yuxuan Zhu, Peng Pu

arXiv 2608.07899首次发表:更新:

AI 中文总结

该研究推出TelemetrySuffBench基准,评估五个前沿语言模型的故障起源诊断能力,发现检测与定位存在差距,安全弃权(不执行)具强模型依赖性,需决策-来源链接与弃权保障。

AI 中文摘要

智能体系统日益暴露执行轨迹,但揭示故障的遥测仍可能不足以识别故障起源。我们推出TelemetrySuffBench,这是一个受控基准,可在证据不足的情况下区分故障检测、故障起源定位和安全弃权(不执行)。该基准构建具有延迟绑定故障的规范多组件轨迹,并将其呈现为配对的粗粒度视图、七因素遥测掩码以及完全相等的模糊起源对。我们使用统一协议、显式候选集、无效输出统计、子组分析和冻结的盲保留集评估五个前沿语言模型。在完整遥测下,各模型的起源步骤Top-1准确率介于33.8%至97.2%之间。元数据、OpenTelemetry兼容视图和OpenInference兼容视图的检测F1值保持99.5%至100%,但将起源步骤准确率限制在最多0.5%,暴露出稳健的检测-定位差距。因素消融进一步表明,移除决策内容会使所有模型的起源步骤准确率降至零,而移除来源也会导致模型相关的大幅损失。在需要弃权(不执行)的丰富模糊输入上,证据门控使三个模型的无支持唯一起源答案减少12.5至48.6个百分点,而两个模型仍回答所有情况,显示安全弃权存在强模型依赖性。冻结保留集的结果在同一生成器家族内重现了核心模式。这些发现表明,终端状态可支持检测,而可靠的因果归因需要显式的决策-来源链接和跨模型有效的弃权(不执行)保障措施。数据集和基准实现可在该httpsURL获取。

英文摘要

Agent systems increasingly expose execution traces, yet telemetry that reveals a failure may still be inadequate for identifying where that failure originated. We introduce TelemetrySuffBench, a controlled benchmark that separates failure detection, fault-origin localization, and safe abstention under insufficient evidence. The benchmark constructs canonical multi-component traces with delayed-binding faults and renders them as paired coarse views, seven-factor telemetry masks, and exact-equal ambiguous origin pairs. We evaluate five frontier language models using unified protocols, explicit candidate sets, invalid-output accounting, subgroup analyses, and a frozen blind holdout. With full telemetry, origin-step Top-1 accuracy ranges from 33.8% to 97.2% across models. Metadata, OpenTelemetry-compatible, and OpenInference-compatible views retain 99.5% to 100% detection F1 while limiting origin-step accuracy to at most 0.5%, exposing a robust detection-localization gap. Factor ablations further show that removing decision content reduces origin-step accuracy to zero for every model, while provenance removal also causes large model-dependent losses. On rich ambiguous inputs that require abstention, evidence gating reduces unsupported unique-origin answers by 12.5 to 48.6 percentage points for three models, whereas two models still answer every case, revealing strong model dependence in safe abstention. Results on the frozen holdout reproduce the central pattern within the same generator family. These findings show that terminal status can support detection, whereas reliable causal attribution requires explicit decision-to-provenance links and abstention safeguards that remain effective across models. The dataset and benchmark implementation are available at https://anonymous.4open.science/r/TelemetrySuffBench-E635/README.md.

Comments10 pages, 3 figures, 4 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑