AI 中文总结
研究微服务跟踪异常检测,提出将事件编码为(端点,调用链)对的方法,实现未见链标记和上下文条件预测。基于此的CHAINLSTM在TrainTicket基准测试中F1值达94.3%,提升了检测性能,为阈值检测提供更大分离余量。
AI 中文摘要
即使每个跨度都正常返回,微服务跟踪在结构上也可能存在异常,比如支付流程悄悄跳过风险检查,每个跨度监视器却认为正常。像DeepLog这样的序列模型通过预测下一个事件来解决此问题,但它们将每个API端点视为无上下文的令牌。我们提出将每个事件编码为(端点,从根到跨度的调用链)对。这一简单改变有两个结果:无需模型推理就能标记未见链,下一个事件预测变为上下文条件式,将微妙的路径异常变为明显的异常值。我们在CHAINLSTM中实现了这一想法,它是一个轻量级双任务LSTM,支持每个事件的在线检测。在TrainTicket基准测试中,CHAINLSTM的F1值达到94.3%(比DeepLog高5.3个百分点),具有可比的延迟召回率和99.1%的路径召回率。案例分析表明,链感知编码将路径异常的预测概率中位数从0.91降至0.002,为基于阈值的检测提供了更大的分离余量。
英文摘要
Microservice traces can be structurally anomalous even when every span returns normally -- a payment flow that silently skips a risk check looks fine to any per-span monitor. Sequence models like DeepLog address this by predicting the next event, but they treat each API endpoint as a context-free token: the same endpoint reached through different invocation chains is mapped to the same vocabulary entry, even when its normal behavior differs across contexts. We propose encoding each event as an (endpoint, root-to-span invocation chain) pair instead. This simple change has two consequences: unseen chains are flagged without model inference, and next-event predictions become context-conditional, turning subtle path anomalies into clear outliers. We instantiate this idea in CHAINLSTM, a lightweight dual-task LSTM supporting per-event online detection. On the TrainTicket benchmark, CHAINLSTM achieves 94.3% F1 (+5.3 pp over DeepLog) with comparable latency recall and 99.1\% path recall. Case analysis shows that chain-aware encoding shifts median prediction probability on path anomalies from 0.91 to 0.002, suggesting a wider separation margin for threshold-based detection.