arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.01975cs.SEcs.CLcs.LGcs.PF

TELLER:面向大语言模型推理的非侵入式跨层根因分析

TELLER: Non-intrusive Cross-Layer Root-Cause Analysis for LLM Inference

Ruilin Xu, Junyi Li, Pengfei Chen, Zongxuan Xie

首次发表
浏览论文内容

中文总结 AI 辅助

TELLER是面向LLM推理的非侵入式跨层根因分析框架,通过跟踪与日志结合实现异常定位与自然语言解释,在多节点GPU工作负载上兼具高效压缩与诊断性能,为LLM推理RCA提供实用支撑。

中文摘要 AI 辅助

大语言模型(LLM)推理已从离线工作负载演变为持续运行的软件服务,但根因分析仍存在难度,因为单个请求会跨越推理引擎、Python/C++后端、主机CUDA API、GPU内核及分布式通信层。现有分析工具仅能暴露原始时间线,而基于日志的诊断常缺失跨层执行语义和请求级结构。我们提出TELLER,一种非侵入式的感知跟踪与日志的LLM推理根因分析框架。TELLER无需修改模型二进制文件,即可收集NVTX/CUPTI跟踪数据和服务日志,随后重建单请求调用链树并将日志行与对应执行步骤对齐。我们引入依赖感知的因果上下文切片,保留父子结构、时间顺序和通信关系,以及跟踪对编码(TPE)分词器,将此类切片压缩为包含父节点、深度和持续时间属性的紧凑结构 token 序列。基于这些表示,TELLER结合数值候选定位与多模态根因模型,共同预测异常步骤、定位可疑算子并生成自然语言解释。在多节点GPU推理工作负载上的实验显示存在明确的压缩-准确率权衡:中等规模的TPE词汇表可将单步骤跟踪长度减少80%以上,同时在横向(跨节点通信)和纵向(节点内执行栈)视图上实现最佳整体性能,而更激进的压缩会大幅降低诊断质量。对低故障先验、强化基线、模态消融、解释质量检查及跟踪开销的进一步分析表明,TELLER为LLM推理根因分析(RCA)提供了实用的分类和证据定位基础。

英文摘要

Large language model (LLM) inference has evolved from an offline workload into a continuously operated software service, yet root-cause analysis remains difficult because a single request spans the inference engine, Python/C++ backend, host CUDA APIs, GPU kernels, and distributed communication. Existing profilers expose raw timelines, while log-based diagnosis often misses cross-layer execution semantics and request-level structure. We present TELLER, a non-intrusive Trace- and Log-aware LLM inference Root-cause analysis framework. TELLER first collects NVTX/CUPTI traces and service logs without modifying model binaries, then reconstructs per-request call-chain trees and aligns log lines with the corresponding execution steps. We introduce a dependency-aware causal-context slice that preserves parent-child structure, temporal order, and communication relations, and a Trace Pair Encoding (TPE) tokenizer that compresses such slices into compact structural token sequences with parent, depth, and duration attributes. On top of these representations, TELLER combines numeric candidate localization with a multimodal root-cause model that jointly predicts abnormal steps, localizes suspicious operators, and generates natural-language explanations. Experiments on multi-node GPU inference workloads show a clear compression-accuracy trade-off: a moderate TPE vocabulary reduces per-step trace length by more than 80% while achieving the best overall performance on both horizontal (cross-node communication) and vertical (within-node execution stack) views, whereas more aggressive compression substantially degrades diagnosis quality. Further analyses under low-fault priors, strengthened baselines, modality ablations, explanation-quality checks, and tracing overhead show that TELLER provides a practical triage and evidence-localization substrate for LLM inference RCA.

发表机构

  • Sun Yat-sen University(中山大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑