arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.23301cs.DC

大规模AI基础设施的精确分布式追踪:时间同步作为可靠可观测性的基础

Accurate Distributed Tracing for Large-Scale AI Infrastructure: Time Synchronization as a Foundation for Reliable Observability

Hesham Elbakoury, Ankur Sharma

首次发表
浏览论文内容

中文总结 AI 辅助

针对大规模AI基础设施中时钟精度不足导致分布式追踪失效的问题,提出TempoTrace系统,通过协同设计PTP时间同步与追踪,降低时间戳不确定性并提升因果顺序保持与故障归因准确性。

中文摘要 AI 辅助

在大规模AI基础设施中,当时钟精度不足时,分布式追踪会静默失效:因果事件顺序错乱,故障归因被破坏,性能诊断不可靠。我们提出了TempoTrace系统,该系统将IEEE 1588v2 PTP时间同步与分布式追踪协同设计,以在异构、多租户GPU集群中保持因果顺序。我们形式化证明了在WAN/云环境下,NTP级时钟在25-30%的有向操作对中产生因果倒置(在LAN NTP下为11.3%)。一个拉普拉斯重尾噪声模型将分析扩展到高斯假设之外,给出加权期望错序率为27.76%,而高斯模型为29.37%,确认了对尾部形状的鲁棒性。TempoTrace通过GPUDirect RDMA硬件时间戳,将GPU到主机的时间戳不确定性从2.1微秒降低到设计目标0.056微秒的残差标准差,并实现亚100纳秒的PTP同步。一个有界的Sketch Vector Clock在最多四个参与者的情况下,以低于10^-5的假阳性率追踪因果关系。一个混合基于规则和XGBoost的诊断引擎在受控验证中(1,800个事件语料库;McNemar p=1.25e-23)实现了宏F1分数0.974。一个形式化证明的多租户模型提供了隔离的逻辑时钟域。对于RoCEv2/ECMP网络,基于P4的带内网络遥测纠正了路径延迟不对称性,实现了99.1%的归因准确率。应用置信策略允许工作负载指定精度要求和降级模式回退。评估结合了物理五节点测量(NTP倒置率44.978%,在拥塞下AIC更偏好拉普拉斯拟合而非高斯拟合)和受控合成验证,配置从512到16,384个H100 GPU,涵盖InfiniBand和RoCEv2网络。

英文摘要

Distributed tracing in large-scale AI infrastructure fails silently when clock accuracy is insufficient: causal events are misordered, fault attribution is corrupted, and performance diagnoses are unreliable. We present TempoTrace, a system that co-designs IEEE 1588v2 PTP time synchronization with distributed tracing to preserve causal ordering across heterogeneous, multi-tenant GPU clusters. We formally prove that NTP-grade clocks produce causal inversions at 25-30% of directed operation pairs under WAN/cloud conditions (11.3% under LAN NTP). A Laplace heavy-tail noise model extends the analysis beyond Gaussian assumptions, giving a weighted expected misorder rate of 27.76% vs. 29.37% Gaussian, confirming robustness to tail shape. TempoTrace reduces GPU-to-host timestamp uncertainty from 2.1 us to a design target of 0.056 us residual standard deviation via GPUDirect RDMA hardware timestamping, with sub-100 ns PTP synchronization. A bounded Sketch Vector Clock tracks causal relationships with false-positive rate below 10^-5 for up to four participants. A hybrid rule-based and XGBoost diagnosis engine achieves macro-F1 0.974 in controlled validation (1,800-incident corpus; McNemar p=1.25e-23). A formally proven multi-tenant model provides isolated logical clock domains. For RoCEv2/ECMP fabrics, P4-based in-band network telemetry corrects path-delay asymmetry, yielding 99.1% attribution accuracy. Application Confidence Policies let workloads specify precision requirements and degraded-mode fallback. Evaluation combines physical five-node measurements (NTP inversion rate 44.978%, Laplace fit preferred over Gaussian by AIC under congestion) and controlled synthetic validation with illustrative configurations from 512 to 16,384 H100 GPUs across InfiniBand and RoCEv2 fabrics.

补充信息

↑