发表机构
University of Southern California(南加州大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对分布式应用程序源级交互式调试难题,提出DDB。通过分布式回溯等方法,解决进程边界、调试状态存活及超时级联等问题。集成代码少,性能良好,在用户研究中故障定位成功率高,为分布式应用调试提供有效方案。
AI 中文摘要
交互式调试是理解程序源级行为的有效工具,可让开发者暂停执行、查看调用栈和检查运行时状态。但交互式调试器专为单进程执行设计,分布式系统中存在诸多问题,如调用栈在进程边界停止、调试状态难在基础设施动态变化中存活,调试器引发的执行暂停会触发灾难性超时级联等。本文提出DDB,它将交互式调试能力扩展到分布式应用程序。针对各挑战给出解决方案,如用分布式回溯(DBT)跨RPC边界重建统一调用栈、用意图保留控制平面管理分布式会话生命周期、用暂停擦除时间(PET)防止超时级联。DDB与RPC框架集成代码仅20 - 60行。在多种框架和多达122个进程上评估,DDB实现了30ms的中位数跨RPC回溯延迟等良好性能,在用户研究中故障定位成功率达100%,中位数定位时间约8分钟。
英文摘要
Interactive debugging is an effective tool for understanding program behavior at the source level, allowing developers to pause execution, navigate the call stack, and inspect runtime state. However, interactive debuggers are designed for single-process execution, and interactive debugging has been widely considered impractical for distributed systems. Call stacks stop at process boundaries, debugging state fails to survive infrastructure dynamics, and, most critically, debugger-induced execution pauses trigger catastrophic timeout cascades that destroy the intended debug flow. Consequently, developers are forced to abandon live hypothesis testing in favor of unwieldy and iterative log-and-redeploy cycles. We present DDB, a source-level interactive debugger that extends interactive debugging capabilities to distributed applications. We show that each of these challenges admits a targeted solution. To bridge disjoint processes, Distributed Backtrace (DBT) embeds compact causality metadata in every RPC and reconstructs a unified call stack across RPC boundaries. To manage the lifecycle of a distributed session, an intent-preserving control plane automatically coordinates and propagates breakpoints across dynamic process sets. To make pausing safe, Pause-Erased Time (PET) virtualizes each process's clock, decoupling logical time from physical pauses and preventing timeout cascades. DDB integrates with an RPC framework in 20-60 lines of code. Evaluated on gRPC, ServiceWeaver, Nu, and Quicksand across up to 122 processes, DDB achieves 30ms median cross-RPC backtrace latency, sub-5 ms time jump under repeated execution pauses, and adds 1-5% throughput overhead, comparable to attaching a single-process debugger. In a controlled user study, DDB achieves a 100% fault localization success rate (compared to 38.5% for baseline tools) with a median localization time of ~8 minutes.
Comments22 pages in total (14 pages for the main body, 5 pages for the appendix), 11 figures, 6 tables. Published at SOSP'26