微服务遥测的动态采样:一种基于强化学习与熵的方法
Dynamic Sampling for Telemetry in Microservices: A Reinforcement Learning and Entropy-Based Approach
浏览论文内容
中文总结 AI 辅助
本文提出RADAR,一种结合强化学习与熵评估的微服务追踪采样智能体,基于OpenTelemetry实现动态采样,在Kubernetes模拟实验中降低带宽97.4%、CPU 99.0%,保留85.6%稀有模式并提升熵25%,验证了熵驱动遥测的可行性。
中文摘要 AI 辅助
微服务架构日益部署于基于云计算的分布式环境中,使得应用开发与维护更具动态性,但也增加了故障排查与可观测性的复杂性。因此,分布式追踪工具对于请求分析与调试至关重要,尽管它们会引入额外的开销,而过量且低效的数据收集可能放大这种开销。本文提出RADAR(用于动态且相关追踪采样的强化学习智能体),该智能体结合强化学习与数据熵评估,基于OpenTelemetry标准,实现对系统监控相关追踪的更高效捕获。RADAR测试不同的采样规则以发现哪种组合最为高效。一个模拟极简在线商店的测试环境,其中多个微服务分布在Kubernetes集群上,作为实验的基础,实验评估了智能体的收敛性以及系统在资源消耗和收集数据质量方面的性能。结果表明,与全量数据收集相比,RADAR将网络带宽消耗降低了97.4%,CPU使用率降低了99.0%,同时优于固定速率采样基线。除了这些资源节省,该方法还保留了对关键场景的可观测性,保留了约85.6%的稀有追踪模式,并将存储信息的平均熵提高了约25%,验证了利用熵自主且高效编排遥测的可行性。
英文摘要
Microservices architectures are increasingly deployed in cloud-based distributed environments, making application development and maintenance more dynamic, but also increasing the complexity of troubleshooting and observability. Distributed tracing tools are therefore essential for request analysis and debugging, despite introducing additional overhead that can be amplified by excessive and inefficient data collection. This article proposes RADAR (Reinforcement Learning Agent for Dynamic And Relevant trace sampling), an agent that combines reinforcement learning with a data entropy assessment to achieve more efficient capture of traces relevant to system monitoring, based on the OpenTelemetry standard. RADAR tests different sampling rules to discover which combination is most efficient. A test environment simulating a minimalist online store with several microservices distributed across a Kubernetes cluster served as the basis for the experiments, which evaluated the agent's convergence and the system's performance in terms of resource consumption and collected data quality. Results showed that RADAR reduced network bandwidth consumption by 97.4% and CPU usage by 99.0% compared to full data collection, also outperforming a fixed-rate sampling baseline. Beyond these resource savings, the approach preserved observability of critical scenarios, retaining approximately 85.6% of rare trace patterns and increasing the average entropy of the stored information by approximately 25%, validating the feasibility of using entropy to orchestrate telemetry autonomously and efficiently.
发表机构
- Federal University of Rio Grande do Sul(南大河联邦大学)
机构由 AI 辅助整理,请以论文原文为准。