AI 中文总结
ORCA是基于可观测性的微服务故障修复流水线,通过遥测数据提炼故障特征定位候选位置,经多智能体生成补丁并通过遥测验证器评估,在575案例基准中成本效益优于所有基线。
AI 中文摘要
微服务故障通常从运营遥测数据中诊断,但自动化程序修复(APR)系统通常从问题报告、局部代码上下文或失败测试开始,这种不匹配导致基于遥测的诊断与补丁生成之间存在差距。本文提出ORCA,一种基于可观测性的微服务故障修复流水线:首先将配对的故障与参考遥测数据的差异提炼为故障特征,再利用该特征识别候选代码和部署配置位置;修复图智能体与探索智能体从这些位置生成统一差异格式的补丁候选;ORCA通过基于遥测的补丁验证器评估生成的补丁,该验证器区分补丁有效性、语法与语义正确性、测试预言完整性及遥测重放。在含575个案例的基准测试中,ORCA在成本效益上优于所有评估的基线。结果表明,运营遥测可从诊断证据转化为可操作的修复上下文:配对遥测支持面向修复的定位,修复图智能体将定位后的代码与配置证据转化为面向大语言模型(LLM)的受限补丁生成上下文,基于遥测的验证则能发现仅通过问题报告或测试评估无法发现的修复结果。
英文摘要
Microservice failures are often diagnosed from operational telemetry. However, automated program repair systems usually start from issue reports, localized code context, or failing tests. This mismatch leaves a gap between telemetry-based diagnosis and patch generation. We present ORCA, an observability-grounded APR pipeline for microservice incidents. ORCA first distills the differences in paired failure and reference telemetry into a fault signature, then uses the signature to identify candidate code and deployment-configuration locations. Repair graph agents and an Exploration agent generate unified-diff patch candidates from these locations. ORCA evaluates generated patches with a Telemetry-Grounded Patch Verifier that separates patch validity, syntactic and semantic correctness, test-oracle integrity, and telemetry replay. On a 575-case benchmark, ORCA outperforms all evaluated baselines in terms of cost-effectiveness. Results show that operational telemetry can be transformed from diagnostic evidence into actionable repair context: paired telemetry supports repair-oriented localization, while repair graph agents convert localized code and configuration evidence into constrained patch-generation context for the LLM. Telemetry-grounded verification then exposes repair outcomes that issue- or test-only evaluation would miss.