arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

分布式系统的服务健康工程

Service Health Engineering for Distributed Systems

Siva Rama Krishna Varma Bayyavarapu

arXiv 2609.08020首次发表:更新:

发表机构

Docusign Inc.(Docusign公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出服务健康工程,通过遥测、工作流完成、依赖行为等实践,结合AI辅助报告架构,确保分布式系统中用户旅程按承诺完成。

AI 中文摘要

分布式系统支撑着许多关键的业务工作流,但服务健康状况往往通过组件仪表盘而非端到端的用户结果来判断。本文提出服务健康工程作为一种实用的可靠性纪律,将遥测、工作流完成、依赖行为、运营就绪和恢复验证联系起来。以文档审批工作流作为运行示例,描述了服务承诺、服务级别指标和目标、看门狗、事件度量、韧性测试和每周服务健康审查如何揭示静默故障和滞留的异步工作。本文还提出了一种人工审查、AI辅助的报告架构,用于汇集服务健康证据,而不使AI成为自主决策者。该方法围绕用户旅程是否按承诺完成,将既有的可靠性实践整合在一起。

英文摘要

Distributed systems support many critical business workflows, but service health is often judged through component dashboards rather than through end-to-end user outcomes. This article presents service health engineering as a practical reliability discipline that connects telemetry, workflow completion, dependency behavior, operational readiness, and recovery validation. Using a document approval workflow as a running example, it describes how service promises, service-level indicators and objectives, watchdogs, incident measures, resiliency testing, and weekly service-health reviews can reveal silent failures and stranded asynchronous work. It also presents a human-reviewed, AI-assisted reporting architecture for assembling service-health evidence without making AI an autonomous decision-maker. The approach brings established reliability practices together around whether user journeys complete as promised.

Comments9 pages, 1 figure. Accepted for publication in IEEE Reliability Magazine

DOI:10.1109/MRL.2026.3727928

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑