AI 中文总结
研究持久工作流执行中升级风险量化问题,提出基于遥测数据的封闭形式概率模型,结合向后和向前项分解风险,全程贝叶斯估计,给出风险分数可信区间,还能划分运行类别,放弃独立性假设后可计算耦合感知分区。
AI 中文摘要
持久工作流引擎通过确定性地重放不可变事件日志来重建执行状态,将每个正在运行的实例与生成其历史记录的代码版本相关联。新的部署可能会使旧版本下启动的运行重放无效,从而破坏状态或停止进度。现有缓解措施将每个更改视为极其危险,不适用于长时间休眠的工作流。我们提出了一种封闭形式的概率模型,仅使用静态结构差异和遥测数据(协议已持久化的事件日志、步骤有效负载、历史路径)来量化从工作流版本\(V_1\)升级到\(V_2\)时正在运行实例的风险,无需预演、沙盒或影子执行。风险沿三个轴(协议、接口、状态迁移)分解,结合了精确的向后(重新水化)项和概率性的向前项。估计全程采用贝叶斯方法,工作流升级风险(WUR)分数带有可信区间,遥测数据稀疏表示不确定性。我们证明零向后风险判定可证明在新版本下安全重新水化,并得出将运行划分为迁移、审查和固定类别的策略。最后我们放弃运行间独立性假设,通过经验耦合图捕获通过钩子、层次结构和共享资源的耦合,舰队风险成为故障传播算子的最小不动点,耦合感知的迁移/固定分区精确计算为最小\(s-t\)割。
英文摘要
Durable workflow engines reconstruct execution state by deterministically replaying an immutable event log, coupling every in-flight run to the code version that produced its history: a new deployment can invalidate the replay of runs started under the old version, silently corrupting state or halting progress. Existing mitigations -- pinning, patch gates, side-by-side deployment -- treat every change as maximally dangerous and drain old versions, untenable for workflows that sleep for weeks. We present a closed-form probabilistic model that quantifies the risk of upgrading in-flight runs from workflow version $V_1$ to $V_2$ using only a static structural diff and telemetry the protocol already persists -- event logs, step payloads, historical paths -- with no dry-run, sandbox, or shadow execution. Risk decomposes along three axes (protocol, interface, state migration) and combines an exact backward (rehydration) term, computed on recorded prefixes modulo trace equivalence of concurrent completions, with a probabilistic forward term from hitting probabilities in an empirically estimated Markov model of control flow. Estimation is Bayesian throughout, so the Workflow Upgrade Risk (WUR) score carries a credible interval and thin telemetry surfaces as uncertainty. We prove that a zero backward-risk verdict certifies safe rehydration under the new version, and derive a policy partitioning runs into migrate, review, and pin classes. Finally we drop the inter-run independence assumption: coupling through hooks, hierarchy, and shared resources is captured by an empirical coupling graph, fleet risk becomes the least fixpoint of a failure-contagion operator, and the coupling-aware migrate/pin partition is computed exactly as a minimum s-t cut.
Comments29 pages, 3 figures