流式状态开销究竟在哪里?Flink 与 Kafka Streams 在 Kafka 上的可复现比较
Where Does Streaming State Cost Go? A Reproducible Comparison of Flink and Kafka Streams on Kafka
浏览论文内容
中文总结 AI 辅助
本文通过受控实验比较 Flink 与 Kafka Streams 在 Kafka 上的恰好一次流处理,揭示状态管理架构对延迟、恢复和正确性的影响,并强调评估容错流系统需综合考量多维度指标。
中文摘要 AI 辅助
本文对使用 Apache Flink 和 Kafka Streams 实现的恰好一次语义 Kafka 管道进行了受控比较,这两种引擎具有不同的状态管理架构。Flink 中的状态管理通过外部存储中的检查点进行,而 Kafka Streams 通过重放 broker 变更日志来恢复本地状态。实验评估了这两种设计对延迟、资源消耗、配置敏感性以及注入故障后恢复的影响。在 50 次正确性试验中,两种引擎均产生了预期输出,每种引擎-工作负载组合各进行 5 次试验(精确 95% 置信区间为 0.929-1.000)。然而,在以固定速率 100 事件/秒进行的 30 分钟试验中,延迟表现出差异。对于无状态工作负载,测得的摄入到输出间隔为 4-6 毫秒,对于窗口化工作负载则为 4.3-6.2 秒。无状态工作负载的中位 p99 为 1.97-1.97 秒,窗口化工作负载的中位 p99 分别为 5.68-9.92 秒。此外,敏感性研究表明,将持久化间隔从 1,000 毫秒增加到 10,000 毫秒会改变 Kafka Streams 和 Flink 各阶段之间的延迟分布。Kafka Streams 的 W3 中位 p99 T2-T1 从 6,085 毫秒增加到 23,155 毫秒;Flink 的 T3-T2 从 732 毫秒增加到 8,702 毫秒。故障注入实验表明,延迟可能干扰不完整的处理行为。结果表明,恰好一次正确性是流处理性能的必要但不充分的度量。Kafka Streams 在有状态 JVM 终止试验中完成了 0/5,在有状态本地卷丢失试验中各完成了 1/5,而相应的 Flink 单元完成了 5/5。开销、位置和大小取决于状态管理、架构、配置和测量位置。因此,对容错流式系统的评估应整体衡量延迟、完成度、恢复行为和正确性。
英文摘要
This paper presents a controlled comparison of exactly-once Kafka pipelines implemented with Apache Flink and Kafka Streams, two engines with different state-management architectures. The state management in Flink occurs through checkpoints in external storage, whereas Kafka Streams restores local state by replaying broker changelogs. The experiments evaluate the effects of the two designs on latency, resource consumption, configuration sensitivity, and recover from injected failures. The engines produced expected outputs across 50 correctness trials, five for each engine-workload combination (exact 95% CI 0.929-1.000). The latencies, however, showed differences when tested in 30 minute trials at a fixed rate of 100 events/sec. The measured ingestion to output interval was 4-6 ms for stateless workloads and 4.3-6.2 s for windowed workloads. The median p99 was 1.97-1.97s for stateless workloads and 5.68-9.92s for windowed respectively. Further, a sensitivity study showed that increasing the durability interval from 1,000 to 10,000 ms shifted the latency distributions in Kafka Streams and Flink between stages. Kafka Streams' W3 median p99 T2-T1 increased from 6,085 to 23,155 ms; Flink's T3-T2 increased from 732 to 8,702 ms. Failure injection experiments demonstrate that latency can interfere with incomplete processing behavior. The results show that exactly-once correctness is a necessary but insufficient measure of stream processing performance. Kafka Streams completed 0/5 stateful JVM-kill and 1/5 in each stateful local-volume-loss trials, while corresponding Flink cells completed 5/5. The cost, placement and magnitude depend on state management, architecture, configuration, and measurement location. Accordingly, evaluations of fault-tolerant streaming systems should holistically measure latency, completion, recovery behavior, and correctness.
发表机构
- Indiana University(印第安纳大学)
机构由 AI 辅助整理,请以论文原文为准。