International Journal of Computer
Trends and Technology

Research Article | Open Access | Download PDF
Volume 74 | Issue 8 | Year 2026 | Article Id. IJCTT-V74I8P101 | DOI : https://doi.org/10.14445/22312803/IJCTT-V74I8P101

Where Does Streaming State Cost Go? A Reproducible Comparison of Flink and Kafka Streams on Kafka


Kiran N Kumar, Santhosh Kumar Saminathan

Received Revised Accepted Published
19 Jun 2026 26 Jul 2026 14 Aug 2026 30 Aug 2026

Citation :

Kiran N Kumar, Santhosh Kumar Saminathan, "Where Does Streaming State Cost Go? A Reproducible Comparison of Flink and Kafka Streams on Kafka," International Journal of Computer Trends and Technology (IJCTT), vol. 74, no. 8, pp. 1-8, 2026. Crossref, https://doi.org/10.14445/22312803/IJCTT-V74I8P101

Abstract

This paper presents a controlled comparison of exactly-once Kafka pipelines implemented with Apache Flink and Kafka Streams, two engines with different state-management architectures. The state management in Flink occurs through checkpoints in external storage, whereas Kafka Streams restores local state by replaying broker changelogs. The experiments evaluate the effects of the two designs on latency, resource consumption, configuration sensitivity, and recover from injected failures. The engines produced expected outputs across 50 correctness trials, five for each engine-workload combination (exact 95% CI 0.929-1.000). The latencies, however, showed differences when tested in 30 minute trials at a fixed rate of 100 events/sec. The measured ingestion to output interval was 4-6 ms for stateless workloads and 4.3-6.2 s for windowed workloads. The median p99 was 1.97-1.97s for stateless workloads and 5.68-9.92s for windowed respectively. Further, a sensitivity study showed that increasing the durability interval from 1,000 to 10,000 ms shifted the latency distributions in Kafka Streams and Flink between stages. Kafka Streams’ W3 median p99 T2-T1 increased from 6,085 to 23,155 ms; Flink’s T3-T2 increased from 732 to 8,702 ms. Failure injection experiments demonstrate that latency can interfere with incomplete processing behavior. The results show that exactly-once correctness is a necessary but insufficient measure of stream processing performance. Kafka Streams completed 0/5 stateful JVM-kill and 1/5 in each stateful local-volume-loss trials, while corresponding Flink cells completed 5/5. The cost, placement and magnitude depend on state management, architecture, configuration, and measurement location. Accordingly, evaluations of fault-tolerant streaming systems should holistically measure latency, completion, recovery behavior, and correctness.

Keywords

Latency Decomposition, Stream Processing Benchmark, Exactly-Once Processing, Flink, Kafka Streams.

References

[1] Apache Flink, Checkpointing. [Online]. Available:
https://nightlies.apache.org/flink/flink-docs-stable/docs/dev/datastream/fault-tolerance/checkpointing/

[2] Apache Flink, Fault Tolerance via State Snapshots. [Online]. Available:
https://nightlies.apache.org/flink/flink-docs-stable/docs/learn-flink/fault_tolerance/

[3] Apache Kafka, Managing Streams Application Topics, 2025. [Online]. Available:
https://kafka.apache.org/40/streams/developer-guide/manage-topics/

[4] Apache Kafka, Architecture, 2025. [Online]. Available: https://kafka.apache.org/31/streams/architecture/

[5] Apache Kafka, Kraft, 2026. [Online]. Available: https://kafka.apache.org/35/operations/kraft/

[6] Apache Kafka, Quickstart, 2026. [Online]. Available: https://kafka.apache.org/quickstart/

[7] Apache Flink, Downloads. [Online]. Available: https://flink.apache.org/downloads/

[8] Maven Central Repository, flink-connector-kafka, 2025. [Online]. Available: 
https://central.sonatype.com/artifact/org.apache.flink/flink-connector-kafka

[9] Pete Tucker et al., “NEXMark: A Benchmark for Queries over Data Streams (Draft),” OGI School of Science and Engineering, Oregon Health and Science University, Technical Report, 2002.
[
Google Scholar]

[10] Sanket Chintapalli et al., “Benchmarking Streaming Computation Engines: Storm, Flink and Spark Streaming,” 2016 IEEE International Parallel and Distributed Processing Symposium Workshops (IPDPSW), Chicago, IL, USA, pp. 1789-1792, 2016.
[
CrossRef] [Google Scholar] [Publisher Link]

[11] Guenter Hesse et al., “ESPBench: The Enterprise Stream Processing Benchmark,” Proceedings of the ACM/SPEC International Conference on Performance Engineering (ICPE), Virtual Event, France, pp. 201-212, 2021.
[
CrossRef] [Google Scholar] [Publisher Link]

[12] Soren Henning, and Wilhelm Hasselbring, “Theodolite: Scalability Benchmarking of Distributed Stream Processing Engines in Microservice Architectures,” Big Data Research, vol. 25, 2021.
[
CrossRef] [Google Scholar] [Publisher Link]

[13] Maycon Viana Bordin et al., “DSPBench: A Suite of Benchmark Applications for Distributed Data Stream Processing Systems,” IEEE Access, vol. 8, pp. 222900-222917, 2020.
[
CrossRef] [Google Scholar] [Publisher Link]

[14] Soren Henning et al., “ShuffleBench: A Benchmark for Large-scale Data Shuffling Operations with Distributed Stream Processing Frameworks,” Proceedings of the 15th ACM/SPEC International Conference on Performance Engineering, London, United Kingdom, pp. 2-13, 2024.
[
CrossRef] [Google Scholar] [Publisher Link]

[15] Adriano Vogel et al., “A Comprehensive Benchmarking Analysis of Fault Recovery in Stream Processing Frameworks,” Proceedings of the 18th ACM International Conference on Distributed and Event-Based Systems, Villeurbanne, France, pp. 171-182, 2024.
[
CrossRef] [Google Scholar] [Publisher Link]

[16] Soren Henning, and Wilhelm Hasselbring, “Benchmarking Scalability of Stream Processing Frameworks Deployed as Microservices in the Cloud,” Journal of Systems and Software, vol. 208, pp. 1-17, 2024.
[
CrossRef] [Google Scholar] [Publisher Link]

[17]  Jeyhun Karimov et al., “Benchmarking Distributed Stream Data Processing Systems,” 2018 IEEE 34th International Conference Data Engineering, Paris, France, pp. 1507-1518, 2018.
[
CrossRef] [Google Scholar] [Publisher Link]