
doi: 10.2139/ssrn.7127185
The increasing number of Internet of Things (IoT) devices generates large volumes of continuous data, creating the need for efficient mechanisms to ingest, process, and store streaming information in real time. This paper presents a comparative study of Apache Kafka and Apache Pulsar for IoT data handling using equivalent Message Queuing Telemetry Transport (MQTT)-based ingestion pipelines. In both cases, telemetry messages are generated by IoT-like producers, transmitted through MQTT, ingested into the streaming platform, optionally processed, and stored in relational databases (DBs). The experiments examine different sink DBs, payload sizes, stream processing mechanisms, and message production rates. The evaluation focuses on end-to-end latency, intermediate pipeline delays, Central Processing Unit (CPU) utilization, and memory consumption. The results show that Apache Kafka achieves lower end-to-end latency and generally lower memory consumption in most evaluated scenarios, especially under baseline ingestion, payload-size variation, and moderate production rates. However, Kafka exhibits significant tail-latency degradation at the highest tested production rate. Apache Pulsar provides useful features such as integrated MQTT protocol handling, built-in schema management, and low processing delay through Pulsar Functions, although its complete end-to-end pipeline shows higher latency in the tested configuration. Overall, the findings show that platform selection should depend on workload intensity, latency requirements, resource constraints, processing needs, and DB persistence behavior.
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 0 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
