April 25th, 2025

What If We Could Rebuild Kafka from Scratch?

Gunnar Morling proposes "Kafka.next," a cloud-optimized version of Kafka, featuring key-centric access, topic hierarchies, concurrency control, extensibility, synchronous commit callbacks, snapshotting, and multi-tenancy for improved data management.

Read original article

FrustrationSkepticismCuriosity

What If We Could Rebuild Kafka from Scratch?

Gunnar Morling discusses the potential for a new version of Kafka, referred to as "Kafka.next," which would be designed from the ground up to better suit cloud environments. He highlights recent developments like KIP-1150 ("Diskless Kafka") and AutoMQ’s Kafka fork, which aim to enhance Kafka's functionality in cloud settings. Morling outlines several desirable features for this new system, including the elimination of partitions in favor of key-centric access, which would allow for more efficient data retrieval and processing. He suggests implementing topic hierarchies for better subscription management, concurrency control to prevent outdated data writes, and broker-side schema support to improve data integrity. Additionally, he emphasizes the importance of extensibility and pluggability, enabling users to customize the system without altering its core. Other proposed features include synchronous commit callbacks for stronger consistency, snapshotting for better state management, and built-in multi-tenancy to support diverse workloads. Morling notes that while some of these features exist in other systems, no single open-source platform currently combines all of them. He invites feedback from others who have experience with Kafka or similar platforms to contribute their ideas.

- Morling proposes a new version of Kafka designed for cloud environments.

- Key features include key-centric access, topic hierarchies, and concurrency control.

- Emphasis on extensibility and pluggability for user customization.

- Synchronous commit callbacks and snapshotting are suggested for improved data management.

- Multi-tenancy is highlighted as essential for modern data systems.

The Essence of Apache Kafka

Apache Kafka is a distributed event-driven architecture that enables efficient real-time data streaming, ensuring fault tolerance and scalability through an append-only log structure and partitioned topics across multiple nodes.

Kafka at the low end: how bad can it get?

The blog post outlines Kafka's challenges as a job queue in low-volume scenarios, highlighting unfair job distribution, increased latency, and recommending caution until improvements from KIP-932 are implemented.

Apache Kafka 4.0 Released

Apache Kafka 4.0.0 has been released, eliminating ZooKeeper, enhancing scalability, and introducing new features like improved consumer performance and queue semantics support, while requiring updated Java versions and removing deprecated APIs.

KIP-1150: Diskless Kafka Topics

KIP-1150 proposes Diskless Topics in Apache Kafka to optimize storage and reduce costs by using object storage, enabling multi-region active-active topics and automatic failover, enhancing Kafka's market competitiveness.

AI: What people are saying

The comments on Gunnar Morling's "Kafka.next" proposal reveal a mix of skepticism and alternative suggestions regarding Kafka's design and functionality.

Many users express frustration with Kafka's complexity and limitations, suggesting that it may not be the best solution for all use cases.
Alternatives to Kafka, such as NATS and Apache Pulsar, are mentioned as potentially simpler and more effective options.
Concerns are raised about Kafka's design choices, particularly regarding partitioning and latency, with some suggesting that it may not be suitable for certain applications.
There is a call for reflection on the purpose of engineering improvements, with some commenters questioning the need for a cloud-optimized Kafka.
Several users mention ongoing developments in the space, including LinkedIn's Northguard and other projects that aim to address Kafka's shortcomings.

24 comments

By @galeaspablo - about 1 hour

Agreed. The head of line problem is worth solving for certain use cases.

But today, all streaming systems (or workarounds) with per message key acknowledgements incur O(n^2) costs in either computation, bandwidth, or storage per n messages. This applies to Pulsar for example, which is often used for this feature.

Now, now, this degenerate time/space complexity might not show up every day, but when it does, you’re toast, and you have to wait it out.

My colleagues and I have studied this problem in depth for years, and our conclusion is that a fundamental architectural change is needed to support scalable per message key acknowledgements. Furthermore, the architecture will fundamentally require a sorted index, meaning that any such a queuing / streaming system will process n messages in O (n log n).

We’ve wanted to blog about this for a while, but never found the time. I hope this comment helps out if you’re thinking of relying on per message key acknowledgments; you should expect sporadic outages / delays.

By @vim-guru - about 7 hours

https://nats.io is easier to use than Kafka and already solves several of the points in this post I believe, like removing partitions, supporting key-based streams, and having flexible topic hierarchies.

By @nitwit005 - about 7 hours

> When producing a record to a topic and then using that record for materializing some derived data view on some downstream data store, there’s no way for the producer to know when it will be able to "see" that downstream update. For certain use cases it would be helpful to be able to guarantee that derived data views have been updated when a produce request gets acknowledged, allowing Kafka to act as a log for a true database with strong read-your-own-writes semantics.

Just don't use Kafka.

Write to the downstream datastore directly. Then you know your data is committed and you have a database to query.

By @Ozzie_osman - about 8 hours

I feel like everyone's journey with Kafka ends up being pretty similar. Initially, you think "oh, an append-only log that can scale, brilliant and simple" then you try it out and realize it is far, far, from being simple.

By @peanut-walrus - about 7 hours

Object storage for Kafka? Wouldn't this 10x the latency and cost?

I feel like Kafka is a victim of it's own success, it's excellent for what it was designed, but since the design is simple and elegant, people have been using it for all sorts of things for which it was not designed. And well, of course it's not perfect for these use cases.

By @tyingq - about 2 hours

He mentions Automq right in the opener. And if I follow the link, they pitch it in a way that sounds very "too good to be true".

Anyone here have some real world experience with it?

By @supermatt - about 5 hours

> "Do away with partitions"

> "Key-level streams (... of events)"

When you are leaning on the storage backend for physical partitioning (as per the cloud example, where they would literally partition based on keys), doesnt this effectively just boil down to renaming partitions to keys, and keys to events?

By @fintler - about 6 hours

Keep an eye out for Northguard. It's the name of LinkedIn's rewrite of Kafka that was announced at a stream processing meetup about a week ago.

By @olavgg - about 5 hours

How many of the Apache Kafka issues are adressed by switching to Apache Pulsar?

I skipped learning Kafka, and jumped right into Pulsar. It works great for our use case. No complaints. But I wonder why so few use it?

By @vermon - about 7 hours

Interesting, if partitioning is not a useful concept of Kafka, what are some of the better alternatives for controlling consumer concurrency?

By @frklem - about 6 hours

"Faced with such a marked defensive negative attitude on the part of a biased culture, men who have knowledge of technical objects and appreciate their significance try to justify their judgment by giving to the technical object the only status that today has any stability apart from that granted to aesthetic objects, the status of something sacred. This, of course, gives rise to an intemperate technicism that is nothing other than idolatry of the machine and, through such idolatry, by way of identification, it leads to a technocratic yearning for unconditional power. The desire for power confirms the machine as a way to supremacy and makes of it the modern philtre (love-potion)." Gilbert Simondon, On the mode of existence of technical objects.

This is exactly what I interpret from these kind of articles: engineering just for the cause of engineering. I am not saying we should not investigate on how to improve our engineered artifacts, or that we should not improve them. But I see a generalized lack of reflection on why we should do it, and I think it is related to a detachment from the domains we create software for. The article suggests uses of the technology that come from so different ways of using it, that it looses coherence as a technical item.

By @elvircrn - about 6 hours

Surprised there's no mention of Redpanda here.

By @mgaunard - about 4 hours

I can't count the number of bad message queues and buses I've seen in my career.

While it would be useful to just blame Kafka for being bad technology it seems many other people get it wrong, too.

By @lewdwig - about 6 hours

Ah the siren call of the ground-up rewrite. I didn’t know how deep the assumption of hard disks underpinning everything is baked into its design.

But don’t public cloud providers already all have cloud-native event sourcing? If that’s what you need, just use that instead of Kafka.

By @0x445442 - about 3 hours

How about logging the logs so I can shell into the server to search the messages.

By @rvz - about 1 hour

Every time another startup falls for the Java + Kafka arguments, it keeps the AWS consultants happier.

Fast forward into 2025, there are many performant, efficient and less complex alternatives to Kafka that save you money, instead of burning millions in operational costs "to scale".

Unless you are at a hundred million dollar revenue company, choosing Kafka in 2025 is doesn't make sense anymore.

By @Spivak - about 8 hours

Once you start asking to query the log by keys, multi-tenancy trees of topics, synchronous commits-ish, and schemas aren't we just in normal db territory where the kafka log becomes the query log. I think you need to go backwards and be like what is the feature a rdbms/nosql db can't do and go from there. Because the wishlist is looking like CQRS with the front queue being durable but events removed once persisted in the backing db where the clients query events from the db.

The backing db in this wishlist would be something in the vein of Aurora to achieve the storage compute split.

By @imcritic - about 7 hours

Since we are dreaming - add ETL there as well!

By @ghuntley - about 6 hours

I'm going to get downvoted for this, but you can literally rebuild Kafka via AI right now in record time using the steps detailed at https://ghuntley.com/z80.

I'm currently building a full workload scheduler/orchestrator. I'm sick of Kubernetes. The world needs better -> https://x.com/GeoffreyHuntley/status/1915677858867105862

By @YetAnotherNick - about 7 hours

I wish there is a global file system with node local disks, which has rule driven affinity to nodes for data. We have two extremes, one like EFS or S3 express which doesn't have any affinity to the processing system, and other what Kafka etc is doing where they have tightly integrated logic for this which makes systems more complicated.

By @Mistletoe - about 7 hours

I know it’s not what the article is about but I really wish we could rebuild Franz Kafka and hear what he thought about the tech dystopia we are in.

>I cannot make you understand. I cannot make anyone understand what is happening inside me. I cannot even explain it to myself. -Franz Kafka, The Metamorphosis

By @hardwaresofton - about 8 hours

See also: Warpstream, which was so good it got acquired by Confluent.

Feels like there is another squeeze in that idea if someone “just” took all their docs and replicated the feature set. But maybe that’s what S2 is already aiming at.

Wonder how long warpstream docs, marketing materials and useful blogs will stay up.

What If We Could Rebuild Kafka from Scratch?