Open-source data streaming platform
You can't break Kwaque with an earthquake.
Kwaque is a next-generation data streaming platform. It is designed to remove the limits that Kafka and systems like it cannot avoid, and it is built first for AI infrastructure, where data grows fast and unevenly.
What is Kwaque?
Modern software is many separate programs working together. One takes payments, another sends email, another answers questions with an AI model. They stay in step by passing each other events: an order was placed, a user signed in, a model returned an answer.
A data streaming platform carries those events. It keeps them in order, stores them safely, and lets any program read them again later. Events are grouped into named streams called topics. Today this job usually goes to Apache Kafka, which most of the world's largest companies run,1 or to a system built to behave like it.
Kwaque does the same job on a different design. It is not a faster copy of Kafka. It changes the parts of the design that make these systems costly to run.
Why Kwaque?
Kafka's design dates from around 2010 and was right for its time. Four of its core decisions have since become limits. They belong to the design itself, so tuning does not remove them, and systems that keep the design in order to stay compatible with Kafka keep the limits too, however fast they run.
-
Capacity is fixed in advance
A topic is divided into a set number of partitions when it is created. The number can be raised later but never lowered, and raising it can move a key to a different partition, which breaks the order of that key's events.2
In Kwaque: a topic divides itself as its traffic grows and joins back together when traffic falls.
-
Adding a machine means moving data
A partition lives whole on particular machines. A new machine takes none of the existing load until partitions are reassigned to it, and every partition that moves is copied in full.2
In Kwaque: a new machine takes new data straight away. Nothing old has to move.
-
Recovery means copying whole partitions
When a machine is lost, every partition it held is one copy short. Restoring those copies is not automatic: an operator or a separate tool has to reassign the partitions, and each one is copied in full.3
In Kwaque: new data goes to healthy machines at once, and the missing copies are restored in the background, one small piece at a time.
-
One controller coordinates everything
A single active controller keeps the record of every topic, partition and machine, and applies changes to it one at a time.4
In Kwaque: that record is divided among many small groups of machines, so no single one stands in the way.
The aim is a platform that a small team can run without a specialist on call, and that follows traffic up and down by itself.
Built for AI infrastructure
These limits cost everyone who runs a streaming platform. They cost AI companies the most, for three reasons.
Data grows several times over in a year
Nobody can choose the right size for a topic that will be many times bigger by next year.
Traffic swings hard within a day
The busiest hour can be many times the quietest. A system sized for the peak sits mostly idle, and one that cannot follow the load falls behind.
Teams are small
Few AI companies want a group of specialists whose whole job is keeping the streaming platform alive.
Kwaque is being built for the parts of an AI product where data has to arrive in order and cannot be lost:
- Usage and billing events, where no model call can go uncounted.
- Messages between services and agents, with the history kept so it can be replayed.
- Live signals for models, delivered while they are still fresh.
Most AI teams are building these pipelines for the first time. They can start on a design made for this kind of load, instead of adapting one that was not.
How it works
This section describes the design Kwaque is being built to. Where it is today says which parts exist.
Topics, ranges and segments
A topic's keys are shared out among ranges. Each range owns one share of the keys and keeps the events for those keys in order. A range stores its events as a series of segments. A segment takes new events until it reaches a set size or age. It is then sealed and never changes again, and the range opens a new one.
Each segment is copied to several machines independently of every other segment. The unit that is copied, placed and repaired is therefore always one small segment, never a topic's whole history.
A topic holds every key
Its ranges each owns a share of the keys
Range B's segments oldest to newest
Copies of the open segment each on a different machine
Ranges split and merge
The number of ranges is not fixed. When a range carries too much traffic, it is closed and two new ranges take over, each with half of its keys. When two neighbouring halves fall quiet, they are closed and a single range takes over both. Only the ranges involved are affected. The rest of the topic carries on.
Events that share a key stay in order throughout, because a range is always closed before its successors accept any data. One limit remains: a single very busy key cannot be divided, since dividing it would break its order.
- A new topic is one range, R1, covering every key.
- Traffic grows. R1 is closed, and R2 and R3 each take half of its keys.
- The lower half is still busy, so R3 is closed and splits into R4 and R5.
- Traffic falls. R4 and R5 are closed and R6 takes over both.
Each column shows all of the topic's keys from top to bottom. Blue marks the ranges created at that step.
New segments go where there is room
Every new segment is given its own set of machines, chosen from those with spare capacity. Because segments are opened all the time, new load spreads across the machines without anyone moving old data. A machine added today starts receiving segments today.
Fixed partitions
The new machine stays empty until whole partitions are copied onto it.
Kwaque
The newest segments go straight to the new machine. The old ones stay where they are.
A failure closes one segment, not the topic
If a machine holding a copy of an open segment fails, Kwaque does not wait for it. It seals that segment at the last point every surviving copy agrees on, and opens the next segment on a healthy set of machines. From then on, new data has its full number of copies again. Anything the sender had not yet had confirmed is sent again, and Kwaque recognises a repeated batch, so nothing is stored twice.
Sealed segments that lost a copy are restored afterwards in the background, at a lower priority than live traffic. Each one is small and never changes, so a repair can be verified, paused and resumed.
Before
Five segments, three copies of each. Segment 5 is open, on C, D and E.
Machine C fails
Segment 5 is sealed. Segment 6 opens on A, B and E, and writing continues.
Repair
Segments 1, 3 and 5 each get a new third copy in the background.
- Copy of a sealed segment
- Copy of the open segment
- Copy restored after the failure
No single controller
The record of what exists and where it lives is cut into a fixed number of slices. Each slice is kept by its own small group of machines, which agree on every change using the Raft consensus algorithm.5 A topic's settings and each of its ranges are placed in slices independently, so even one very large topic spreads its bookkeeping across many groups.
One small directory remains, saying which machines keep which slice. It changes only when the set of machines changes, and takes no part in day-to-day work.
One controller
Every change to any topic, partition or machine goes through one active controller.
Kwaque
A topic's settings and each of its ranges land in different slices, so changes to them do not queue behind one another.
Confirmed means on disk, on every copy
A machine holding a copy of the open segment writes each batch of events to two files, a write-ahead log and the segment itself, and flushes both to disk. The sender gets its confirmation only after every copy has done so. Each batch carries a checksum, and after a crash Kwaque verifies what it finds on disk before using it.
- The application sends a batch to the machine that leads the open segment, here machine 2. That machine passes it to the other copies.
- Every copy writes the batch to its write-ahead log and to its segment file, and flushes both to disk.
- Only when every copy reports both files safely on disk does the application get its confirmation.
One thread for each processor core
Kwaque is written in C++ on the Seastar framework. Each processor core runs a single thread with its own memory, its own connections and its own queue of work. Cores share nothing and pass messages to one another instead of locking shared data, so they never wait on each other, and there is no garbage collector to pause them.6
Segment data is read and written directly, without going through the operating system's cache, and each kind of work has a fixed memory budget. Both keep behaviour predictable when the system is under load.
Core 1
memoryconnectionswork queueCore 2
memoryconnectionswork queueCore 3
memoryconnectionswork queueCore 4
memoryconnectionswork queueNothing is shared between cores. When one needs something from another, it sends a message.
What it costs
The case for Kwaque is about cost as much as design. This section prices a Kafka deployment line by line, then shows which lines Kwaque is meant to change. Every input is stated, and every outside figure links to its source.
What Kafka costs to run
The model uses Amazon's list prices7, 8 for a common setup: three copies of the data in three zones, seven days of data kept, and one application reading it. Zones are the separate data centres inside a cloud region, and traffic between them is charged for. An engineer costs $200,000 a year with benefits and overhead.15, 16, 17 Disks are kept half full, which is not a pessimistic guess: Honeycomb reported running its Kafka disks 20% full.12
At 100 megabytes of events a second (MB/s), that comes to about $986,000 a year. People are the largest line until a deployment is very large. After that it is storage and the traffic between zones. Servers, the part a faster engine reduces, are 2% to 4% of the total.
The result is within 20% of Confluent's own worked example9 and of an independent engineer's estimate.14 Grab has said that traffic between zones was half the cost of its Kafka platform.11
The number of engineers is the softest input. Confluent, the main Kafka vendor, puts an early production deployment at two full-time engineers and cites a streaming team of seven to ten at Lyft.10 A study it commissioned from Forrester counted 10.5 people running the platform at a large enterprise.13 The model uses half an engineer at 10 MB/s, two at 100 and eight at 1,000.
10 MB/s $159,000 a year
100 MB/s $986,000 a year
1,000 MB/s $7.46 million a year
- People
- Storage
- Traffic between zones
- Servers
Show the arithmetic for 100 MB/s
| Line | How it is worked out | A year |
|---|---|---|
| People | 2 engineers × $200,000 | $400,000 |
| Storage | 181,440 GB × $0.08 a month × 12, doubled because disks are half full | $348,365 |
| Traffic between zones | 3,153,600 GB × 3.33 crossings × $0.02 | $210,240 |
| Servers | $2,281 a month × 12 | $27,372 |
| Total | $985,977 |
100 MB/s is 3,153,600 GB a year. Seven days of it is 60,480 GB, and three copies make 181,440 GB. Storage is $0.08 per GB a month.8 Traffic between zones is $0.01 per GB in each direction, so $0.02 per crossing.7 Each GB crosses twice to make the other two copies, and writers and readers each sit in a different zone two times out of three, which gives 3.33 crossings. The server figure is taken from Confluent's example for this workload.9 An engineer is a $142,750 mid-range salary15, 16 multiplied by 1.4 for benefits and overhead.17
What Kwaque aims to change
Kwaque's design goes after three of the four lines.
- People. Nobody plans capacity or moves data by hand. Target: an operations team 60% smaller.
- Storage. Segments go wherever there is room, so disks fill evenly. Target: 75% full instead of 50%.
- Servers. Adding a machine moves no data, so less spare capacity is kept running. Target: half as many.
- Traffic between zones. Unchanged. Kwaque still keeps three copies in three zones.
The chart shows the target case, at about $616,000 a year against $986,000 for Kafka, and a best case, which takes a team 80% smaller on top of the storage and server targets.
A faster engine on Kafka's design is shown for comparison on generous terms: a third of the servers and an operations team 25% smaller.
These are targets, not measurements. Kwaque does not yet run across machines, so none of this has been measured.
Kafka as usually run $986,000
A faster engine on the same design $868,000, 12% less
Kwaque, target $616,000, 38% less
Kwaque, best case $536,000, 46% less
- People
- Storage
- Traffic between zones
- Servers
| Setup | A year | Saving |
|---|---|---|
| Kafka as usually run | $7.46 million | |
| A faster engine on the same design | $6.88 million | 8% |
| Kwaque, target | $5.20 million | 30% |
| Kwaque, best case | $4.88 million | 35% |
What the estimate rests on
About two thirds of the saving in the target case comes from the smaller operations team, and that is the least proven input. How big the cut is depends on how much of a team's work is tied to the number of separate clusters it runs, because this design is meant to let one large cluster replace many.
The first table shows the result for three sizes of cut.
The one outside figure is from the Forrester study: 10.5 people before a company moved to a managed service and 3.5 after, a cut of 67%.13 That is a different change from the one Kwaque makes, so it is a reference point and not evidence.
Kwaque does not cut everything. Three copies still cross between zones and still sit on disk. Copying between zones alone is about 31% of the infrastructure bill, and it stays. Systems that keep data in cloud object storage do cut it, in exchange for slower delivery.
The comparison covers running costs at list prices. It leaves out discounts, and it leaves out the one-off cost of moving existing applications, which is real because Kwaque has its own protocol.
| Team smaller by | A year | Saving |
|---|---|---|
| 40% | $696,000 | $290,000 (29%) |
| 60%, the target | $616,000 | $370,000 (38%) |
| 80%, the best case | $536,000 | $450,000 (46%) |
| Line | Saving | Share |
|---|---|---|
| People | $240,000 | 65% |
| Storage | $116,000 | 31% |
| Servers | $14,000 | 4% |
| Traffic between zones | $0 | 0% |
Where it is today
Kwaque is early in its development. This list separates what exists from what is planned.
- Working today
- The storage engine for a single machine. It writes each batch to a write-ahead log and to a segment, confirms it only once it is safely on disk, takes checkpoints so that a restart is quick, and rebuilds its state after a crash. The server starts, reports its health and shuts down cleanly, but it does not accept data over the network yet.
- How it is tested
- By breaking it on purpose. The disk, network and clock can be replaced with simulated ones, so tests can crash the engine in the middle of a write, lose data that was never made safe, and replay any failure exactly. Every change also runs under fuzzing and memory-error checkers.
- In development
- The protocol for sending and reading data, and everything that involves more than one machine: copying segments, dividing up the record of the system, and splitting and merging ranges.
- Kafka compatibility
- None today. Kwaque has its own protocol, so applications written for Kafka will not work with it unchanged.
- Production use
- Not yet. Do not trust it with real data.
Kwaque runs on Linux, on x86-64 and ARM processors. It is released under the Apache 2.0 licence, and all of the work happens in the open on GitHub.
Who is building it
I'm Vikram Aditya Verma. I started Kwaque in August 2026 and I'm building it on my own.
If you run a streaming system today and know what hurts about it, or you are building data pipelines for an AI product, I'd like to hear from you. You can open an issue on GitHub or reach me directly.
- Vikram Aditya Verma
- GitHub
- VikramAditya33
Sources
The outside facts and figures on this page, grouped by subject. The numbers match the small references in the text.
Kafka and the design
- Apache Kafka project home page. The project's own figure for how widely Kafka is used. kafka.apache.org
- Apache Kafka documentation, basic operations. Changing the number of partitions, adding machines and reassigning partitions. kafka.apache.org/43/operations/basic-kafka-operations
- Apache Kafka improvement proposal KIP-46, Self Healing Kafka. The copies held by a lost machine are not moved elsewhere automatically. cwiki.apache.org, KIP-46
- Apache Kafka documentation, KRaft. One controller is active at a time. kafka.apache.org/43/operations/kraft
- Diego Ongaro and John Ousterhout, In Search of an Understandable Consensus Algorithm. The Raft paper. raft.github.io/raft.pdf
- Seastar, shared-nothing design. One thread for each core, with no memory shared between them. seastar.io/shared-nothing
Cost model
- Amazon Web Services, data transfer price list for US East. $0.01 per GB in each direction between zones. pricing.us-east-1.amazonaws.com, AWSDataTransfer
- Amazon Web Services, block storage pricing. $0.08 per GB a month for general-purpose volumes. aws.amazon.com/ebs/pricing
- Confluent, Understanding and optimizing your Kafka costs, part 1: infrastructure. The worked example behind the server line, and a cross-check on the whole model. confluent.io/blog, Kafka costs part 1
- Confluent, Understanding and optimizing your Kafka costs, part 2: development and operations. How many engineers a deployment needs. confluent.io/blog, Kafka costs part 2
- Grab Engineering, on removing the cost of traffic between zones. Traffic between zones was half the cost of its Kafka platform. engineering.grab.com/zero-traffic-cost
- Honeycomb, on scaling Kafka for its observability pipelines. How full its Kafka disks and processors run. honeycomb.io/blog/scaling-kafka-observability-pipelines
- Forrester, Total Economic Impact study of Confluent. Commissioned by Confluent. People running the platform before and after a move to a managed service. tei.forrester.com/go/IBM/confluent
- Stanislav Kozlovski, The Brutal Truth about Apache Kafka Cost Calculators. An independent engineer's estimate of the same costs. bigdata.2minutestreaming.com
- Levels.fyi, site reliability engineer pay in the United States. levels.fyi
- Robert Half, site reliability engineer salary. roberthalf.com
- US Bureau of Labor Statistics, Employer Costs for Employee Compensation. The share of an employee's cost that is benefits, behind the 1.4 multiplier. bls.gov, ECEC release