Architecture economics · Kafka

Comparing the total cost of self-managed Kafka, Amazon MSK, and Confluent Cloud

Choosing how to run Kafka is not only a comparison of infrastructure or service prices. Each implementation changes what the provider operates, what the engineering team continues to own, and where the organization spends money over time.

BBRmap compares those economics across three specified implementation paths: self-managed Apache Kafka on AWS, Amazon MSK, and Confluent Cloud. The model combines infrastructure and vendor charges with setup engineering effort and ongoing operational work.

Explore the Kafka Demo

A fair Kafka cost comparison goes beyond the provider bill

The provider bill is only one part of the comparison. BBRmap puts three economic components on the same monthly basis.

Infrastructure and vendor charges

Compute, storage, data movement, managed capacity, and other applicable service charges for the specified implementation.

Setup engineering effort

One-time work required to design, provision, secure, automate, observe, and prepare the platform for production. The model converts this effort to cost and allocates it over 12 months.

Ongoing engineering effort

Recurring work such as platform operations, capacity management, incident response, lifecycle management, governance, and other responsibilities that continue after implementation.

The comparison makes the economic consequences of different implementation choices more visible while leaving architecture suitability as a separate decision.

The three implementations modeled

Self-managed Apache Kafka on AWS EC2 + EBS gp3

The modeled self-managed path uses AWS EC2 broker instances and EBS gp3 storage. The customer operates the Kafka infrastructure directly and retains the associated setup and ongoing responsibilities represented in the model.

Amazon MSK Provisioned

The modeled MSK configuration uses Standard brokers. AWS manages broker infrastructure while customer-side architecture, integration, governance, and platform operating work remains.

Confluent Cloud Enterprise

The modeled Enterprise service uses eCKU capacity. More of the Kafka platform is managed by the provider, while the customer continues to own surrounding architecture, integration, governance, security, reliability, and usage responsibilities represented in the model.

These are specified modeled implementation paths, not every possible Amazon MSK or Confluent Cloud configuration.

A worked example — Baseline EDA

The Baseline event-driven architecture (EDA) workload has 10 MB/s peak ingress and 10 MB/s peak read throughput, 100 leader partitions, 1 environment, and 7-day retention in US East (N. Virginia). Monthly ingress and read volumes are 1,000 GB and 500 GB respectively. No additional disaster-recovery posture is specified.

Baseline EDA modeled monthly total cost of ownership
ImplementationModeled monthly TCO
Self-managed Apache Kafka on AWS EC2 + EBS gp3$31,135
Amazon MSK Provisioned$22,733
Confluent Cloud Enterprise$24,307

USD, rounded to whole dollars. Values use the same saved Baseline result as the Demo.

Amazon MSK Provisioned has the lowest modeled monthly TCO in this example. Its infrastructure cost is higher than self-managed Kafka, but lower modeled engineering costs more than offset that difference.

These results are specific to this scenario and its assumptions. They are not general vendor price claims or a recommendation to choose an architecture based on cost alone.

Open the Baseline EDA Demo

Why the costs differ

Managed infrastructure changes the location of cost

A higher provider bill can accompany lower retained engineering work. In Baseline EDA, Confluent Cloud has lower modeled setup and ongoing engineering costs than MSK, but those savings do not offset its higher infrastructure and vendor costs. Managed services shift responsibilities; they do not eliminate customer engineering work.

Workload characteristics affect the implementations differently

The Demo explores peak ingress, peak read throughput, leader partitions, monthly read volume, and environment count. These inputs affect capacity, usage charges, and engineering work through different modeled relationships. Infrastructure can scale in discrete broker or managed-capacity units, creating step changes rather than a smooth increase in cost.

Retention and storage are also part of the saved workload assumptions. Capacity policy and workload distribution need validation before applying any modeled step to a production architecture.

Engineering effort can materially affect the comparison

Setup and ongoing work are explicit economic components, valued at $1,300 per effort-day. In Baseline EDA, these components account for approximately 98% of MSK’s modeled total. The comparison therefore depends strongly on effort estimates and their valuation, not just infrastructure prices. Lower modeled effort does not necessarily produce an equivalent headcount reduction or cash saving.

What happens as the workload changes?

The second saved scenario, Large/read-heavy, illustrates a workload with higher throughput and read volume, more partitions and environments, and longer retention. Comparing the saved scenarios shows how a different workload profile changes each implementation’s economics.

The analysis shows how a defined workload changes each implementation's economics. Crossover points depend on the workload and assumptions. The current cost-driver explorer is available for Baseline EDA.

Explore cost drivers in the Demo

What the comparison does — and does not — say

Results are modeled estimates under stated assumptions. Rates and vendor costs are not based on supplier quotes; public pricing is used where available. Current prices, discounts, traffic charges, and commercial terms need confirmation.

Engineering effort uses explicit effort assumptions and $1,300 per effort-day. Setup is allocated over 12 months; ongoing effort is recurring. Validate both the effort and cost basis against your organization.

Technical suitability still depends on performance, reliability and recovery, security, compliance, organizational capability, migration, contractual terms, and other requirements beyond TCO. The Demo supplies saved examples, not supplier quotes or live calculations from edited inputs.

Inspect the model behind the example

About BBRmap

BBRmap is designed to compare the full cost of technically viable software implementation alternatives before a team commits to one. Kafka event streaming is the first implemented Demo.

The approach extends the comparison beyond licensing and infrastructure costs to include setup engineering effort and recurring operational burden.

Learn more about BBRmap · Explore other use cases

Questions or feedback?

The assumptions are intended to be visible and challengeable. If you have experience with Kafka economics, are evaluating similar alternatives, or see something in the model that should be questioned, I'd value your perspective.

Send feedback or email feedback@bbrmap.com.