Designs Kafka streaming pipelines end to end - partition-count sizing math, key choice, consumer-group sizing, offset strategy, delivery semantics, poll tuning, and dead-letter handling - with production configs. Use when someone asks "how many partitions should this topic have", "my consumer group keeps rebalancing", "consumer lag keeps growing", "how do I handle poison messages", or is designing a new topic or event-driven pipeline. Do NOT use for Spark batch or PySpark job tuning - use spark-jobs instead; do NOT use for HTTP webhook ingestion reliability - use webhook-receiver-hardener instead.
Click to play with sound.
---
name: Kafka Pipelines
description: Designs Kafka streaming pipelines end to end - partition-count sizing math, key choice, consumer-group sizing, offset strategy, delivery semantics, poll tuning, and dead-letter handling - with production configs. Use when someone asks "how many partitions should this topic have", "my consumer group keeps rebalancing", "consumer lag keeps growing", "how do I handle poison messages", or is designing a new topic or event-driven pipeline. Do NOT use for Spark batch or PySpark job tuning - use spark-jobs instead; do NOT use for HTTP webhook ingestion reliability - use webhook-receiver-hardener instead.
---
# Kafka Pipelines
Partition count is the one Kafka decision that is nearly permanent: it caps consumer parallelism forever, and increasing it later rehashes keys and breaks per-key ordering for every keyed consumer. The two costly mistakes this skill prevents are under-partitioning a topic that later gets hot, and committing offsets before processing - which silently loses records on the first crash.
## Operating procedure
### Step 1: Gather inputs
Collect these before choosing any number. Label estimates as guesses and revisit after load testing.
1. Peak produce throughput (MB/s and messages/s) - peak, not average. Default guess: 3x the average.
2. Average message size.
3. Per-record consumer processing time, including downstream calls (DB writes, HTTP). This usually dominates.
4. Ordering requirement: which records must stay ordered relative to each other? That defines the key.
5. Replay window: how far back must consumers be able to re-read? Drives retention.
6. Delivery requirement: at-least-once (default) or exactly-once.
7. Growth expectation over the topic's lifetime.
### Step 2: Size the partition count
… load the full skill through Skill Me