Urgent.News

What's breaking now, across thousands of outlets.

Tech

Kafka Interview Prep (Part 2) โ€” When Things Break in Production

This is Part 2 of a two-part Kafka prep sheet. Part 1 covered the fundamentals โ€” partitions, brokers, delivery semantics, ISR, KRaft. This part assumes you're comfortable with that and shifts into the kind of questions that actually separate mid-level from senior/lead candidates: real-world design reasoning and what happens when production breaks . ๐Ÿ“Ž Haven't read Part 1 yet? Start there first โ†’โ€ฆ

This is Part 2 of a two-part Kafka prep sheet. It covers the kind of questions that separate mid-level from senior/lead candidates, focusing on real-world design reasoning and what happens when production breaks. The previous part covered the fundamentals such as partitions, brokers, delivery semantics, and exactly-once semantics.

A practical example is Zomato's live location tracking for delivery partners. The pipeline involves GPS updates from the delivery partner app sent to a Kafka topic called "location-updates". The updates are processed in a stream, where they are validated, filtered, enriched with additional data, and then used for business logic.

The latest location for each order is stored in Redis, an in-memory key-value store, to achieve fast read times. The customer app can either poll Redis for updates or subscribe via WebSockets for real-time data.

Interview questions in senior/lead roles focus on understanding why Kafka is the right tool for the job and how to handle production failures. For instance, streaming events instead of synchronous API calls decouples producers from consumers, allowing the system to scale and remain responsive. A message broker like Kafka enables multiple independent downstream consumers (such as map updates, ETA recalculations, fraud detection, and analytics) to consume the same data stream without the producer needing to know who's listening.

When something breaks in production, it's crucial to understand how Kafka handles fault tolerance and horizontal scalability. For example, if a consumer goes down, it doesn't lose events; they sit durably in the topic until a replacement consumer picks them up. This resiliency ensures that despite failures, the system continues to operate smoothly. The key takeaway for senior/lead interviews is to demonstrate a deep understanding of why each component is used and how they fit together into a coherent, scalable system.

Written by urgent.news from Dev.to's reporting โ€” not their text. Machine-written โ€” may contain errors; check the original before relying on it.

Read the original at dev.to โ†’

More in Tech

operationId and tags in OpenAPI: naming conventions that keep generated SDKs and docs usable

Open a generated SDK where the methods are named getV1UsersByIdGet , postV1UsersPost , and usersGet2 , and the cost of careless operationId s is immediate: nobody can discover anything, and everyโ€ฆ

  • OperationId and tag fields are crucial naming conventions in OpenAPI
  • Poorly chosen operationId leads to broken callers and flat documentation sidebar
  • Tags should form a small, stable taxonomy aligned to business domains

Health checks for APIs: /healthz, /readyz, /livez, version, and metrics in OpenAPI

A single /health endpoint that checks the database, the cache, and three downstream services, and returns 500 if any of them is briefly unreachable, is actively harmful.

  • The /health endpoint can be /livez, /readyz, /livez, version, and metrics in OpenAPI
  • /livez probe should only check if the process is alive, not if dependencies are reachable
  • /readyz probe verifies if the instance can serve requests right now, stopping traffic if not

More from Thursday 8 October โ†’