{
  "id": 12788704,
  "title": "Kafka Interview Prep (Part 2) — When Things Break in Production",
  "url": "https://urgent.news/2026/10/08/kafka-interview-prep-part-2-when-things-break-in-production",
  "topic": "tech",
  "section": "Tech",
  "published": "2026-10-08T04:44:00.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/tejaswipandava/kafka-interview-prep-part-2-when-things-break-in-production-2ljf"
  },
  "original_language": "en",
  "account": "This is Part 2 of a two-part Kafka prep sheet. It covers the kind of questions that separate mid-level from senior/lead candidates, focusing on real-world design reasoning and what happens when production breaks. The previous part covered the fundamentals such as partitions, brokers, delivery semantics, and exactly-once semantics.\n\nA practical example is Zomato's live location tracking for delivery partners. The pipeline involves GPS updates from the delivery partner app sent to a Kafka topic called \"location-updates\". The updates are processed in a stream, where they are validated, filtered, enriched with additional data, and then used for business logic. The latest location for each order is stored in Redis, an in-memory key-value store, to achieve fast read times. The customer app can either poll Redis for updates or subscribe via WebSockets for real-time data.\n\nInterview questions in senior/lead roles focus on understanding why Kafka is the right tool for the job and how to handle production failures. For instance, streaming events instead of synchronous API calls decouples producers from consumers, allowing the system to scale and remain responsive. A message broker like Kafka enables multiple independent downstream consumers (such as map updates, ETA recalculations, fraud detection, and analytics) to consume the same data stream without the producer needing to know who's listening.\n\nWhen something breaks in production, it's crucial to understand how Kafka handles fault tolerance and horizontal scalability. For example, if a consumer goes down, it doesn't lose events; they sit durably in the topic until a replacement consumer picks them up. This resiliency ensures that despite failures, the system continues to operate smoothly. The key takeaway for senior/lead interviews is to demonstrate a deep understanding of why each component is used and how they fit together into a coherent, scalable system.",
  "summary": "This is Part 2 of a two-part Kafka prep sheet. Part 1 covered the fundamentals — partitions, brokers, delivery semantics, ISR, KRaft. This part assumes you're comfortable with that and shifts into the kind of questions that actually separate mid-level from senior/lead candidates: real-world design reasoning and what happens when production breaks . 📎 Haven't read Part 1 yet? Start there first →…",
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}