{
  "id": 12429008,
  "title": "Un SLO para emails de CI en Kubernetes",
  "url": "https://urgent.news/2026/10/06/un-slo-para-emails-de-ci-en-kubernetes",
  "topic": "tech",
  "section": "Tech",
  "published": "2026-10-06T17:23:19.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/alexcarteruk/un-slo-para-emails-de-ci-en-kubernetes-22e4"
  },
  "original_language": "es",
  "account": "Automated email verification in CI pipelines often fails in mysterious ways. The job appears green, but the verification email never arrives in the inbox specified by the test. A small measurable SLO turns this mystery into an operable incident. The symptoms are: CI green, email missing.\n\nIn a Kubernetes pipeline, the test typically performs three actions: creates a user, requests a link, and waits for the message. When the wait expires, it looks like an application failure. However, there are several layers between the request and the assertion: the service publishes the message, the queue accepts or rejects the work, the worker processes the message, and the test inbox makes it visible. The test locates the correct message and consumes the code. If all stages share a single timeout, the team won't know where the email got lost. Moreover, retrying the entire job adds noise and slows down the queue.\n\nThe first SRE lesson is simple: don't call an email sent a five-state process. Define the SLO before touching Kubernetes. For an integration suite, a reasonable SLO might be: 99% of messages appear in the correct inbox within 30 seconds. The important thing is to have a start, an end, and a clear window. In this case: Start - API confirms the send request was accepted. End - the test finds a message with the expected run_id, recipient, and type. Error - the limit is exceeded or a message from another execution appears.\n\nDon't just measure average latency. A high p95 is often a better explanation of the pipeline experience, and the percentage of messages that never appear shows if there is real loss. For a fixture created with tempmail.so or a local stub, the rule is the same: each message needs identity and traceability. The important alert is not \"a pod is restarting,\" but \"the rate of localized messages has dropped and the error budget is being consumed.\"\n\nSeparate delivery, reading, and diagnosis. In logs, record a stable job ID, for example ci-2026-10-07-1842, but don't put full tokens or personal data. A useful metric can have these dimensions: environment and namespace; provider or transport used; message_type, such as signup or password_reset; result, with values found, timeout, or duplicate. The reading should also be explicit. A test shouldn't ask \"are there any new emails?\" but should look for the message with the run_id and advance a cursor or a known received_at. In this way, an old message doesn't accidentally make the pipeline pass.\n\nWhen processing has several stages, an observable queue helps much more than just increasing the timeout. This pattern of visible queues for long email jobs allows distinguishing provider delays from stalled workers. Minimum instrumentation for reliable email involves adding four timestamps to the event: requested_at, accepted_at, processed_at, visible_at. With them, you can calculate where the budget is consumed. If accepted_at is missing, the problem is with the API or connection. If processed_at is missing, check the queue and worker. If processed_at exists but visible_at is not, look at the inbox adapter or execution isolation. In Kubernetes, add at least a queue depth and age metric; a retry and discard counter; structured logs by run_id; an alert when p95 visibility exceeds the SLO; and a retention limit for fixtures and logs. Dashboards become noise if they only add charts. Each panel should answer a question: \"Did the message get lost before or after the worker?\" If the flow uses an agent or tool that triggers the send, document the input and output contract. Bot contracts are useful even outside of LLMs: they make clear what accepting an operation means and what evidence must be left.\n\nWhen writing a runbook, don't rely on heroes. When an alert fires, the operator should follow a short sequence: confirm if the failure affects one namespace or the whole environment; search the run_id in the API, queue, worker, and inbox; compare accepted_at with processed_at and visible_at; review duplicates, retries, and message age; pause new deployments if the error budget continues to fall; save the execution receipt before cleaning data. Don't delete the fixture before knowing what happened. The cleanup should wait for a minimum receipt. A query like tempmail.so may appear in reports, but it shouldn't mix with metric names or system keys. A failure that happens only in the middle of the night deserves the same evidence as one during business hours. Sometimes a quick alert can be resolved with a glance; other times the delay only appears after several retries.",
  "summary": "Las pruebas de email en CI fallan de una forma engañosa: el job termina en verde, pero el mensaje de verificación nunca llega a la bandeja que el test está leyendo. Un SLO pequeño y medible convierte ese misterio en un incidente operable. El síntoma: CI verde, correo perdido En un pipeline de Kubernetes es habitual que el test haga tres cosas: crea un usuario, solicita un enlace y espera el…",
  "key_points": [
    "Define SLO for email verification in CI pipelines",
    "Measure email delivery latency with start, end, and error points",
    "Instrument queue with timestamps to identify bottlenecks"
  ],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}