{
  "id": 10849894,
  "title": "How We Test LLM Features So They Don't Regress in Production",
  "url": "https://urgent.news/2026/09/30/how-we-test-llm-features-so-they-dont-regress-in-production",
  "topic": "ai",
  "section": "AI",
  "published": "2026-09-30T03:53:00.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/lycore/how-we-test-llm-features-so-they-dont-regress-in-production-56io"
  },
  "original_language": "en",
  "account": "Last year, a client received an LLM-powered classification feature for a Django application. Initially, it performed well, but three weeks later, a routine prompt tweak caused it to misclassify a specific edge case — one that the team had thoroughly tested during development. The issue went unnoticed for four days because the existing automated test suite only confirmed that the endpoint returned a 200 and that the response was valid JSON. It did not verify the accuracy of the classification itself. This incident highlighted the need for automated checks to prevent such regressions in production.\n\nStandard unit tests and integration tests, built on determinism, are insufficient for LLM features. While unit tests assert that a function returns the expected output for a given input, LLMs are not deterministic. The same prompt can yield slightly different wording on different runs, and what constitutes a correct answer is often subjective rather than exact. This creates a gap where regressions can hide.\n\nTo address this, the team employs a multi-faceted evaluation (eval) system for their LLM features. Evals are tests that assess the quality of an LLM's output, rather than just its shape. There are three types of evals: exact match, model-graded, and human-labelled golden sets. The exact match evaluates whether the output contains a specific string or matches a particular value, ideal for structured outputs like classification and extraction. Model-graded uses a second LLM call to judge whether the output meets a criterion, offering flexibility but also adding cost and latency. Human-labelled golden sets use a curated set of inputs with known-correct outputs, maintaining the highest signal but being the most expensive to build and maintain.\n\nIn their Django setup, the eval runner is a management command that processes a fixed dataset of JSON fixtures. The command imports the LLM classification function, loads the dataset, and compares the model's output against the expected labels. If the model's label matches the expected label, the test passes; otherwise, it records the failure along with the input, expected output, actual output, and confidence score. The evaluation results are then displayed, including the pass rate. If the pass rate falls below a 90% threshold, the command raises an exit error, prompting a review of the prompt or model before deployment. This rigorous eval stack ensures that LLM features maintain accuracy and reliability in production environments.",
  "summary": "We shipped an LLM-powered classification feature for a client last year. It worked well. Three weeks later, after a routine prompt tweak, it started miscategorising a specific edge case — one that the team had explicitly tested for during development. Nobody noticed for four days. The problem was not the prompt change. The problem was that we had no automated check that would have caught it. Our…",
  "key_points": [
    "LLM-powered classification feature misclassified edge case post-prompt tweak",
    "Existing automated tests only checked 200 response and valid JSON",
    "Multi-faceted evaluation system prevents production regressions"
  ],
  "editors_take": "The team's multi-faceted evaluation system for LLM features closes the gap where regressions can hide, ensuring accuracy and reliability in production environments by catching subtle output changes.",
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}