{
  "id": 10904004,
  "title": "Production ML Is Not About Models—It Is About Pipelines",
  "url": "https://urgent.news/2026/09/30/production-ml-is-not-about-models-it-is-about-pipelines",
  "topic": "tech",
  "section": "Tech",
  "published": "2026-09-30T09:05:05.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/robert_kawoski_20bd638fd2/production-ml-is-not-about-models-it-is-about-pipelines-5g6k"
  },
  "original_language": "en",
  "account": "Nine out of ten machine learning projects that show promise in a notebook fail to reach durable production use. The issue is not typically the model's accuracy, but the absence of a robust pipeline that can handle real-world production traffic, schema changes, and inputs that go out of distribution. The actual engineering achievement is a model that maintains performance under these challenging conditions, which is more about system design than model architecture. Across multiple engagements, the consistent lesson is that the pipeline – not the model – is the product. Models trained on last year's data encode outdated patterns, and as customer behavior shifts, upstream systems change data emission, and seasonality alters input distributions, this relentless data drift renders production models vulnerable. Adding latency constraints or operational complexity compounds the problem. The solution lies in a pipeline designed to anticipate and absorb these inevitable realities, rather than relying on a single \"better\" model. The root cause of most production ML failures is the discrepancy between training features and serving features. A data scientist might compute a rolling 30-day average using one query, while the production system might use a slightly different method, time zone, or null-handling approach. To prevent this class of bugs, teams adopt feature stores and versioned datasets, ensuring consistent feature computation and traceability from training to serving. Automated retraining as a scheduled, automated pipeline stage, rather than a manual effort, is crucial. Every new model version must undergo A/B testing before replacing the current production model to catch regressions early and build evidence of improved performance on live data. Fallback mechanisms are essential, especially in high-stakes fintech and e-commerce contexts, to prevent silent failures. When a model is uncertain or the input is out-of-distribution, it should fall back to a deterministic rule, route to human review, or even decline the decision entirely. Explainability and audit trails are equally important, providing clear insight into why a model made a specific prediction, which is vital for regulatory or financially sensitive applications. Lastly, production ML requires its own observability layer, monitoring feature drift, prediction distribution shifts, and business metric correlations to ensure the system is working as intended.",
  "summary": "Roughly nine out of ten machine learning projects that show promise in a notebook never make it to durable production use — not because the model was wrong, but because nobody built the system around it. A model that scores well on a held-out test set is a research result. A model that keeps scoring well after three months of real, shifting production traffic, survives a schema change in an…",
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}