Urgent.News

What's breaking now, across thousands of outlets.

Tech

Wpipe: Resilience for Long-Running Engineering with Atomic Checkpoints

Wpipe: Resilience for Long-Running Engineering with Atomic Checkpoints Day 10 of the Wisrovi Open Source Architecture Series. Why is failure recovery in data and processing pipelines still a manual, brittle task? Automate resilience, not just the flow. In modern orchestration frameworks, task status is often reduced to a status string in a remote database: Scheduled, Running, or Failed. But for…

Day 10 of the Wisrovi Open Source Architecture Series explores the importance of failure recovery in data and processing pipelines. Modern orchestration frameworks often reduce task status to a simple status string, but senior engineers handling long-running or mission-critical workflows need more context. The wpipe Checkpoint engine provides this missing level of context by saving the exact state of memory at the moment of failure. This allows engineers to resume workflows without re-computing from scratch.

The wpipe architecture can be implemented as a SaaS cloud orchestrator or a sovereign local-first engine. The local engine, running as a single Python process with an embedded SQLite WAL engine, offers zero-config setup and immediate resume capabilities. Atomic data checkpoints, achieved through SQLite WAL logs, enable seamless task re-dispatch from the remote server. This local-first approach provides network resilience, immune to partitions, and requires minimal resource footprint.

A practical example demonstrates atomic checkpointing in just 10 lines of code. By using the wpipe Pipeline and Step classes, developers can create an Atomic Checkpointing system in just ten lines of Python code. This lightweight library is suitable for edge devices, Raspberry Pi, and CI/CD runs. The source repository and PyPI page provide further details on implementing this resilient data processing workflow.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in Tech

CrowdStrike outage explained: 21 fields, 20 slots, 8.5M blue screens

The CrowdStrike outage of July 19, 2024 is the biggest IT failure most of us have lived through, and its root cause fits in one sentence: a detection template declared 21 input fields, and the code…

  • CrowdStrike outage affected 8.5 million Windows devices globally
  • Single line of code in detection template caused 20 vs 21 field mismatch
  • Outage led to 7,000 flight cancellations and $500M in damages

Wpipe: Resilience for long-duration engineering with Atomic Checkpoints

Wpipe: Resiliencia para ingeniería de larga duración con Checkpoints Atómicos Día 10 de la serie técnica Wisrovi Open Source Architecture.

  • Wpipe offers atomic checkpoints for long-duration engineering resilience
  • Modern orchestration frameworks lack full context for failure recovery
  • Wpipe's SQLite-based persistence enables in-situ task recovery

More from Tuesday 29 September →