{
  "id": 11970360,
  "title": "I built an immune system for my one-person SaaS (so it survives while I sleep)",
  "url": "https://urgent.news/2026/10/04/i-built-an-immune-system-for-my-one-person-saas-so-it-survives-while",
  "topic": "tech",
  "section": "Tech",
  "published": "2026-10-04T17:25:53.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/mystiqueracing/i-built-an-immune-system-for-my-one-person-saas-so-it-survives-while-i-sleep-5g67"
  },
  "original_language": "en",
  "account": "The author operates two SaaS platforms as a solo developer. Their execution environment, a sandboxed virtual machine, frequently hard-restarts, causing all processes to terminate while files survive. Frustrated by these unexpected interruptions, the author developed a multi-layered \"immune system\" to ensure their platforms remain operational even when they are not actively managing them.\n\nLayer 1 establishes proper organization: daemons never run from the temporary directory /tmp, and all state lives in a persistent artifact directory. When a daemon restarts, it resumes from its state file containing information such as queue position, deduplication sets, and cooldown timestamps.\n\nLayer 2 introduces a watchdog daemon that checks the health of each registered daemon every 10 minutes. If a daemon appears dead, the watchdog relaunches it using setsid nohup. This prevents false positives where the watchdog incorrectly identifies a dead daemon based on a simple pgrep command.\n\nThe author discovered a bug where pgrep could mistakenly match the shell process if it contained the daemon's name in its command line. To resolve this, daemons are launched via setsid, making them session leaders with a PID that differs from their SID. When checking a daemon's aliveness, the author runs a subprocess to obtain its SID and compares it with the PID obtained from the initial pgrep call.\n\nIn a separate instance, the author learned to never use pkill -f X from a shell whose command line includes X, as it would kill the shell itself. To prevent this, the author uses a bracket trick, pkill -f [watchdog.py].\n\nLayer 3 involves site checks performed every 30 minutes through the watchdog. The watchdog sends GET requests to every public URL of both sites. If three consecutive homepage failures are detected, the system automatically redeploy the known-good static build using Wrangler pages deploy. The deploy configuration is stored in a JSON registry, allowing for easy auditing and ensuring that any mistakes, such as pointing to the wrong export directory, are caught early.\n\nLayer 4 provides an additional safety net in case the entire VM fails. This outer layer does not run on the VM; instead, a Cloudflare Worker triggered by a cron job every six hours fetches eight URLs across both sites and writes the results to KV storage. This worker is independent of the sandbox and continues to function even if everything else on the VM is lost.\n\nLayer 5 addresses network self-healing. The author blocks access to multiple domains, including services like X, Bluesky, Reddit, and Google News. Instead of implementing per-script hacks, every fetch goes through a function that first tries a direct connection and falls back to a Cloudflare Pages relay if necessary. When a new domain gets blocked, the author simply extends the allowlist regex, and all tools automatically adapt. The author's worker lives at two labels deep in the domain structure to avoid any potential subdomain blockages.\n\nIn summary, the author's immune system incorporates layered protections, including proper file organization, watchdog monitoring, daemon launches with setsid, site check auto-redeployment, independent Cloudflare Worker monitoring, and network self-healing through a layered approach. This robust system has ensured the continued operation of the author's two SaaS platforms, even during unexpected VM restarts or failures.",
  "summary": "I run two production sites as a one-person operation. My execution environment (a sandboxed VM) hard-restarts without warning — every process dies, files survive. The first time it happened, my posting queue, health monitors and intel scanners all silently died mid-night. I found out hours later. Never again. Here's the immune system I built, layer by layer. Layer 1: everything lives in the right…",
  "key_points": [
    "Author creates multi-layered immune system for SaaS platforms",
    "Layer 1 ensures proper organization and persistent state",
    "Layer 2 uses watchdog daemon to monitor and relaunch daemons"
  ],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}