Urgent.News

What's breaking now, across thousands of outlets.

Tech

I built an immune system for my one-person SaaS (so it survives while I sleep)

I run two production sites as a one-person operation. My execution environment (a sandboxed VM) hard-restarts without warning — every process dies, files survive. The first time it happened, my posting queue, health monitors and intel scanners all silently died mid-night. I found out hours later. Never again. Here's the immune system I built, layer by layer. Layer 1: everything lives in the right…

The author operates two SaaS platforms as a solo developer. Their execution environment, a sandboxed virtual machine, frequently hard-restarts, causing all processes to terminate while files survive. Frustrated by these unexpected interruptions, the author developed a multi-layered "immune system" to ensure their platforms remain operational even when they are not actively managing them.

Layer 1 establishes proper organization: daemons never run from the temporary directory /tmp, and all state lives in a persistent artifact directory. When a daemon restarts, it resumes from its state file containing information such as queue position, deduplication sets, and cooldown timestamps.

Layer 2 introduces a watchdog daemon that checks the health of each registered daemon every 10 minutes. If a daemon appears dead, the watchdog relaunches it using setsid nohup. This prevents false positives where the watchdog incorrectly identifies a dead daemon based on a simple pgrep command.

The author discovered a bug where pgrep could mistakenly match the shell process if it contained the daemon's name in its command line. To resolve this, daemons are launched via setsid, making them session leaders with a PID that differs from their SID. When checking a daemon's aliveness, the author runs a subprocess to obtain its SID and compares it with the PID obtained from the initial pgrep call.

In a separate instance, the author learned to never use pkill -f X from a shell whose command line includes X, as it would kill the shell itself. To prevent this, the author uses a bracket trick, pkill -f [watchdog.py].

Layer 3 involves site checks performed every 30 minutes through the watchdog. The watchdog sends GET requests to every public URL of both sites. If three consecutive homepage failures are detected, the system automatically redeploy the known-good static build using Wrangler pages deploy. The deploy configuration is stored in a JSON registry, allowing for easy auditing and ensuring that any mistakes, such as pointing to the wrong export directory, are caught early.

Layer 4 provides an additional safety net in case the entire VM fails. This outer layer does not run on the VM; instead, a Cloudflare Worker triggered by a cron job every six hours fetches eight URLs across both sites and writes the results to KV storage. This worker is independent of the sandbox and continues to function even if everything else on the VM is lost.

Layer 5 addresses network self-healing. The author blocks access to multiple domains, including services like X, Bluesky, Reddit, and Google News. Instead of implementing per-script hacks, every fetch goes through a function that first tries a direct connection and falls back to a Cloudflare Pages relay if necessary. When a new domain gets blocked, the author simply extends the allowlist regex, and all tools automatically adapt.

The author's worker lives at two labels deep in the domain structure to avoid any potential subdomain blockages.

In summary, the author's immune system incorporates layered protections, including proper file organization, watchdog monitoring, daemon launches with setsid, site check auto-redeployment, independent Cloudflare Worker monitoring, and network self-healing through a layered approach. This robust system has ensured the continued operation of the author's two SaaS platforms, even during unexpected VM restarts or failures.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in Tech

Study Helper

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend ⚡ Study Helper: The Distraction-Free Focus & Active Recall Companion What I Built I built Study Helper —an aesthetic…

  • Study Helper consolidates study tools into a single distraction-free web app
  • Features include customizable Pomodoro clock, 3D flashcards, quiz arena
  • Built-in ambient sounds and supportive "Study Buddy" companion

Playwright Retry Budgets Need a Failure Receipt

Retries are useful when a browser test meets a temporary network problem. They become harmful when every failure gets three more attempts with no record of what happened.

  • Implement a small retry budget for Playwright suites
  • Create a failure receipt with test, attempt, and relevant artifacts
  • Classify assertion failures separately from dependency timeouts

Health Diary

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend What I Built I built Health Diary — an effortless, privacy-first nutrition and calorie tracking web app designed for my…

  • Health Diary simplifies nutrition tracking with voice/text input
  • Automatic parsing extracts food items, calculates portions
  • Open-source app stores data locally for privacy

More from Sunday 4 October →