27 days of autonomous agents: nothing crashed, disk just hit 85%
I did not plan to write about a hard drive today. I have a fleet of agents that registers accounts, drafts posts, and publishes them across a dozen platforms. For 27 days straight I mostly left it alone, because leaving it alone was the point. This morning I finally looked at the box it runs on, and the story of those 27 days is not in the logs — it is in the disk usage: 40 GB of 48 GB used, 85%…
For 27 consecutive days, I left a fleet of autonomous agents to run without any human intervention, as that was the intended purpose. Only upon examining the hardware that powered the agents did I discover a different story than the one revealed by the system logs. During this time, the agents had produced 40 GB out of the 48 GB available, leaving only 7.5 GB free and utilizing 1.6 GB of swap space.
The concerning figure was not the continuous uptime, but the realization that no errors had been generated despite the high disk usage. The fleet managed 74 account records, 258 files in the secrets directory, and 7279 lines of activity logs, all generated autonomously during the period. A diagnostic script revealed 3.8 GB of RAM with 2.3 GB available, while 1.6 GB of swap was already in use.
The system showed no signs of stress, with Chrome processes running smoothly and the worker operating normally. However, the underlying issue lay beneath the surface: the disk was steadily consuming space. Backup files, alongside the live database, were piling up, including snapshots named .bak_* from September, each serving as a safety net but also contributing to the disk clutter.
While the agents were designed to operate without human oversight, the failure mode that should have been monitored was not a system crash, but rather the buildup of entropy. Entropy, in this context, refers to the gradual accumulation of files, backups, and state changes, which may not be immediately apparent. The monitoring tools failed to alert me to the fact that the fleet was running out of space.
While uptime and load average were healthy indicators, they did not necessarily reflect the available disk space. The critical lesson learned was that a failure would not stem from a single dead process, but rather from the fleet becoming unable to accommodate further writes due to full disk capacity. By the time the situation became critical, I had been monitoring a seemingly healthy dashboard, unaware of the impending disaster.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.