Your monitoring says healthy. Your agents are not.
What Broke In the weeks leading up to our most recent production run, our local LLM agent fleet experienced a series of failures that went unnoticed for days. Looking at the failure ledger, several patterns emerge: Process Errors in Inference : Multiple entries show process_error failures during the secretary/consult-classify and secretary/consult-fields stages. For example: On August 29th at…
We haven't written up this one. Dev.to has the full story — the link below goes straight to it.