I Thought the Model Drifted. My Cache Key Was Serving Tuesday.
Have you ever watched an LLM endpoint return a clean answer that belonged to a different prompt entirely? I spent forty-eight hours blaming sampling noise, temperature, and a free model that would not sit still. The request logs looked honest enough, and the health check on the box stayed green the whole time. The bug was quieter than that: a cache key that hashed the user message and ignored…
We haven't written up this one. Dev.to has the full story — the link below goes straight to it.