llama-server's sleep mode loses or crashes on a request that arrives just before it sleeps
TL;DR : llama-server --sleep-idle-seconds N unloads the model after N idle seconds and is documented to reload it for "any new incoming task". A request handler checks that the server is awake when it starts, then tokenizes the prompt, then queues the task. If the idle timer fires in between, the server goes to sleep anyway. On llama.cpp b11368 (CPU, Gemma 3 1B), a 17,653-token prompt sent 46 to…
The llama-server's sleep mode can fail or crash when a request arrives just before it goes to sleep. This issue was reported with llama.cpp b11368 on CPU, Gemma 3 1B. A 17,653-token prompt sent 46 to 34 milliseconds before the server fell asleep remained in the queue with no response until a second request woke the server 32 times out of 32.
When the request was sent 34 to 3 milliseconds before sleep, it crashed the server with SIGSEGV inside the tokenizer, which was reading the vocabulary that had just been freed by the sleep. With a two-word prompt, the window required to avoid the issue was too short, resulting in no responses in 93 out of 93 tries. Sending a one-token completion before the main request avoided both the crash and the hang in 123 out of 123 attempts.
The root cause is that the server's check for sleep mode is triggered before the request is processed, specifically when a request handler checks that the server is awake, tokenizes the prompt, and queues the task. This check fails if a request arrives while the server is in the process of going to sleep. The documentation states that new tasks will trigger the model to reload, but this isn't true if the request arrives slightly before the server has fallen asleep.
The problem is more pronounced when the server has a vision projector loaded (--mmproj) and speculative decoding is enabled, leading to a lower chance of the issue occurring.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.