Ollama num_ctx Truncated 287 of 400 Prompts and Never Told Me
My local RAG bot answered 46% of my test questions correctly. Llama 3.1 8B on Ollama, a 128K context window on the model card, retrieved chunks that I had checked by hand. The right paragraph was in the prompt every single time. The model just never saw it. Ollama's num_ctx was set to 2048 tokens, and Ollama quietly chopped the front off every prompt longer than that. No error, no field in the…
Ollama, a locally hosted Llama 3.1 8B model with a 128K context window, silently truncated prompts longer than its allocated context length of 2048 tokens. This behavior was undocumented and not communicated to the user. Out of 400 test prompts, 287 (72%) were truncated, resulting in a significant drop in accuracy from 46% to 46%.
The issue stemmed from the model discarding the beginning of prompts that exceeded the context limit, discarding crucial instructions and context. Detecting prompt truncation required comparing the `prompt_eval_count` with the `num_ctx` setting; if they matched, truncation was occurring. The fix involved setting `num_ctx` to a larger value, such as 8192, either per-request or in a Modelfile.
This increased VRAM usage by about 1 GB and slightly increased response latency, but improved accuracy to 81% on the same set of prompts. The problem highlighted the importance of understanding how context windows are managed in LLM inference pipelines.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.