The Same Prompt Twice Is Double Quota: A Verifiable Cache for Free Model Endpoints
Free model endpoints bill quota per token. Send the same prompt twice, and you pay twice. Retries protect you from failures. They do not protect you from duplicates. A cache layer does. This tutorial builds one from zero. Every stage ends with a verification step. No frameworks. No dependencies. One Node.js file. The problem CI jobs repeat prompts. Tests re-run the same summarization. Previews…
Stage 2 — Add the cache
The cache key is a SHA-256 hash of the HTTP method, URL, and raw request body. The same request will always produce the same cache key. The cache is implemented as a JavaScript Map object, which stores keys and values. When a request is received, the cache key is generated using the crypto module's SHA-256 hash function. The cache key is then used to look up the corresponding response in the cache.
If a matching entry is found, it's a cache hit and the stored response is returned to the client. This avoids sending the same request to the upstream model endpoint twice, which would incur two quota charges. If no matching cache entry exists, the request is forwarded to the upstream endpoint, the response is stored in the cache using the generated key, and then returned to the client.
This cache hit/miss tracking is implemented with two integer variables, hits and misses. The TTL_MS variable sets the time-to-live for cached responses in milliseconds, defaulting to 60,000 milliseconds (1 minute) if not specified via an environment variable.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.