DeepSeek V4.1 Currently 1500x Cheaper Than Usual
DeepSeek V4.1 Flash is currently listed at about 1,500 times cheaper than DeepSeek’s usual off-peak input rate on some third-party … Read More The post DeepSeek V4.1 Currently 1500x Cheaper Than Usual appeared first on ProPakistani .
DeepSeek's V4.1 Flash model is currently being offered at prices that are approximately 1,500 times lower than the company's usual off-peak input rate on certain third-party platforms. On OpenRouter, the model can be accessed for $0.0001 per million input tokens, while DeepSeek's standard off-peak input price is $0.15 per million tokens.
This significant price difference has sparked discussions on Reddit, with users highlighting Relace and Open Inference's "shockingly low" listed prices. Open Inference's listing on OpenRouter is around $0.00011 per million input tokens and $0.36 per million output tokens. The disparity in pricing applies specifically to input costs, not the overall pricing structure.
Relace still lists the output cost at $0.60 per million tokens, which aligns with DeepSeek's off-peak output rate. The price gap is attributed to various factors such as cache hits, routing, latency, and the number of output tokens utilized in a job. DeepSeek's V4.1 Flash model is a sparse mixture-of-experts model released in September 2026, built on the company's Causal Encoder-Decoder architecture.
With 552 billion parameters, it activates 8 billion parameters on input and 16 billion on output. Designed for coding, terminal work, computer-use agents, and long-context tasks, the model boasts a 1 million token context window. OpenRouter lists a context window of 1 million tokens. The reduced input prices on third-party platforms are not indicative of a full price overhaul, as output costs remain relatively close to DeepSeek's off-peak rates.
While the lower input prices may be appealing for high-volume workloads, output costs still remain near the standard DeepSeek off-peak price of $0.60 per million tokens. Developers considering these third-party endpoints must weigh the potential cost savings against factors such as reliability and speed. DeepSeek has previously adjusted pricing, including raising costs by up to 4 times in the past.
For now, the practical takeaway is narrow and easily verifiable. On current OpenRouter listings, Relace and Open Inference are advertising DeepSeek V4.1 Flash input rates near $0.0001 per million tokens, roughly 1,500 times lower than DeepSeek's off-peak input price of $0.15 per million tokens. Output pricing remains closer to the standard rates.
Written by urgent.news from ProPakistani's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.