DeepSeek Launches V4.1-Flash With Lower Memory and API Costs
DeepSeek V4.1-Flash promises lower memory use and API costs, but buyers should test its performance, compatibility and total deployment expenses. The post DeepSeek Launches V4.1-Flash With Lower Memory and API Costs appeared first on TechRepublic .
DeepSeek has unveiled V4.1-Flash, its latest AI model within the V4.1 architecture family. This new model promises lower memory usage and reduced API costs while maintaining strong performance. At 552 billion total parameters, V4.1-Flash activates approximately 8 billion parameters per input token and 16 billion per output token, significantly fewer than its predecessors.
Operating on a novel Causal Encoder-Decoder architecture enhanced with new pretraining methods and reinforcement learning, the model delivers superior results without utilizing the full model for each request. Featuring a one-million-token context window, V4.1-Flash is particularly efficient for long-running conversations and agent workloads.
Benchmarks show V4.1-Flash outperforming V4 Pro and several competitors on coding, cybersecurity, and agent evaluations. For example, it scored 90.6 on Terminal-Bench 2.1, exceeding GPT-5.6 Sol's 88.8, Kimi K3's 88.3, and V4 Pro's 87.9. Additionally, V4.1-Flash scored 88.1 on Cybergym, surpassing V4 Pro's 83.3 and Kimi K3's 80. Despite these improvements, DeepSeek cautions that the model's performance might differ in real-world deployments and advises users to test it thoroughly before adopting it.
DeepSeek has also slashed API pricing, offering off-peak cached input as low as 0.02 yuan per million tokens. To streamline operations, starting September 14, DeepSeek will route V4 Pro requests to V4.1-Flash at the Flash rate until V4.1-Pro becomes available, retiring older V4-Flash and V4-Flash-Vision-Exp endpoints. Organizations using these endpoints should test V4.1-Flash prior to the routing change, especially if their applications rely on consistent output formats, latency needs, or a specific model version.
Written by urgent.news from TechRepublic's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.
This story
This is one outlet's version. Read the fullest account.
- DeepSeek V4.1 Flash Launches With Lower Prices and Native Vision propakistani.pk