{
  "id": 6874024,
  "title": "DeepSeek Launches V4.1-Flash With Lower Memory and API Costs",
  "url": "https://urgent.news/2026/09/11/deepseek-launches-v4-1-flash-with-lower-memory-and-api-costs",
  "topic": "ai",
  "section": "AI",
  "published": "2026-09-11T18:10:32.000Z",
  "source": {
    "name": "TechRepublic",
    "slug": "techrepublic",
    "url": "https://www.techrepublic.com/article/news-deepseek-v4-1-flash-costs-apac-china/"
  },
  "original_language": "en",
  "account": "DeepSeek has unveiled V4.1-Flash, its latest AI model within the V4.1 architecture family. This new model promises lower memory usage and reduced API costs while maintaining strong performance. At 552 billion total parameters, V4.1-Flash activates approximately 8 billion parameters per input token and 16 billion per output token, significantly fewer than its predecessors. Operating on a novel Causal Encoder-Decoder architecture enhanced with new pretraining methods and reinforcement learning, the model delivers superior results without utilizing the full model for each request. Featuring a one-million-token context window, V4.1-Flash is particularly efficient for long-running conversations and agent workloads. Benchmarks show V4.1-Flash outperforming V4 Pro and several competitors on coding, cybersecurity, and agent evaluations. For example, it scored 90.6 on Terminal-Bench 2.1, exceeding GPT-5.6 Sol's 88.8, Kimi K3's 88.3, and V4 Pro's 87.9. Additionally, V4.1-Flash scored 88.1 on Cybergym, surpassing V4 Pro's 83.3 and Kimi K3's 80. Despite these improvements, DeepSeek cautions that the model's performance might differ in real-world deployments and advises users to test it thoroughly before adopting it. DeepSeek has also slashed API pricing, offering off-peak cached input as low as 0.02 yuan per million tokens. To streamline operations, starting September 14, DeepSeek will route V4 Pro requests to V4.1-Flash at the Flash rate until V4.1-Pro becomes available, retiring older V4-Flash and V4-Flash-Vision-Exp endpoints. Organizations using these endpoints should test V4.1-Flash prior to the routing change, especially if their applications rely on consistent output formats, latency needs, or a specific model version.",
  "summary": "DeepSeek V4.1-Flash promises lower memory use and API costs, but buyers should test its performance, compatibility and total deployment expenses. The post DeepSeek Launches V4.1-Flash With Lower Memory and API Costs appeared first on TechRepublic .",
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 2,
    "also_reported_by": [
      {
        "outlet": "ProPakistani",
        "title": "DeepSeek V4.1 Flash Launches With Lower Prices and Native Vision",
        "url": "https://urgent.news/2026/09/11/deepseek-v4-1-flash-launches-with-lower-prices-and-native-vision",
        "published": "2026-09-11T13:56:33.000Z"
      }
    ]
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}