{
  "id": 4640038,
  "title": "DeepSeek’s first vision model vs. Gemini 3.7 Flash: It comes down to spend vs. speed",
  "url": "https://urgent.news/2026/08/31/deepseeks-first-vision-model-vs-gemini-3-7-flash-it-comes-down-to",
  "topic": "ai",
  "section": "AI",
  "published": "2026-08-31T12:00:00.000Z",
  "source": {
    "name": "The New Stack",
    "slug": "the-new-stack",
    "url": "https://thenewstack.io/deepseek-gemini-vision-comparison/"
  },
  "original_language": "en",
  "account": "DeepSeek and Google have both released new vision models, V4 Flash Vision Exp and Gemini 3.7 Flash, respectively. These models are designed to understand images, such as charts, invoices, and production logs, in addition to text. DeepSeek's model is priced at $0.22 per million input tokens and $0.66 per million output tokens, while Gemini's model costs $0.75 and $3.75 per million tokens respectively. Both models were tested on three image-based tasks: chart reading, invoice audit, and incident diagnosis.\n\nIn the chart reading test, the models correctly identified the quarter with exceeded costs, the revenue segment that grew every quarter, and estimated the company's total revenue. In the invoice audit test, both models correctly identified the incorrect line items and calculated the correct total due. During the incident diagnosis test, both models accurately pinpointed the root cause of the outage and recommended the first action to restore service. However, Gemini was significantly faster and more cost-effective than DeepSeek.\n\nAccuracy was identical for both models, with both providing correct answers to all questions. However, Gemini was faster, answering in an average of 7.2 seconds compared to DeepSeek's 16.8 seconds, and at a lower cost, with a total bill of $0.0122 compared to DeepSeek's $0.0039. These results highlight the importance of considering both cost and speed when choosing a vision model for image input tasks.",
  "summary": "DeepSeek released V4 Flash Vision Exp on August 21, its first model that accepts image input. Image input means a The post DeepSeek’s first vision model vs. Gemini 3.7 Flash: It comes down to spend vs. speed appeared first on The New Stack .",
  "key_points": [
    "DeepSeek's V4 Flash Vision Exp costs $0.22 per million input tokens, $0.66 per million output tokens",
    "Gemini 3.7 Flash costs $0.75 per million input tokens, $3.75 per million output tokens",
    "Gemini is faster and cheaper than DeepSeek in image-based tasks"
  ],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}