DeepSeek-v4-flash-vision-exp
Article URL: https://api-docs.deepseek.com/guides/vision/ Comments URL: https://news.ycombinator.com/item?id=49386163 Points: 340 # Comments: 112
The DeepSeek-v4-flash-vision-exp model enables the processing of images alongside textual inputs. This capability allows users to describe pictures, extract text from screenshots, analyze charts, and more. The model supports JPEG, PNG, GIF, and WebP image formats, with the desired format automatically determined by the file's content, not its name or MIME type.
To supply an image to the model, three distinct methods are available, all adhering to the standard OpenAI-compatible Chat Completions format. This format involves an array of blocks rather than a single string. These methods also extend to the Responses API, where images are embedded within input_image content parts. The base URL for these examples is https://api.deepseek.com.
Encoding the image and embedding it directly as a data: URL is the simplest approach for local files. However, this method counts toward the 48 MiB limit for the request body. If an external http(s) link is used instead, the model downloads the image. The URL must not exceed 8192 characters, with the image file limited to 32 MiB and must complete downloading within 60 seconds.
Alternatively, images can be uploaded once using the Files API, then referenced by file_id in subsequent requests. This method is preferable for images reused across multiple requests or when the inline limit of 48 MiB is exceeded. Images referenced via file_id can be up to 64 MiB and are not limited to 32 MiB per image. The file content block uses either a file_id or base64 encoded image data, but not both simultaneously.
For image_url inputs, an optional detail field can be set to adjust how the image is processed. Images counted as tokens in the request, with their dimensions determining the number of tokens. After resizing, each image is limited to 384 tokens, regardless of its original dimensions. When multiple images are included in a request, each is billed independently under this rule.
To gauge the token cost of an image of a specific size, refer to the image token calculator on the Token & Token Usage page. Regarding storage and upload quotas for files uploaded via the Files API, see the Files API: Limits.
Beyond the OpenAI-compatible endpoint, images can also be sent through the Anthropic-compatible /messages endpoint, accessible at https://api.deepseek.com/anthropic. The difference here lies in the shape of the image content block, which uses an image block with a source object containing a type of base64, url, or file. This mirrors the OpenAI methods, with the same limits and semantics applying.
Lastly, the deepseek-v4-flash-vision-exp model accepts images via the OpenAI-compatible Responses API, using input_image parts similar to the Chat Completions format. The same detail field semantics and restrictions apply, with only the content part shape differing.
Written by urgent.news from Hacker News Best's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.