DeepSeek adds vision to v4-flash-exp, accepts images up to 64 MiB
DeepSeek-v4-flash-vision-exp

DeepSeek's new deepseek-v4-flash-vision-exp model accepts images alongside text, supporting JPEG, PNG, GIF, and WebP. Images can be sent as base64 data URLs, external URLs, or via the Files API, with a max of 600 images per request. The model caps image tokens at 384 per image, resizing larger images to roughly 800x800.
Only vision models (deepseek-v4-flash-vision-exp) accept images; other models return a 400 error.
- ciberado
DS being unable to precisely view Playwright screenshots is the only thing I really miss from Sonnet. This is promising.
> Images are converted into tokens based on their dimensions, and these tokens are billed together with your text tokens.
> Before inference, every image is automatically resized:
> - Images with a total pixel count below roughly 384×384 are scaled up while preserving their aspect ratio.
> - Larger images are scaled down while preserving their aspect ratio so that the total pixel count after resizing is roughly that of an 800×800 image.
> As a result, there is an upper bound of 384 tokens per image: for example, a 2000×2000 image and a 5000×5000 image consume the same number of tokens after resizing. When a request contains multiple images, each image is counted independently under the same rule—there is no separate calculation for multi-image requests.
400 tokens per image results in 2,500 images per dollar, if I’m not mistaken.
edit: format.
- leumon
It fails the simple clock test for me which Qwen3.8 27B got (nearly) right.
given an image of a clock https://files.catbox.moe/kgwa5e.png
I asked it "what time does the clock show?" (both on reasoning: high)
DS answered: The clock shows *5:10* (and 45 seconds).
Here is the breakdown:
* *Hour hand (red, shortest):* Pointing at the *5*.
* *Minute hand (green, longest):* Pointing at the *2*, which represents 10 minutes.
* *Second hand (blue, medium):* Pointing at the *9*, which represents 45 seconds.
Qwen answered: The clock shows *8:10* (with the red second hand on the 5, i.e. *8:10:25*).
- *Hour hand* (short, blue) → 8
- *Minute hand* (long, green) → 2 (10 minutes)
- *Second hand* (thin, red) → 5 (25 seconds)
Correct answer is 08:09:25.
- LorenDB
I've heard that DeepSeek v4 Flash 0731 has frequently assumed that it has vision capabilities and then resorts to inventing text-based image analysis tools when it finds that it actually can't see. In that case, this is a great upgrade for the model.
Anecdotally, I had to tell 0731 to refrain from viewing screenshots since it kept breaking its sessions by trying to read images.
- cjg007
I've been using v4-flash without vision for this months — it's my go-to for code tasks. Now with vision, I'm wondering: if this model can do everything the text-only version does (plus see images), why keep the text-only one around?
Is it just cost/latency? Or is there something text-only does better?
- meetpateltech
News announcement with benchmarks: https://api-docs.deepseek.com/news/news260821/
- zmmmmm
> Larger images are scaled down while preserving their aspect ratio, so that the total pixel count after resizing is roughly that of an 800×800 image.
It's useful but for OCR and a lot of other applications it needs to be a bit higher (eg: putting in a full A4 / Letter sized page)
- BrucecarlL
Congratulations! DeepSeek has finally gained eyes — the dark days are about to be behind us.
- shangyu1994
Looks like multimodal training is really useful, app developers might need to consider adapter multimodal agents