DeepSeek adds vision to v4-flash-exp, accepts images up to 64 MiB

DeepSeek-v4-flash-vision-exp

DeepSeek adds vision to v4-flash-exp, accepts images up to 64 MiB

DeepSeek's new deepseek-v4-flash-vision-exp model accepts images alongside text, supporting JPEG, PNG, GIF, and WebP. Images can be sent as base64 data URLs, external URLs, or via the Files API, with a max of 600 images per request. The model caps image tokens at 384 per image, resizing larger images to roughly 800x800.

Only vision models (deepseek-v4-flash-vision-exp) accept images; other models return a 400 error.
  1. ciberado

    DS being unable to precisely view Playwright screenshots is the only thing I really miss from Sonnet. This is promising.

    > Images are converted into tokens based on their dimensions, and these tokens are billed together with your text tokens.

    > Before inference, every image is automatically resized:

    > - Images with a total pixel count below roughly 384×384 are scaled up while preserving their aspect ratio.

    > - Larger images are scaled down while preserving their aspect ratio so that the total pixel count after resizing is roughly that of an 800×800 image.

    > As a result, there is an upper bound of 384 tokens per image: for example, a 2000×2000 image and a 5000×5000 image consume the same number of tokens after resizing. When a request contains multiple images, each image is counted independently under the same rule—there is no separate calculation for multi-image requests.

    400 tokens per image results in 2,500 images per dollar, if I’m not mistaken.

    edit: format.

  2. leumon

    It fails the simple clock test for me which Qwen3.8 27B got (nearly) right.

    given an image of a clock https://files.catbox.moe/kgwa5e.png

    I asked it "what time does the clock show?" (both on reasoning: high)

    DS answered: The clock shows *5:10* (and 45 seconds).

    Here is the breakdown:

    * *Hour hand (red, shortest):* Pointing at the *5*.

    * *Minute hand (green, longest):* Pointing at the *2*, which represents 10 minutes.

    * *Second hand (blue, medium):* Pointing at the *9*, which represents 45 seconds.

    Qwen answered: The clock shows *8:10* (with the red second hand on the 5, i.e. *8:10:25*).

    - *Hour hand* (short, blue) → 8

    - *Minute hand* (long, green) → 2 (10 minutes)

    - *Second hand* (thin, red) → 5 (25 seconds)

    Correct answer is 08:09:25.

  3. LorenDB

    I've heard that DeepSeek v4 Flash 0731 has frequently assumed that it has vision capabilities and then resorts to inventing text-based image analysis tools when it finds that it actually can't see. In that case, this is a great upgrade for the model.

    Anecdotally, I had to tell 0731 to refrain from viewing screenshots since it kept breaking its sessions by trying to read images.

  4. cjg007

    I've been using v4-flash without vision for this months — it's my go-to for code tasks. Now with vision, I'm wondering: if this model can do everything the text-only version does (plus see images), why keep the text-only one around?

    Is it just cost/latency? Or is there something text-only does better?

  5. meetpateltech

    News announcement with benchmarks: https://api-docs.deepseek.com/news/news260821/

  6. zmmmmm

    > Larger images are scaled down while preserving their aspect ratio, so that the total pixel count after resizing is roughly that of an 800×800 image.

    It's useful but for OCR and a lot of other applications it needs to be a bit higher (eg: putting in a full A4 / Letter sized page)

  7. BrucecarlL

    Congratulations! DeepSeek has finally gained eyes — the dark days are about to be behind us.

  8. shangyu1994

    Looks like multimodal training is really useful, app developers might need to consider adapter multimodal agents

More from this day

2026-08-21