$ cat wiki/models/deepseek-v4-flash-vision-exp.md
DeepSeek V4-Flash-Vision-Exp
modelupdated 2026-08-22created 2026-08-22
Spec
| Attribute | Value |
|---|---|
| Developer | DeepSeek |
| Released | 2026-08-21 (experimental / API access) |
| Announced | 2026-08-21 |
| Context window | unknown |
| Pricing | unknown |
| License | unknown |
| Availability | DeepSeek API |
Every unknown above is genuinely unstated in the coverage read, not omitted. | |
| The model is an experimental release; no model card, price schedule, weight | |
| drop or licence was surfaced. Its base, DeepSeek V4-Flash, is MIT | |
| open-weight, but nothing read says this vision variant ships under the same terms | |
| (source). |
Release Date
2026-08-21, accessible through the DeepSeek API. Framed by coverage as arriving while DeepSeek prepares for an IPO (source).
What it is
An experimental multimodal (vision) build of DeepSeek's flagship text-only DeepSeek V4-Flash — it can analyze and act on visual prompts such as images and screenshots. DeepSeek states it matches V4-Flash on text (agents, reasoning, world knowledge) and adds vision on top (source).
Benchmarks
Vendor-stated, as relayed by coverage — the model card was not readable from this environment. DeepSeek positions the model's multimodal agentic capability as "close to" Opus 4.8, and reports it beating Opus-4.8 on three benchmarks (source):
| Benchmark | Margin over Opus 4.8 (vendor-stated) |
|---|---|
| DeepSWE | +1.3 |
| Agents' Last Exam | +1.6 |
| ZeroBench | +1.0 |
| Per Eval Harness Configuration, these are vendor-run figures with **no | |
| published harness configuration** and no absolute scores given, so they are | |
| recorded as claims about a (model, harness) pair rather than as model properties. | |
| The comparison is a ~1-point margin on three benchmarks against a model roughly | |
| a year old (Claude Opus 4.8, 2026-05-28); Anthropic's current frontier is | |
| Claude Opus 5. |
Use Cases
- Screenshot / image-grounded agentic tasks (the multimodal agentic setting DeepSeek benchmarks)
- Cost-sensitive multimodal serving, if it inherits V4-Flash's pricing profile (unconfirmed)
Compared To
| Model | Modality | Open? |
|---|---|---|
| DeepSeek V4-Flash-Vision-Exp | text + vision (experimental) | unknown |
| DeepSeek V4-Flash | text | Yes (MIT) |
| Claude Opus 4.8 | text + vision | Proprietary |