Report on the performance and evaluation of DeepSeek-V4-Flash-Vision-Exp
1. Summary
DeepSeek-V4-Flash-Vision-Exp is an experimental multimodal model that incorporates a vision module into the DeepSeek-V4-Flash architecture. According to the Hugging Face Model Card, multimodal agent capabilities have been significantly improved compared to existing models.
The vision capabilities of this model are designed with a focus on task execution for agents. It is capable of directly reading visual information such as charts, software interfaces, and web screenshots. Performance on text tasks is reported to be equivalent to conventional models.
2. Bunrin Bench (BUNRIN LABO Original Test)
No data available (scheduled). For the methodology, refer to About Bunrin Bench.
3. Various Benchmarks
Based on figures from the Hugging Face Model Card.
| Benchmark | V4-Flash-Vision-Exp | V4-Flash-0731 | Opus-4.8 |
| :--- | :---: | :---: | :---: |
| Text Agent | | |
| Terminal Bench 2.1 | 83.9 | 82.7 | 85.0 |
| NL2Repo | 57.7 | 54.2 | 69.7 |
| Cybergym | 75.3 | 76.7 | 78.3 |
| DeepSWE | 59.3 | 54.4 | 58.0 |
| Toolathlon-Verified | 75.9 | 70.3 | 76.2 |
| DSBench-Hard | 63.6 | 59.6 | 71.7 |
| AutomationBench (Public) | 25.7 | 25.1 | 27.2 |
| Multimodal Agent | | |
| ApexBench (Pass@1) | 36.5 | 26.2† | 39.4 |
| Agents' Last Exam | 27.3 | 25.2† | 25.7 |
| Chartography | 64.3 | - | 65.0 |
| ZeroBench (Pass@5) | 35.0 | - | 34.0 |
† In ApexBench and Agents' Last Exam, V4-Flash-0731 was evaluated by ignoring multimodal elements within the input.
4. Official announcements
2026-09-05: llama.cpp announced support for Qwen3.8-Flash-Next and Nemotron-3-Puzzle, as well as the addition of video input options.
2026-09-02: FlashLabs Inc. reported that it has released "DeepSeek-V4-Flash-Vision-Uncensored," a GFUQquantized build with relaxed content filtering.
2026-08-31: Model weights were uploaded to Hugging Face.
2026-08-21: DeepSeek released deepseek-v4-flash-vision-exp in preview on its API platform.
5. Real-world performance (Community reception)
A user on Reddit r/LocalLLaMA reported operation in an environment using two RTX 6000 GPUs.
antirez demonstrated the model running at high speeds on a Mac M5 Max. The operation is fast.
Image inputs are compressed to 800x800 pixels. Consequently, there are reports that it is not suitable for reading fine details.
The latest version of LM Studio supports speculative decoding methods such as DSpark. This is said to reduce generation latency.
6. Recommended parameters
According to annotations in the Hugging Face Model Card:
temperature = 1.0top_p = 0.95- reasoning effort:
max - Agent framework: DeepSeek Harness (minimal mode)
7. Sources
- Hugging Face Model Card (Official)
- Deepseek v4 Flash Vision is out... (Reddit r/LocalLLaMA, 2026-09-01)
- Adding images expanded the text. Why DeepSeek's new model stirred the local AI community - shiritomo (Google News: DeepSeek, 2026-09-01)
- deepseek-ai/DeepSeek-V4-Flash-Vision-Exp · Hugging Face (Reddit r/LocalLLaMA, 2026-09-01)
- DeepSeek-v4-flash-vision-exp (HN 498pt, 155 comments) (HN Search (backfill), 2026-08-21)
- v0.4.0 (llama.cpp, 2026-09-05)
- FlashLabs announces GGUF version of DeepSeek-V4-Flash-Vision-Uncensored for OrcaRouter - News Media VOIX (Google News: DeepSeek, 2026-09-04)
- Testing Deepseek-V4-Flash-Vision-EXP on Dual RTX6000 build (Reddit r/LocalLLaMA, 2026-09-03)
- OrcaRouter releases GGUF quantized build of "DeepSeek-V4-Flash-Vision-Uncensored" model for security research - Livedoor News (Google News: DeepSeek, 2026-09-03)
- OrcaRouter begins providing frontier models: Google "Gemini 3.8 Flash," Alibaba "Qwen3.8 Max," and Anthropic "Claude Fable 5.1" - PR TIMES (Google News: Gemini, 2026-09-03)
- OrcaRouter releases GGUF quantized build of "DeepSeek-V4-Flash-Vision-Uncensored" model for security research (PR TIMES) - Mainichi Shimbun (Google News: DeepSeek, 2026-09-03)
- OrcaRouter releases GGUF quantized build of "DeepSeek-V4-Flash-Vision-Uncensored" model for security research - Jiji.com (Google News: DeepSeek, 2026-09-03)
- OrcaRouter releases GGUF quantized build of "DeepSeek-V4-Flash-Vision-Uncensored" model for security research - Excite (Google News: DeepSeek, 2026-09-03)
- OrcaRouter releases GGUF quantized build of "DeepSeek-V4-Flash-Vision-Uncensored" model for security research - PR TIMES (Google News: DeepSeek, 2026-09-03)
- OrcaRouter releases GGUF quantized build of "DeepSeek-V4-Flash-Vision-Uncensored" model for security research - Sankei News (Google News: DeepSeek, 2026-09-03)
- Meta announces "Muse Spark 1.3," finally achieving performance that catches up to cutting-edge models from Anthropic and OpenAI - GIGAZINE (Google News: OpenAI, 2026-09-03)
- Vision support merged for DeepSeek-V4-Flash-Vision-Exp (Reddit r/LocalLLaMA, 2026-09-03)
- Vision added to antirez's DS4, DeepSeek V4 Flash runs locally on M5 Max - Pasquale Pillitteri (Google News: DeepSeek, 2026-09-02)
- DeepSeek releases "DeepSeek-V4-Flash-Vision-Exp" with image recognition as an open model, achieving performance equivalent to Claude Opus 4.8 in tasks including image recognition (GIGAZINE, 2026-09-02)
- Anthropic releases "Claude Fable 5.1" and "Mythos 5.1," introducing AI watermarking - Yahoo! News (Google News: Anthropic, 2026-09-02)
- Comparable to Opus 4.8, image-capable "DeepSeek-V4-Flash-Vision-Exp" released for free (PC Watch) - Yahoo! News (Google News: DeepSeek, 2026-09-02)
- Comparable to Opus 4.8, image-capable "DeepSeek-V4-Flash-Vision-Exp" released for free - Excite (Google News: DeepSeek, 2026-09-02)
- Got DeepSeek-V4-Flash-Vision running reliably on 2× RTX PRO 6000 Blackwell (SM120) with SGLang — had to patch 3 separate (Reddit r/LocalLLaMA, 2026-09-02)
- Comparable to Opus 4.8, image-capable "DeepSeek-V4-Flash-Vision-Exp" released for free (PC Watch) - Yahoo! News (Google News: DeepSeek, 2026-09-02)
- Comparable to Opus 4.8, image-capable "DeepSeek-V4-Flash-Vision-Exp" released for free - Chiba TV (Google News: DeepSeek, 2026-09-02)
- Comparable to Opus 4.8, image-capable "DeepSeek-V4-Flash-Vision-Exp" released for free - PC Watch (Google News: DeepSeek, 2026-09-02)
- All currently popular local models in one table + Opus 4.8 results (Reddit r/LocalLLaMA, 2026-09-01)
- DeepSeek's first open-source multimodal model arrives: these eyes are not for showing images to humans, but for working for agents - news.aibase.com (Google News: DeepSeek, 2026-09-01)
- LM Studio's AI agent "Bionic" Linux version released (PC Watch) - Yahoo! News (Google News: LM Studio, 2026-09-01)
- Generative AI News Yesterday [Monday, August 31, 2026]: Check official AI announcements in 3 minutes - TECH NOISY (Google News: DeepSeek, 2026-09-01)
- DeepSeek releases "DeepSeek-V4-Flash-Vision-Exp" with image recognition as an open model, achieving performance equivalent to Claude Opus 4.8 in tasks including image recognition - au Web Portal (Google News: DeepSeek, 2026-08-31)
- OrcaRouter releases "GLM-5.3-MLX," enabling a single Mac to run the massive 1.5TB+ LLM "GLM-5.3" - PR TIMES (Google News: MLX, 2026-08-31)
- OpenClaw 2026.8.1 (OpenClaw, 2026-08-31)
- Free LM Studio accelerates inference with DFlash/DSpark/MTP (PC Watch) - Yahoo! News (Google News: LM Studio, 2026-08-28)
- Free LM Studio accelerates inference with DFlash/DSpark/MTP (PC Watch) - Yahoo! News (Google News: LM Studio, 2026-08-28)
- OrcaRouter releases "Qwen3.8-Flash-Next-Uncensored," a security research model based on Alibaba's LLM "Qwen3.8-Flash-Next" - PR TIMES (Google News: MLX, 2026-08-28)
- OrcaRouter releases MLX quantized build for Apple Silicon of the 320B-class open-weight LLM "GLM-5.3-Flash" in 5 variants from "2-bit Lite" to "6-bit" - PR TIMES (Google News: MLX, 2026-08-28)