DeepSeek-V4-Flash-Vision-Expの性能と評価に関するレポート
1. 総括
DeepSeek-V4-Flash-Vision-Expは、DeepSeek-V4-Flashのアーキテクチャに視覚モジュールを組み込んだ実験的マルチモーダルモデルだ。Hugging Face モデルカードによれば、既存モデルと比較してマルチモーダルエージェント能力が大幅に向上している。
このモデルの視覚機能は、エージェント向けのタスク実行に重点を置いて設計されている。チャートやソフトウェアのインターフェース、ウェブスクリーンショットなどの視覚情報を直接読み取ることが可能だ。テキストタスクの性能は、従来のモデルと同等である。
2. ブンリンベンチ(ブンリンラボ独自テスト)
データなし(実施予定)。手法はブンリンベンチについてを参照。
3. 各種ベンチマーク
Hugging Face モデルカードの数値に基づく。
| ベンチマーク | V4-Flash-Vision-Exp | V4-Flash-0731 | Opus-4.8 |
| :--- | :---: | :---: | :---: |
| テキストエージェント | | |
| Terminal Bench 2.1 | 83.9 | 82.7 | 85.0 |
| NL2Repo | 57.7 | 54.2 | 69.7 |
| Cybergym | 75.3 | 76.7 | 78.3 |
| DeepSWE | 59.3 | 54.4 | 58.0 |
| Toolathlon-Verified | 75.9 | 70.3 | 76.2 |
| DSBench-Hard | 63.6 | 59.6 | 71.7 |
| AutomationBench (Public) | 25.7 | 25.1 | 27.2 |
| マルチモーダルエージェント | | |
| ApexBench (Pass@1) | 36.5 | 26.2† | 39.4 |
| Agents' Last Exam | 27.3 | 25.2† | 25.7 |
| Chartography | 64.3 | - | 65.0 |
| ZeroBench (Pass@5) | 35.0 | - | 34.0 |
† ApexBench・Agents' Last Exam において、V4-Flash-0731 は入力内のマルチモーダル要素を無視して評価されている。
4. 公式の発表
2026-09-05:llama.cppは、Qwen3.8-Flash-NextとNemotron-3-Puzzleのサポート、動画入力オプションの追加などを発表した。
2026-09-02:FlashLabs株式会社は、コンテンツフィルタリングを緩和したGFUQ量子化版「DeepSeek-V4-Flash-Vision-Uncensored」を公開したと報じている。
2026-08-31:Hugging Faceにモデルの重みがアップロードされた。
2026-08-21:DeepSeekは、APIプラットフォームにてdeepseek-v4-flash-vision-expを先行公開した。
5. 実際の性能(コミュニティ評価)
Reddit r/LocalLLaMAのユーザーは、2枚のRTX 6000を用いた環境での動作を報告している。
antirez氏は、Mac M5 Max上で本モデルが高速に動作する様子を披露した。動作は速い。
画像入力は800×800ピクセルに圧縮される。そのため、細部の読み取りには不向きであるとの報告がある。
LM Studioの最新版は、DSparkなどの投機的デコーディング手法に対応している。これにより、生成の待ち時間を短縮できるという。
6. おすすめパラメータ
Hugging Face モデルカードの注釈による。
temperature = 1.0top_p = 0.95- reasoning effort:
max - エージェントフレームワーク: DeepSeek Harness (minimal mode)
7. 情報源
- Hugging Face モデルカード(公式)
- Deepseek v4 Flash Vision is out...(Reddit r/LocalLLaMA、2026-09-01)
- 画像を足したらテキストも伸びた。DeepSeekの新モデルがローカルAI勢をざわつかせた理由 - shiritomo(Google News: DeepSeek、2026-09-01)
- deepseek-ai/DeepSeek-V4-Flash-Vision-Exp · Hugging Face(Reddit r/LocalLLaMA、2026-09-01)
- DeepSeek-v4-flash-vision-exp(HN 498pt・155コメント)(HN Search (backfill)、2026-08-21)
- v0.4.0(llama.cpp、2026-09-05)
- FlashLabsがOrcaRouter向けにDeepSeek-V4-Flash-Vision-UncensoredのGGUF版を発表 - ニュースメディアVOIX(Google News: DeepSeek、2026-09-04)
- Testing Deepseek-V4-Flash-Vision-EXP on Dual RTX6000 build(Reddit r/LocalLLaMA、2026-09-03)
- OrcaRouter、セキュリティ研究向けモデル「DeepSeek-V4-Flash-Vision-Uncensored」のGGUF量子化版を公開 - ライブドアニュース(Google News: DeepSeek、2026-09-03)
- OrcaRouter、Google「Gemini 3.8 Flash」、Alibaba「Qwen3.8 Max」、Anthropic「Claude Fable 5.1」のフロンティアモデルを提供開始 - PR TIMES(Google News: Gemini、2026-09-03)
- OrcaRouter、セキュリティ研究向けモデル「DeepSeek-V4-Flash-Vision-Uncensored」のGGUF量子化版を公開(PR TIMES) - 毎日新聞(Google News: DeepSeek、2026-09-03)
- OrcaRouter、セキュリティ研究向けモデル「DeepSeek-V4-Flash-Vision-Uncensored」のGGUF量子化版を公開 - 時事ドットコム(Google News: DeepSeek、2026-09-03)
- OrcaRouter、セキュリティ研究向けモデル「DeepSeek-V4-Flash-Vision-Uncensored」のGGUF量子化版を公開 - Excite エキサイト(Google News: DeepSeek、2026-09-03)
- OrcaRouter、セキュリティ研究向けモデル「DeepSeek-V4-Flash-Vision-Uncensored」のGGUF量子化版を公開 - PR TIMES(Google News: DeepSeek、2026-09-03)
- OrcaRouter、セキュリティ研究向けモデル「DeepSeek-V4-Flash-Vision-Uncensored」のGGUF量子化版を公開 - 産経ニュース(Google News: DeepSeek、2026-09-03)
- Metaが「Muse Spark 1.3」を発表、ついにAnthropicやOpenAIの最先端モデルに追いつく性能を達成 - GIGAZINE(Google News: OpenAI、2026-09-03)
- Vision support merged for DeepSeek-V4-Flash-Vision-Exp(Reddit r/LocalLLaMA、2026-09-03)
- antirezのDS4に視覚が加わる、DeepSeek V4 FlashがM5 Maxのローカルで動く - Pasquale Pillitteri(Google News: DeepSeek、2026-09-02)
- DeepSeekが画像認識に対応した「DeepSeek-V4-Flash-Vision-Exp」をオープンモデルとして公開、画像認識を含むタスクでClaude Opus 4.8と同等の性能(GIGAZINE、2026-09-02)
- Anthropic、「Claude Fable 5.1」「Mythos 5.1」公開 AI透かしも導入 - Yahoo!ニュース(Google News: Anthropic、2026-09-02)
- Opus 4.8匹敵、画像対応「DeepSeek-V4-Flash-Vision-Exp」無償公開 (PC Watch) - Yahoo!ニュース(Google News: DeepSeek、2026-09-02)
- Opus 4.8匹敵、画像対応「DeepSeek-V4-Flash-Vision-Exp」無償公開 - Excite エキサイト(Google News: DeepSeek、2026-09-02)
- Got DeepSeek-V4-Flash-Vision running reliably on 2× RTX PRO 6000 Blackwell (SM120) with SGLang — had to patch 3 separate(Reddit r/LocalLLaMA、2026-09-02)
- Opus 4.8匹敵、画像対応「DeepSeek-V4-Flash-Vision-Exp」無償公開(PC Watch) - Yahoo!ニュース(Google News: DeepSeek、2026-09-02)
- Opus 4.8匹敵、画像対応「DeepSeek-V4-Flash-Vision-Exp」無償公開 - 千葉テレビ放送(Google News: DeepSeek、2026-09-02)
- Opus 4.8匹敵、画像対応「DeepSeek-V4-Flash-Vision-Exp」無償公開 - PC Watch(Google News: DeepSeek、2026-09-02)
- All currently popular local models in one table + Opus 4.8 results(Reddit r/LocalLLaMA、2026-09-01)
- DeepSeekの最初のオープンソースマルチモーダルモデル登場:この目は人間のために画像を見せるのではなく、エージェントのために働くため - news.aibase.com(Google News: DeepSeek、2026-09-01)
- LM StudioのAIエージェント「Bionic」、Linux版が登場(PC Watch) - Yahoo!ニュース(Google News: LM Studio、2026-09-01)
- 生成AIニュース 昨日【2026年8月31日(月)】のAI公式発表を3分でチェック - TECH NOISY(Google News: DeepSeek、2026-09-01)
- DeepSeekが画像認識に対応した「DeepSeek-V4-Flash-Vision-Exp」をオープンモデルとして公開、画像認識を含むタスクでClaude Opus 4.8と同等の性能 - au Webポータル(Google News: DeepSeek、2026-08-31)
- OrcaRouter、1.5TB超の大規模LLM「GLM-5.3」を1台のMacで実行できる「GLM-5.3-MLX」を公開 - PR TIMES(Google News: MLX、2026-08-31)
- OpenClaw 2026.8.1(OpenClaw、2026-08-31)
- 無料のLM Studio、DFlash/DSpark/MTPで推論を高速化(PC Watch) - Yahoo!ニュース(Google News: LM Studio、2026-08-28)
- 無料のLM Studio、DFlash/DSpark/MTPで推論を高速化 (PC Watch) - Yahoo!ニュース(Google News: LM Studio、2026-08-28)
- OrcaRouter、Alibaba発LLM「Qwen3.8-Flash-Next」をベースにしたセキュリティ研究向けモデル「Qwen3.8-Flash-Next-Uncensored」を公開 - PR TIMES(Google News: MLX、2026-08-28)
- OrcaRouter、320B級オープンウェイトLLM「GLM-5.3-Flash」のApple Silicon向けMLX量子化版を「2-bit Lite」〜「6-bit」の5バリアントで公開 - PR TIMES(Google News: MLX、2026-08-28)