Summary

On August 21, 2026, DeepSeek released DeepSeek-V4-Flash-Vision-Exp on its API platform, an experimental multimodal variant that adds image understanding while matching V4-Flash on text tasks including agents, reasoning, and world knowledge. Images are billed at up to 384 tokens each with no price premium over V4-Flash rates, and DeepSeek shipped Harness 0.1.1 with native support the same day alongside a free Files API.

What changed

DeepSeek launched deepseek-v4-flash-vision-exp on the same API endpoint as V4-Flash, accepting images alongside text at existing V4-Flash token rates (images up to 384 tokens each). It says the model matches V4-Flash on text-only work and makes a major leap on multimodal agent benchmarks, approaching Opus 4.8. Same-day it released Harness 0.1.1 with out-of-the-box support and a free Files API (upload once, reference by file_id).

Why it matters

It pushes capable multimodal agent perception down to a low-cost tier with no pricing premium, pressuring frontier multimodal pricing and making vision-enabled agents cheaper to build. Bundling same-day Harness support and a Files API shows DeepSeek shipping the model and its agent runtime together rather than leaving integration to third parties.

Evidence excerpt

DeepSeek-V4-Flash-Vision-Exp is now live on the DeepSeek API Platform... matches DeepSeek-V4-Flash on text capabilities including agents, reasoning, and world knowledge. On multimodal agent benchmarks it makes a major leap, bringing performance close to Opus-4.8. Released alongside a free Files API and version 0.1.1 of the DeepSeek Harness agent framework with native support for the new model.

Sources