DeepSeek-V4-Flash-Vision-Exp

deepseek

DeepSeek-V4-Flash-Vision-Exp is DeepSeek's experimental multimodal model in the V4-Flash family, adding visual understanding to the V4-Flash architecture. It serves a 1M-token (1,048,576) context window and supports image input with visual grounding, tool calling, structured/JSON output, and configurable reasoning effort (low/high/max, or disabled).

Context
1.0M
Max output
384K
Input price
$0.44/1M tokens
Output price
$1.32/1M tokens

Capabilities

vision
tool call
structured output
reasoning
json mode
streaming
fine tuning
batch

Details

Model IDdeepseek-ai/DeepSeek-V4-Flash-Vision-Exp
Provider DeepInfra
Creator DeepSeek
Family—
Licensedeepseek
Parameters—
Statusactive
Input modalitiestext, image
Output modalitiestext
Architecture—
Knowledge cutoff—
Training data cutoff—
Release date—
Deprecation date—
Typechat
Reasoning tokensYes
Max input—
Open weightYes
Sourceofficial
Last updated

Tools

Function Callingfunction_calling
Call external functions and APIs

Endpoints

Chat CompletionsPOST
Generate chat responses with messageshttps://api.deepinfra.com/v1/openai/v1/chat/completions

Pricing

Input
$0.44
Output
$1.32
Cache write
—
Cache read
$0.01
↓90%$0.14
Batch in
—
Batch out
—
Price history · $/1M cache read
$0.14
$0.014
Sep 3Sep 10
DateInputOutputCache readChange
Sep 3, 2026$0.44$1.32$0.14—
Sep 10, 20268e0a0b7$0.44$1.32$0.014↓90%

Cache read is ↓90% since Sep 3, 2026, from $0.14 to $0.014 per 1M tokens.

Price across providers

ProviderInputOutputSince launch
DeepInfracheapest$0.44$1.32—
Fireworks AI$1$1↑355%

API

GET/v1/models/deepinfra/deepseek-ai/DeepSeek-V4-Flash-Vision-Exp