DeepSeek-V4-Flash-Vision-Exp is DeepSeek's experimental multimodal model in the V4-Flash family, adding visual understanding to the V4-Flash architecture. It serves a 1M-token (1,048,576) context window and supports image input with visual grounding, tool calling, structured/JSON output, and configurable reasoning effort (low/high/max, or disabled).
function_callinghttps://api.deepinfra.com/v1/openai/v1/chat/completions| Date | Input | Output | Cache read | Change |
|---|---|---|---|---|
| Sep 3, 2026 | $0.44 | $1.32 | $0.14 | — |
| Sep 10, 20268e0a0b7 | $0.44 | $1.32 | $0.014 | ↓90% |
Cache read is ↓90% since Sep 3, 2026, from $0.14 to $0.014 per 1M tokens.
| Provider | Input | Output | Since launch |
|---|---|---|---|
| DeepInfracheapest | $0.44 | $1.32 | — |
| Fireworks AI | $1 | $1 | ↑355% |
/v1/models/deepinfra/deepseek-ai/DeepSeek-V4-Flash-Vision-ExpDeepSeek-V4-Flash-Vision-Exp is DeepSeek's experimental multimodal model in the V4-Flash family, adding visual understanding to the V4-Flash architecture. It serves a 1M-token (1,048,576) context window and supports image input with visual grounding, tool calling, structured/JSON output, and configurable reasoning effort (low/high/max, or disabled).