Introducing DeepSeek V4.1 FlashDeepSeekReleased September 10, 2026

DeepSeek V4.1 Flash

DeepSeek current native-multimodal Flash model with an asymmetric 552B-parameter MoE architecture.

multimodal-llmOpen WeightsPeak: USD 0.30 / 1M input tok; USD 1.20 / 1M output tokContext: 1M (1,000,000 tokens)

Technical Specifications

Architecture Type
Causal Encoder-Decoder MoE
Total Parameters
552B
Active Parameters (MoE)
8B input / 16B output
Context Window
1M (1,000,000 tokens)
Max Output Tokens
384K (384,000 tokens)
Knowledge Cutoff
Frontier (Current)
Supported Modalities
text, vision, reasoning
License & Access
Open Weights

Deep Architectural Overview

DeepSeek V4.1 Flash is called with deepseek-flash. It supports thinking and non-thinking modes, native visual understanding, a 1M-token context window, and a 384K maximum output.

Strengths & Considerations

Core Strengths
  • Native visual understanding
  • 1M-token context
  • 8B active input and 16B active output parameters

Token & API Pricing

Input Tokens (1M)$0.30
Output Tokens (1M)$1.20
Cached Input (1M)$0.0060
Pricing is verified directly against DeepSeek's developer documentation and API rate sheets.

Similar & Alternative Models

Explore other frontier models from DeepSeek and comparable reasoning engines.

Browse all models
DeepSeek

DeepSeek V4 Pro API model, version DeepSeek-V4-Pro-0813, with thinking and non-thinking modes.

1M ctxPeak: USD 1.32 / 1M input tok; USD 3.96 / 1M output tok