Catalog
DeepSeekrouted

DeepSeek-V4.1-Flash

DeepSeek-V4.1-Flash (MIT licence, September 2026) is a multimodal mixture-of-experts model and a different network from V4-Flash: a 552B-parameter backbone that activates 8B per token while reading the prompt and 16B while writing. DeepSeek designed it for input-heavy agentic work — code agents, tool use, agents that look at screens — keeping memory use small even at million-token context. Reasoning effort is a continuous dial from 1 to 100 rather than a few fixed levels. It is fulfilled through global providers under our European contract, and requests go only to backends that do not retain prompts or completions.

model id: deepseek-v4.1-flash

from openai import OpenAI client = OpenAI(    api_key="hb-...",    base_url="https://api.heabsy.com/v1",) r = client.chat.completions.create(    model="deepseek-v4.1-flash",    messages=[{"role": "user", "content": "Hello"}],)print(r.choices[0].message.content)
Context
1M tokens
Serving
Global providers · EU contract
Tool calling
Reasoning
Structured output
Vision
Input
$0.15/ 1M
Output
$0.60/ 1M

Price list, updated 12 September 2026. No minimum spend, no subscription.

About the model

Developer
DeepSeek
Released
September 2026
Licence
MIT
Architecture
mixture of experts · 552B parameters, 8–16B active per token
Input
text, images
Reasoning
selectable effort
Languages named by the developer
none named on the model card

Source: the developer's model card, as of 17 September 2026.

Which DeepSeek model to choose

DeepSeek V4 Flash is the April 2026 preview. DeepSeek-V4-Flash-0731 is its official release (284B parameters, 13B active, text only) and the sensible default for code and tool agents at moderate cost.

DeepSeek-V4-Pro-0813 is the flagship: 1.6T parameters with 49B active. DeepSeek recommends allowing up to 384K output tokens at high and max effort — our maximum output for this model is listed under technical details.

DeepSeek-V4.1-Flash (September 2026) is a different network, not a point release: 552B parameters, 8B active while reading and 16B while writing, and it accepts images. Take it for input-heavy agents; keep the preview only if you need to reproduce earlier results.

All four are open weights under the MIT licence.

Licence: what it means for your company

MIT. A permissive open-source licence: you may use, modify and sell products built on the model, including commercially; when you redistribute the weights, the licence and copyright notices go with them.

A summary of the licence, not legal advice.

Technical details

Context window
1M tokens
Max output
65,536 tokens
Throughput limit
200K tok/min
Tokenizer
DeepSeek
Streaming
yes
Open weights
deepseek-ai/DeepSeek-V4.1-Flash
Supported parameters
max_tokens
1 … 65536
temperature
0 … 2
top_p
0 … 1
frequency_penalty
-2 … 2
presence_penalty
-2 … 2
stop
array
seed
integer
Bought for you, billed by us

This model is fulfilled through global providers under our European contract: one DPA with an EU company, one invoice for every model in the catalog, no US counterparty on your paperwork. Compute may run outside the EEA — if you need strict EEA-only processing, use our own-hardware model.

Questions about this model

Where is DeepSeek-V4.1-Flash computed?

Outside the EEA, through a global provider under our European contract. You get one DPA with an EU company and one invoice, but the compute itself does not run in the EEA — we state that instead of hiding it.

Does DeepSeek-V4.1-Flash store prompts and completions?

No. This model runs with zero data retention: request content is not written to durable storage and nothing is used for training.

How large is the DeepSeek-V4.1-Flash context window?

1M tokens, with up to 65,536 tokens of output.

What does DeepSeek-V4.1-Flash cost?

$0.15 per million input tokens and $0.60 per million output tokens. No minimum spend and no subscription; the same price list covers the API and the invoice.

Is there a throughput limit on DeepSeek-V4.1-Flash?

Yes, 200K tokens per minute. Higher limits are a question of volume, not of principle.

Which languages does DeepSeek-V4.1-Flash support?

The developer's model card names no languages. Test the model on your own material before relying on it in a particular language.

Under which licence is DeepSeek-V4.1-Flash released?

MIT. A permissive open-source licence: you may use, modify and sell products built on the model, including commercially; when you redistribute the weights, the licence and copyright notices go with them.

Can reasoning be switched off in DeepSeek-V4.1-Flash?

The developer lets you choose how much the model reasons; the card does not document turning reasoning off completely.

Compare with similar models

ModelDeveloperParametersContextLicenceInput / 1M
DeepSeek-V4.1-FlashDeepSeek552B / 8–161MMIT$0.15
DeepSeek-V4-Flash-0731DeepSeek284B / 13B1.3MMIT$0.06
DeepSeek V4 FlashDeepSeek284B / 13B1MMIT$0.25
DeepSeek-V4-Pro-0813DeepSeek1.6T / 49B1MMIT$0.80

30 more models behind the same key and the same invoice.

Browse the catalog