Catalog
Googlerouted

Gemma 4 26B A4B

Gemma 4 26B A4B is Google DeepMind's fast mixture-of-experts Gemma (Apache 2.0): 25.2B parameters with 3.8B active per token. It handles reasoning, coding and agentic workflows with native function calling, and reads documents, scans, charts and screenshots as images. Google states out-of-the-box support for more than 35 languages; thinking is optional and off unless you enable it. It is fulfilled through global providers under our European contract, and requests go only to backends that do not retain prompts or completions.

model id: gemma-4-26b-a4b-it

from openai import OpenAI client = OpenAI(    api_key="hb-...",    base_url="https://api.heabsy.com/v1",) r = client.chat.completions.create(    model="gemma-4-26b-a4b-it",    messages=[{"role": "user", "content": "Hello"}],)print(r.choices[0].message.content)
Context
256K tokens
Serving
Global providers · EU contract
Tool calling
Reasoning
Structured output
Vision
Input
$0.07/ 1M
Cached input
$0.07/ 1M
Output
$0.34/ 1M

Price list, updated 12 September 2026. No minimum spend, no subscription.

About the model

Developer
Google DeepMind
Licence
Apache-2.0
Architecture
mixture of experts · 25.2B parameters, 3.8B active per token
Input
text, images
Reasoning
switchable per request
Knowledge cutoff
January 2025
Languages named by the developer
35+ languages

Source: the developer's model card, as of 17 September 2026.

Gemma 4 31B or 26B A4B?

Both come from Google DeepMind under Apache 2.0, read images as well as text, support more than 35 languages out of the box and keep thinking off unless you enable it.

Gemma 4 31B is dense (30.7B parameters) and scores higher on the card: 85.2 % on MMLU Pro and 80.0 % on LiveCodeBench v6. Gemma 4 26B A4B is a mixture of experts with 3.8B of 25.2B parameters active, built for speed, at 82.6 % and 77.1 %. Gemma-4-26B-A4B Uncensored is a community build of the 26B with refusals removed.

Licence: what it means for your company

Apache-2.0. A permissive open-source licence: you may use, modify and sell products built on the model, including commercially; when you redistribute the weights, the licence and copyright notices go with them.

A summary of the licence, not legal advice.

Benchmarks reported by the developer

MMLU Pro
82.6 %
GPQA Diamond
82.3 %
MMMLU
86.3 %
LiveCodeBench v6
77.1 %

Figures from the developer's model card, not our measurement. Source: huggingface.co/google/gemma-4-26B-A4B-it

Technical details

Context window
256K tokens
Max output
65,536 tokens
Throughput limit
200K tok/min
Tokenizer
Other
Streaming
yes
Open weights
google/gemma-4-26b-a4b-it
Supported parameters
max_tokens
1 … 65536
temperature
0 … 2
top_p
0 … 1
frequency_penalty
-2 … 2
presence_penalty
-2 … 2
stop
array
seed
integer
Bought for you, billed by us

This model is fulfilled through global providers under our European contract: one DPA with an EU company, one invoice for every model in the catalog, no US counterparty on your paperwork. Compute may run outside the EEA — if you need strict EEA-only processing, use our own-hardware model.

Questions about this model

Where is Gemma 4 26B A4B computed?

Outside the EEA, through a global provider under our European contract. You get one DPA with an EU company and one invoice, but the compute itself does not run in the EEA — we state that instead of hiding it.

Does Gemma 4 26B A4B store prompts and completions?

No. This model runs with zero data retention: request content is not written to durable storage and nothing is used for training.

How large is the Gemma 4 26B A4B context window?

256K tokens, with up to 65,536 tokens of output.

What does Gemma 4 26B A4B cost?

$0.07 per million input tokens and $0.34 per million output tokens. No minimum spend and no subscription; the same price list covers the API and the invoice.

Is there a throughput limit on Gemma 4 26B A4B?

Yes, 200K tokens per minute. Higher limits are a question of volume, not of principle.

Which languages does Gemma 4 26B A4B support?

The developer states support for more than 35 languages without listing them. Whether your language is among them is not documented — test on your own material first.

Under which licence is Gemma 4 26B A4B released?

Apache-2.0. A permissive open-source licence: you may use, modify and sell products built on the model, including commercially; when you redistribute the weights, the licence and copyright notices go with them.

Can reasoning be switched off in Gemma 4 26B A4B?

Yes. Per the developer, thinking can be switched on or off for each request.

Compare with similar models

ModelDeveloperParametersContextLicenceInput / 1M
Gemma 4 26B A4BGoogle DeepMind25.2B / 3.8B256KApache-2.0$0.07
Gemma 4 31BGoogle DeepMind30.7B256KApache-2.0$0.14
Mistral Small 24B (2501)Mistral AI24B32KApache-2.0$0.05
Cydonia 24B v4.1TheDrummer24B128K$0.42

30 more models behind the same key and the same invoice.

Browse the catalog