Catalog
Metarouted

Llama 3.1 8B Instruct

Llama 3.1 8B Instruct is Meta's small dense model from July 2024, released under the Llama 3.1 Community License. It is an assistant-style chat model with official support for eight languages — English, German, French, Italian, Portuguese, Hindi, Spanish and Thai — and no thinking mode. An older generation, but widely integrated and cheap: a baseline for simple chat, rewriting and classification. It is fulfilled through global providers under our European contract, and requests go only to backends that do not retain prompts or completions.

model id: llama-3.1-8b-instruct

from openai import OpenAI client = OpenAI(    api_key="hb-...",    base_url="https://api.heabsy.com/v1",) r = client.chat.completions.create(    model="llama-3.1-8b-instruct",    messages=[{"role": "user", "content": "Hello"}],)print(r.choices[0].message.content)
Context
128K tokens
Serving
Global providers · EU contract
Tool calling
Reasoning
Structured output
Vision
Input
$0.08/ 1M
Cached input
$0.10/ 1M
Output
$0.12/ 1M

Price list, updated 12 September 2026. No minimum spend, no subscription.

About the model

Developer
Meta
Released
July 2024
Licence
Llama 3.1 Community License
Architecture
dense · 8B parameters
Input
text
Reasoning
no reasoning mode
Languages named by the developer
English, German, French, Italian, Portuguese, Hindi, Spanish, Thai

Source: the developer's model card, as of 17 September 2026.

Licence: what it means for your company

Llama 3.1 Community License. Meta's own community licence, not an open-source licence: commercial use is allowed under its terms together with Meta's Acceptable Use Policy, and downloading the weights requires accepting both. Per the card, use in languages Meta does not list is the responsibility of whoever deploys the model.

A summary of the licence, not legal advice.

Technical details

Context window
128K tokens
Max output
65,536 tokens
Throughput limit
200K tok/min
Tokenizer
Other
Streaming
yes
Open weights
meta-llama/llama-3.1-8b-instruct
Supported parameters
max_tokens
1 … 65536
temperature
0 … 2
top_p
0 … 1
frequency_penalty
-2 … 2
presence_penalty
-2 … 2
stop
array
seed
integer
Bought for you, billed by us

This model is fulfilled through global providers under our European contract: one DPA with an EU company, one invoice for every model in the catalog, no US counterparty on your paperwork. Compute may run outside the EEA — if you need strict EEA-only processing, use our own-hardware model.

Questions about this model

Where is Llama 3.1 8B Instruct computed?

Outside the EEA, through a global provider under our European contract. You get one DPA with an EU company and one invoice, but the compute itself does not run in the EEA — we state that instead of hiding it.

Does Llama 3.1 8B Instruct store prompts and completions?

No. This model runs with zero data retention: request content is not written to durable storage and nothing is used for training.

How large is the Llama 3.1 8B Instruct context window?

128K tokens, with up to 65,536 tokens of output.

What does Llama 3.1 8B Instruct cost?

$0.08 per million input tokens and $0.12 per million output tokens. No minimum spend and no subscription; the same price list covers the API and the invoice.

Is there a throughput limit on Llama 3.1 8B Instruct?

Yes, 200K tokens per minute. Higher limits are a question of volume, not of principle.

Which languages does Llama 3.1 8B Instruct support?

The developer names English, German, French, Italian, Portuguese, Hindi, Spanish, Thai. Languages not on that list may still work, but they are not claimed — test on your own material first.

Under which licence is Llama 3.1 8B Instruct released?

Llama 3.1 Community License. Meta's own community licence, not an open-source licence: commercial use is allowed under its terms together with Meta's Acceptable Use Policy, and downloading the weights requires accepting both. Per the card, use in languages Meta does not list is the responsibility of whoever deploys the model.

Can reasoning be switched off in Llama 3.1 8B Instruct?

Llama 3.1 8B Instruct has no reasoning mode — it answers directly, which keeps latency and output tokens low.

Compare with similar models

ModelDeveloperParametersContextLicenceInput / 1M
Llama 3.1 8B InstructMeta8B128KLlama 3.1 Community License$0.08
Llama 3.3 70B InstructMeta70B128KLlama 3.3 Community License$0.15
Granite 4.2 8BIBM8B128KApache-2.0$0.09
Qwen3.5 9BQwen (Alibaba)9B256KApache-2.0$0.15

30 more models behind the same key and the same invoice.

Browse the catalog