Nemotron 3 Super 120B A12B
Nemotron 3 Super 120B A12B is NVIDIA's larger hybrid model (NVIDIA Nemotron Open Model License, March 2026): 120B parameters with 12B active, mixing state-space, mixture-of-experts and attention layers. It is aimed at agentic workflows, long-context reasoning, tool use, retrieval-augmented generation and high-volume jobs such as ticket automation; the card reports 91.75 on the RULER long-context test at one million tokens. Supported languages include English, German, French, Italian, Spanish, Japanese and Chinese. It is fulfilled through global providers under our European contract, and requests go only to backends that do not retain prompts or completions.
model id: nemotron-3-super-120b-a12b
from openai import OpenAI client = OpenAI( api_key="hb-...", base_url="https://api.heabsy.com/v1",) r = client.chat.completions.create( model="nemotron-3-super-120b-a12b", messages=[{"role": "user", "content": "Hello"}],)print(r.choices[0].message.content)
curl https://api.heabsy.com/v1/chat/completions \ -H "Authorization: Bearer $HEABSY_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "nemotron-3-super-120b-a12b", "messages": [{"role": "user", "content": "Hello"}] }'
import OpenAI from "openai"; const client = new OpenAI({ apiKey: process.env.HEABSY_API_KEY, baseURL: "https://api.heabsy.com/v1",}); const r = await client.chat.completions.create({ model: "nemotron-3-super-120b-a12b", messages: [{ role: "user", content: "Hello" }],});console.log(r.choices[0].message.content);
Price list, updated 12 September 2026. No minimum spend, no subscription.
About the model
- Developer
- NVIDIA
- Released
- March 2026
- Licence
- NVIDIA Nemotron Open Model License
- Architecture
- mixture of experts · 120B parameters, 12B active per token
- Input
- text
- Reasoning
- switchable per request
- Languages named by the developer
- English, French, German, Italian, Japanese, Spanish, Chinese
Source: the developer's model card, as of 17 September 2026.
Licence: what it means for your company
NVIDIA Nemotron Open Model License. NVIDIA's own open model licence rather than a standard open-source licence. NVIDIA marks the model as ready for commercial use; the terms for redistributing weights or derived models are in the licence itself.
A summary of the licence, not legal advice.
Benchmarks reported by the developer
- RULER (1M)
- 91.75
Figures from the developer's model card, not our measurement. Source: huggingface.co/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16
Technical details
- Context window
- 256K tokens
- Max output
- 65,536 tokens
- Throughput limit
- 200K tok/min
- Tokenizer
- Other
- Streaming
- yes
- Open weights
- nvidia/nemotron-3-super-120b-a12b
- max_tokens
- 1 … 65536
- temperature
- 0 … 2
- top_p
- 0 … 1
- frequency_penalty
- -2 … 2
- presence_penalty
- -2 … 2
- stop
- array
- seed
- integer
This model is fulfilled through global providers under our European contract: one DPA with an EU company, one invoice for every model in the catalog, no US counterparty on your paperwork. Compute may run outside the EEA — if you need strict EEA-only processing, use our own-hardware model.
Questions about this model
›Where is Nemotron 3 Super 120B A12B computed?
Outside the EEA, through a global provider under our European contract. You get one DPA with an EU company and one invoice, but the compute itself does not run in the EEA — we state that instead of hiding it.
›Does Nemotron 3 Super 120B A12B store prompts and completions?
No. This model runs with zero data retention: request content is not written to durable storage and nothing is used for training.
›How large is the Nemotron 3 Super 120B A12B context window?
256K tokens, with up to 65,536 tokens of output.
›What does Nemotron 3 Super 120B A12B cost?
$0.13 per million input tokens and $0.62 per million output tokens. No minimum spend and no subscription; the same price list covers the API and the invoice.
›Is there a throughput limit on Nemotron 3 Super 120B A12B?
Yes, 200K tokens per minute. Higher limits are a question of volume, not of principle.
›Which languages does Nemotron 3 Super 120B A12B support?
The developer names English, French, German, Italian, Japanese, Spanish, Chinese. Languages not on that list may still work, but they are not claimed — test on your own material first.
›Under which licence is Nemotron 3 Super 120B A12B released?
NVIDIA Nemotron Open Model License. NVIDIA's own open model licence rather than a standard open-source licence. NVIDIA marks the model as ready for commercial use; the terms for redistributing weights or derived models are in the licence itself.
›Can reasoning be switched off in Nemotron 3 Super 120B A12B?
Yes. Per the developer, thinking can be switched on or off for each request.
Compare with similar models
| Model | Developer | Parameters | Context | Licence | Input / 1M |
|---|---|---|---|---|---|
| Nemotron 3 Super 120B A12B | NVIDIA | 120B / 12B | 256K | NVIDIA Nemotron Open Model License | $0.13 |
| Nemotron 3.5 Lightning | NVIDIA | 30B / 3B | 256K | OpenMDW-1.1 | $0.12 |
| Laguna S 2.1 | Poolside | 118B / 8B | 1M | OpenMDW-1.1 | $0.09 |
| gpt-oss-120b | OpenAI | 117B / 5.1B | 128K | Apache-2.0 | $0.11 |
30 more models behind the same key and the same invoice.
Browse the catalog