Qwen3.8 27B
Qwen3.8 27B, official RadixArk NVFP4 build, served on GPUs we run inside the European Economic Area. Strong at coding, tool use and agentic work. This is the model behind every measurement on the inference page: ~176 tokens/s per stream via MTP speculative decoding, first token in 0.3–0.9 s, and prefix-cache affinity routing that keeps a conversation on the machine that already holds its context — 91–97% cache hit on agentic workloads, with cached input billed at a fraction of the fresh price. Prompts and completions are not written to our logs or durable storage.
model id: qwen38
from openai import OpenAIclient = OpenAI(api_key="hb-...",base_url="https://api.heabsy.com/v1",)r = client.chat.completions.create(model="qwen38",messages=[{"role": "user", "content": "Hello"}],)
Gateway price list, synced 2026-08-29. No minimum spend, no subscription.
This model runs on GPUs we run in the EEA (Italy and Norway). Prompts and completions are not written to our logs or durable storage, nothing is used for training, and request content is processed in the EEA. It is the model behind every measurement on the inference page — 176 tok/s per stream, sub-second first token, 91–97% prefix-cache hit on agentic workloads.
13 more models behind the same key and the same invoice.
Browse the catalog