Catalog
Qwenour hardware · EEA

Qwen3.8 27B

Qwen3.8 27B, official RadixArk NVFP4 build, served on GPUs we run inside the European Economic Area. Strong at coding, tool use and agentic work. This is the model behind every measurement on the inference page: ~176 tokens/s per stream via MTP speculative decoding, first token in 0.3–0.9 s, and prefix-cache affinity routing that keeps a conversation on the machine that already holds its context — 91–97% cache hit on agentic workloads, with cached input billed at a fraction of the fresh price. Prompts and completions are not written to our logs or durable storage.

model id: qwen38

api.heabsy.com · OpenAI SDK
from openai import OpenAI
client = OpenAI(
api_key="hb-...",
base_url="https://api.heabsy.com/v1",
)
r = client.chat.completions.create(
model="qwen38",
messages=[{"role": "user", "content": "Hello"}],
)
Context
128K tokens
Serving
RTX 5090 · EEA — Italy & Norway
Tool calling
Reasoning
Structured output
Vision
Input
$0.40/ 1M
Cached input
$0.05/ 1M
Output
$3.00/ 1M

Gateway price list, synced 2026-08-29. No minimum spend, no subscription.

Computed in the EEA, never logged

This model runs on GPUs we run in the EEA (Italy and Norway). Prompts and completions are not written to our logs or durable storage, nothing is used for training, and request content is processed in the EEA. It is the model behind every measurement on the inference page — 176 tok/s per stream, sub-second first token, 91–97% prefix-cache hit on agentic workloads.

13 more models behind the same key and the same invoice.

Browse the catalog