Heabsy AI Platform · Catalog

Every model here is one key, one bill, one European counterparty.

OpenAI- and Anthropic-compatible API. Per-million pricing straight from the gateway's price list — no minimum spend, no subscription.

Our hardware — computed in the EEA

Served on GPUs we operate in Italy and Norway. Zero data retention, prefix caching, every speed figure on the inference page measured here. Strict EEA-only processing lives in this tier.

Routed — bought for you, billed by us

Fulfilled through global providers under our European contract: you sign one DPA with an EU company and get one invoice for every model, instead of a US counterparty per vendor. Compute may run outside the EEA — the catalog says so instead of hiding it.

Qwen
our hardware · EEA

Qwen3.8 27B

Official NVFP4 build on our RTX 5090 fleet in the EEA — the model every measurement on this page was taken on.

128K contextTool callingStructured output
Input
$0.40
Cached
$0.05
Output
$3.00
/ 1M tok
Z.ai
routed

GLM 5.3 Flash

GLM-5.3-Flash is a native multimodal model from Z.ai

1.3M contextReasoningTool callingVision
Input
$0.08
Output
$0.25
/ 1M tok
openai
routed

gpt-oss-120b

gpt-oss-120b is an open-weight, 117B-parameter Mixture-of-Experts (MoE) language model from OpenAI designed for high-reasoning, agentic, and general-purpose pro

128K contextReasoningTool callingStructured output
Input
$0.11
Output
$0.51
/ 1M tok
DeepSeek
routed

DeepSeek V4 Flash 0423

DeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with 284B total parameters and 13B activated parameters, supporting a 1M-tok

1M contextReasoningTool callingStructured output
Input
$0.25
Output
$0.50
/ 1M tok
tencent
routed

Hy3

Hy3 is a 295B-parameter Mixture-of-Experts model from Tencent (21B active, 192 experts with top-8 routing) built for reasoning, agentic workflows, and real-worl

256K contextReasoningTool callingStructured output
Input
$0.40
Output
$1.58
/ 1M tok
xiaomi
routed

MiMo-V2.5

MiMo-V2.5 is a native omnimodal model by Xiaomi

1M contextReasoningTool callingVision
Input
$0.42
Output
$0.84
/ 1M tok
Mistral
routed

Mistral Small 4

Mistral Small 4 is the next major release in the Mistral Small family, unifying the capabilities of several flagship Mistral models into a single system

256K contextReasoningTool callingVision
Input
$0.45
Output
$1.80
/ 1M tok
minimax
routed

MiniMax M3

MiniMax-M3 is a multimodal foundation model from MiniMax

1M contextReasoningTool callingVision
Input
$0.90
Output
$3.60
/ 1M tok
DeepSeek
routed

DeepSeek V4 Pro 0423

DeepSeek V4 Pro is a large-scale Mixture-of-Experts model from DeepSeek with 1.6T total parameters and 49B activated parameters, supporting a 1M-token context w

1M contextReasoningTool callingStructured output
Input
$1.83
Output
$3.66
/ 1M tok
Qwen
routed

Qwen3 Coder Plus

Qwen3 Coder Plus is Alibaba's proprietary version of the Open Source Qwen3 Coder 480B A35B

1M contextTool callingStructured output
Input
$1.95
Output
$9.75
/ 1M tok
Moonshot AI
routed

Kimi K2.7 Code

MoonshotAI: Kimi K2.7 Code is a coding-focused model in Moonshot AI's Kimi K2 family, built to complete end-to-end programming tasks reliably over long contexts

256K contextReasoningTool callingVision
Input
$1.98
Output
$10.20
/ 1M tok
Z.ai
routed

GLM 5.2

GLM 5.2 is a large-scale reasoning model from Z.ai

1M contextReasoningTool callingStructured output
Input
$3.57
Output
$11.22
/ 1M tok
Z.ai
routed

GLM 5.3

GLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks

1.3M contextReasoningTool callingStructured output
Input
$4.20
Output
$13.20
/ 1M tok
Moonshot AI
routed

Kimi K3

Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI

1M contextReasoningTool callingVision
Input
$9.00
Output
$45.00
/ 1M tok

Prices from the gateway price list, synced 2026-08-29. The bill always follows the gateway, and this page follows the gateway — never the other way round. Machine-readable: api.heabsy.com/openrouter/v1/models (own hardware).

Missing a model you need?

The routed catalog grows on request, and open models can be brought onto our own EEA hardware when the workload justifies a dedicated fleet.

Talk to us