Heabsy · Modellkatalog

Jedes Modell hier: ein Schlüssel, eine Rechnung, ein europäischer Vertragspartner.

OpenAI- und Anthropic-kompatible Schnittstelle. Preise je Million Token direkt aus der Preisliste des Gateways — keine Mindestabnahme, kein Abonnement.

Eigene Hardware — Verarbeitung im EWR

Ausgeliefert auf Karten, die wir im Europäischen Wirtschaftsraum betreiben. Keine Protokollierung, Prefix-Cache, und jede Tempoangabe auf der Infrastrukturseite wurde hier gemessen. Wer ausschließlich Verarbeitung im EWR braucht, bleibt in dieser Ebene.

Weitergereicht — für Sie eingekauft, von uns abgerechnet

Erfüllt über globale Anbieter unter unserem europäischen Vertrag: Sie schließen einen Auftragsverarbeitungsvertrag mit einer EU-Gesellschaft und bekommen eine Rechnung für alle Modelle statt eines US-Vertragspartners je Anbieter. Die Berechnung kann außerhalb des EWR laufen — der Katalog sagt das, statt es zu verschweigen.

Alle Modelle, ein Schlüssel, eine Rechnung

Qwen
our hardware · EEA

Qwen3.8 27B

Official NVFP4 build on our own fleet in the EEA — the model every measurement on this page was taken on.

128K contextTool callingStructured output
Input
$0.40
Cached
$0.05
Output
$3.00
/ 1M tok
Z.ai
routed

GLM 5.3 Flash

GLM-5.3-Flash is a native multimodal model from Z.ai

1.3M contextReasoningTool callingVision
Input
$0.08
Output
$0.25
/ 1M tok
openai
routed

gpt-oss-120b

gpt-oss-120b is an open-weight, 117B-parameter Mixture-of-Experts (MoE) language model from OpenAI designed for high-reasoning, agentic, and general-purpose pro

128K contextReasoningTool callingStructured output
Input
$0.11
Output
$0.51
/ 1M tok
DeepSeek
routed

DeepSeek V4 Flash 0423

DeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with 284B total parameters and 13B activated parameters, supporting a 1M-tok

1M contextReasoningTool callingStructured output
Input
$0.25
Output
$0.50
/ 1M tok
tencent
routed

Hy3

Hy3 is a 295B-parameter Mixture-of-Experts model from Tencent (21B active, 192 experts with top-8 routing) built for reasoning, agentic workflows, and real-worl

256K contextReasoningTool callingStructured output
Input
$0.40
Output
$1.58
/ 1M tok
xiaomi
routed

MiMo-V2.5

MiMo-V2.5 is a native omnimodal model by Xiaomi

1M contextReasoningTool callingVision
Input
$0.42
Output
$0.84
/ 1M tok
Mistral
routed

Mistral Small 4

Mistral Small 4 is the next major release in the Mistral Small family, unifying the capabilities of several flagship Mistral models into a single system

256K contextReasoningTool callingVision
Input
$0.45
Output
$1.80
/ 1M tok
minimax
routed

MiniMax M3

MiniMax-M3 is a multimodal foundation model from MiniMax

1M contextReasoningTool callingVision
Input
$0.90
Output
$3.60
/ 1M tok
DeepSeek
routed

DeepSeek V4 Pro 0423

DeepSeek V4 Pro is a large-scale Mixture-of-Experts model from DeepSeek with 1.6T total parameters and 49B activated parameters, supporting a 1M-token context w

1M contextReasoningTool callingStructured output
Input
$1.83
Output
$3.66
/ 1M tok
Qwen
routed

Qwen3 Coder Plus

Qwen3 Coder Plus is Alibaba's proprietary version of the Open Source Qwen3 Coder 480B A35B

1M contextTool callingStructured output
Input
$1.95
Output
$9.75
/ 1M tok
Moonshot AI
routed

Kimi K2.7 Code

MoonshotAI: Kimi K2.7 Code is a coding-focused model in Moonshot AI's Kimi K2 family, built to complete end-to-end programming tasks reliably over long contexts

256K contextReasoningTool callingVision
Input
$1.98
Output
$10.20
/ 1M tok
Z.ai
routed

GLM 5.2

GLM 5.2 is a large-scale reasoning model from Z.ai

1M contextReasoningTool callingStructured output
Input
$3.57
Output
$11.22
/ 1M tok
Z.ai
routed

GLM 5.3

GLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks

1.3M contextReasoningTool callingStructured output
Input
$4.20
Output
$13.20
/ 1M tok
Moonshot AI
routed

Kimi K3

Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI

1M contextReasoningTool callingVision
Input
$9.00
Output
$45.00
/ 1M tok

Preise aus der Preisliste des Gateways, Stand 2026-08-29. Die Rechnung folgt dem Gateway, und diese Seite folgt der Rechnung — nie umgekehrt.api.heabsy.com/openrouter/v1/models

Fehlt ein Modell, das Sie brauchen?

Der weitergereichte Katalog wächst auf Anfrage, und offene Modelle lassen sich auf unsere eigene Hardware im EWR holen, sobald die Last eine eigene Flotte rechtfertigt.

Sprechen Sie uns an

Häufige Fragen

Welches Modell soll ich nehmen?+

Für Dokumentenarbeit auf Deutsch entscheidet das der Pilot an Ihren eigenen Unterlagen, nicht die Rangliste: bei deutschen Verträgen und Handbüchern unterscheiden sich Modelle deutlich stärker, als die Benchmarks vermuten lassen. Wo ausschließlich Verarbeitung im EWR gefordert ist, bleibt nur das Modell auf eigener Hardware.

Was ist der Unterschied zwischen eigener Hardware und weitergereicht?+

Beim Modell auf eigener Hardware betreiben wir die Karten selbst im Europäischen Wirtschaftsraum: Anfragen werden nicht protokolliert, nichts wird für Training verwendet, die Verarbeitung bleibt im EWR. Weitergereichte Modelle kaufen wir bei globalen Anbietern unter unserem europäischen Vertrag ein — Sie haben einen Auftragsverarbeitungsvertrag und eine Rechnung, aber die Berechnung kann außerhalb des EWR laufen. Der Katalog sagt bei jedem Modell, was zutrifft.

Ändern sich die Preise?+

Sie folgen der Preisliste des Gateways und werden von dort übernommen, nicht von Hand gepflegt. Der Stand steht unter dem Katalog. Abgerechnet wird immer das, was das Gateway berechnet — diese Seite folgt der Rechnung und nicht umgekehrt.

Gibt es eine Mindestabnahme?+

Nein, weder Mindestumsatz noch Abonnement. Bezahlt wird nach Token, getrennt nach frischer Eingabe, Eingabe aus dem Cache und Ausgabe. Bei agentischer Arbeit ist der Cache-Anteil der größte Posten, und er kostet einen Bruchteil.

Fehlt ein Modell, das wir brauchen?+

Der weitergereichte Katalog wächst auf Anfrage. Offene Modelle lassen sich zusätzlich auf unsere eigene Hardware holen, sobald die Last eine eigene Flotte rechtfertigt — das ist eine Frage des Volumens, nicht des Prinzips.

Zuletzt aktualisiert: