SOVINFRAPrivate modeless inference · EU-hosted.
Sovereign AI stackSEPTEMBER 2026

Built for high-volume workloads

One OpenAI-compatible endpoint. Let the router pick, or call a model by name. Five operated models, one usage counter, one bill.

QWEN 3.8-27B FP8 · H200Code, reasoning, vision and tool calling
GEMMA 4 31B BF16 · H200Text, vision and extraction
WHISPER LARGE V3FR / EN audio coverage
BGE-M3Embeddings and semantic search
KOKOROFR / EN speech synthesis

Status

First token from ~50 ms in Europe · Redundant hosting designed for 99.9999% availability

Service status →

Arena

Compare AI. Measured live.

Launch Arena

Built for self-hosted agents

One URL. Five models. One meter.

The simplicity of a unified model router, with every model served on European GPUs operated by SovInfra.
01

Migrate in one line

Point any OpenAI-compatible client to SovAPI. Hermes Agent and OpenClaw work today without a custom integration.

02

Five operated models

Qwen 3.8-27B, Gemma 4 31B, Whisper, BGE-M3 and Kokoro are tested, monitored and served by SovInfra. More arrive in September.

03

Public latency proof

Measure real request latency in the Arena without creating an account.

04

European by design

No intermediary, no data retention and no training on customer data. Built for repetitive, context-heavy agent loops.

Automatic routing

Let SovAPI choose
"model": "sovapi"
Automatic routing across Gemma 4 31B and Qwen 3.8-27B. Whisper, BGE-M3 and Kokoro are also available through SovAPI.

Direct model call

Gemma 4 31B
"model": "gemma-4-31b"
Text, vision and structured extraction.

Direct model call

Qwen 3.8-27B
"model": "qwen3.8-27b"
Code, reasoning, vision and tool calling.

Benchmarks

Results published by the model developers, according to their evaluation protocols.

Qwen researchGoogle GemmaAnthropic publicationsDeepSeek publications

Run the Arena

Simple usage pricing

Pricing.

Billed by measured usage, with no commitment. Start with 1B free tokens.

Pay-as-you-go

$0.12/ M tokens in
Your first billion tokens are free.No credit card required.
Gemma 4 31B BF16 · H200
$0.12 in · $0.38 out
Qwen 3.8-27B FP8 · H200
$0.12 in · $0.38 out · $0.04 cached
Whisper Large v3
Transcription, 99+ languages$0.00048 / audio minute
Kokoro
Speech synthesis FR/EN$0.62 / M charactersOther languages on request
BGE-M3
Embeddings$0.01 / M input tokens
Usage
One usage counter · One invoice
Support
Email response < 15 min · 08:00–20:00 CET
Data retention
None

One API. Let the router pick, or call a model by name.

Get your free API key

EU hosting · GDPR native · No data training · OpenAI-compatible API

Built for European workloads

Security. Compliance.

01

EU-hosted

All inference runs on European GPUs. No data leaves the EU.

02

GDPR native

No training on user data. End-to-end encryption. DPA available.

03

NemoClaw protected

Action sandboxing, strict access policy, prompt injection protection. Powered by NVIDIA.

04

Auditable

Our infrastructure stack is open for inspection on Codeberg. Full transparency on security and routing.

1 billion free tokens

No credit card required

Create API key →