Built for high-volume workloads
One OpenAI-compatible endpoint. Let the router pick, or call a model by name. Five operated models, one usage counter, one bill.
Status
First token from ~50 ms in Europe · Redundant hosting designed for 99.9999% availability
Arena
Compare AI. Measured live.
Built for self-hosted agents
One URL. Five models. One meter.
The simplicity of a unified model router, with every model served on European GPUs operated by SovInfra.Migrate in one line
Point any OpenAI-compatible client to SovAPI. Hermes Agent and OpenClaw work today without a custom integration.
Five operated models
Qwen 3.8-27B, Gemma 4 31B, Whisper, BGE-M3 and Kokoro are tested, monitored and served by SovInfra. More arrive in September.
Public latency proof
Measure real request latency in the Arena without creating an account.
European by design
No intermediary, no data retention and no training on customer data. Built for repetitive, context-heavy agent loops.
Automatic routing
Let SovAPI choose"model": "sovapi"Automatic routing across Gemma 4 31B and Qwen 3.8-27B. Whisper, BGE-M3 and Kokoro are also available through SovAPI.Direct model call
Gemma 4 31B"model": "gemma-4-31b"Text, vision and structured extraction.Direct model call
Qwen 3.8-27B"model": "qwen3.8-27b"Code, reasoning, vision and tool calling.Benchmarks
Results published by the model developers, according to their evaluation protocols.
Qwen researchGoogle GemmaAnthropic publicationsDeepSeek publications
Run the ArenaSimple usage pricing
Pricing.
Billed by measured usage, with no commitment. Start with 1B free tokens.Pay-as-you-go
$0.12/ M tokens in- Gemma 4 31B BF16 · H200
- $0.12 in · $0.38 out
- Qwen 3.8-27B FP8 · H200
- $0.12 in · $0.38 out · $0.04 cached
- Whisper Large v3
- Transcription, 99+ languages$0.00048 / audio minute
- Kokoro
- Speech synthesis FR/EN$0.62 / M charactersOther languages on request
- BGE-M3
- Embeddings$0.01 / M input tokens
- Usage
- One usage counter · One invoice
- Support
- Email response < 15 min · 08:00–20:00 CET
- Data retention
- None
One API. Let the router pick, or call a model by name.
Get your free API keyEU hosting · GDPR native · No data training · OpenAI-compatible API
Built for European workloads
Security. Compliance.
EU-hosted
All inference runs on European GPUs. No data leaves the EU.
GDPR native
No training on user data. End-to-end encryption. DPA available.
NemoClaw protected
Action sandboxing, strict access policy, prompt injection protection. Powered by NVIDIA.
Auditable
Our infrastructure stack is open for inspection on Codeberg. Full transparency on security and routing.
1 billion free tokens
No credit card required