Page 1 of 2
DEDICATED CAPACITY IN EUROPE,
AT SERVERLESS PRICES
ONE MINUTE · FIVE QUESTIONS
Tell us about your workload. We tune the engine to how you actually use it: cache, context and concurrency. For qualified production volumes, dedicated GPU capacity is reserved for your environment, at the same per-token price as serverless.
Pricing and performance
Gemma 4 31B
:
$0.12/M input · $0.38/M output
Qwen 3.8-27B :
$0.12/M input · $0.38/M output · $0.04/M cached
Whisper Large v3 :
$0.00048/min
Kokoro :
$0.62/M characters
BGE-M3 :
$0.01/M input tokens
First token in as little as ~50 ms in Europe · Infrastructure designed for 99.9999% availability.
Up to 10 billion tokens available for production workload testing. No credit card required.
Founder access: 20% off
public pricing for 36 months · Available until September 30.
How it works
1. Tell us about your setup by answering the five questions below.
2. Try it on your real traffic using your free allowance.
3. Go live with your settings: we size the engine to your cache, context and concurrency. For production volumes, a written proposal sets out the GPU capacity reserved for your environment.
Work email
*
Estimated monthly volume
*
A
Under 1 billion tokens
B
1 to 10 billion
C
10 to 50 billion
D
Over 50 billion
E
Not sure yet
Primary use case
*
A
Document extraction
B
Conversational or agents
C
Search or RAG
D
Inference built into your product
E
Other
Model or service you have in mind
*
A
Gemma 4 31B
B
Qwen 3.8 27B
C
Whisper Large v3
D
Kokoro
E
BGE-M3
F
Other
Your main constraint right now
EU hosting · No retention · OpenAI-compatible API
RECEIVE MY FOUNDER ACCESS