Pricing
Inference at scale, priced for builders
From single-model experiments to production clusters. Choose the compute tier that maps to your ambition.
Sandbox
$
0
/ month
- Up to 3 active models
- 10K inference calls / month
- Community Discord access
- Public model registry
- Standard latency
Most popular
Pro
$
20
/ seat / month
- Unlimited active models
- 500K inference calls / month
- Priority support (Slack)
- Private model registry
- Custom checkpoint storage
- GPU-accelerated inference
Cluster
Custom
pricing
- Unlimited inference calls
- Dedicated GPU cluster
- 99.95% uptime SLA
- On-premise deployment
- Custom fine-tuning pipeline
- Dedicated solutions engineer