Frontier accuracy.
Small-model prices.
Train a specialized small model in under 24 hours. Pay nothing to start, pay per GPU hour in production, or run everything on your own infrastructure.
Free
Try for free
$0
Train a model and see the results on your own hardware.
- 2 full training runs
- Task description + as few as 10 examples to start
- Download the trained model and self-host it for R&D
- Community support on Slack
Production
Most popularHosted inference
~$0.04 / 1M tokens*
A dedicated GPU serving your model. You're billed per GPU hour of uptime, not per token; the per-token numbers are just what that works out to.
- Dedicated H100 endpoint running your distilled model
- Unlimited tokens, throughput is yours 24/7
- Effective cost around $0.04 / 1M tokens at sustained load
- Priority support
* Blended 3:1: three input tokens per output token.
Talk to us →Enterprise
Private deployments
Custom
Run distil labs models entirely inside your own cloud or on-premise.
- Deploy on your own infrastructure: your cloud, your VPC, or air-gapped
- Built for sensitive data: nothing leaves your environment
- Edge and on-device options down to 100M parameters
- Retraining from your own traces, on your own hardware
- SLAs, procurement and security review support
The same tokens, elsewhere
Effective price per 1M tokens at sustained load on one H100.
Blended price / 1M tokens*
Cheaper on every token. The 4B runs roughly 5× under Gemini Flash-Lite and 12× under GPT-5.4 Nano, and because you pay per GPU hour your effective rate falls further as traffic grows.
* Blended 3:1: three input tokens per output token. Effective rates at sustained load, derived from prefill and decode throughput on one H100.
1
Train free
2 runs, 10+ examples, a model in under 24 hours.
2
Deploy
Hosted endpoint, or your own hardware.
3
Improve
Production traces feed retraining automatically.
Questions
What happens at low traffic?+
A dedicated GPU makes sense above roughly 30% utilization. Below that, start on the free tier and self-host.
Which models can you distill into?+
Open-source students from 100M to ~9B parameters, including the Qwen3 family. We only use open-source teacher models, so your model is yours, with no licensing strings.
Is the free tier really free?+
Yes. Two complete training runs, downloadable weights, and local deployment for evaluation. You only pay when you go to production.
Can we buy more training runs without a production plan?+
Yes. Training credit packs are available (currently $1,000 for 10 runs).Buy credits →
Train your first model free.
10 examples, one task description, a model in under 24 hours. Pay only when you go to production.