Frontier accuracy.
Small-model prices.

Train a specialized small model in under 24 hours. Pay nothing to start, pay per GPU hour in production, or run everything on your own infrastructure.

Free

Try for free

$0

Train a model and see the results on your own hardware.

  • 2 full training runs
  • Task description + as few as 10 examples to start
  • Download the trained model and self-host it for R&D
  • Community support on Slack
Start training →

Production

Most popular

Hosted inference

~$0.04 / 1M tokens*

A dedicated GPU serving your model. You're billed per GPU hour of uptime, not per token; the per-token numbers are just what that works out to.

  • Dedicated H100 endpoint running your distilled model
  • Unlimited tokens, throughput is yours 24/7
  • Effective cost around $0.04 / 1M tokens at sustained load
  • Priority support

* Blended 3:1: three input tokens per output token.

Talk to us →

Enterprise

Private deployments

Custom

Run distil labs models entirely inside your own cloud or on-premise.

  • Deploy on your own infrastructure: your cloud, your VPC, or air-gapped
  • Built for sensitive data: nothing leaves your environment
  • Edge and on-device options down to 100M parameters
  • Retraining from your own traces, on your own hardware
  • SLAs, procurement and security review support
Contact sales →

The same tokens, elsewhere

Effective price per 1M tokens at sustained load on one H100.

Blended price / 1M tokens*

Your distilled model, Qwen3 4B$0.04
Your distilled model, Qwen3.5 9B$0.08
Gemini 2.5 Flash-Lite$0.18
GPT-5.4 Nano$0.46
Claude Haiku 4.5$2.00

Cheaper on every token. The 4B runs roughly 5× under Gemini Flash-Lite and 12× under GPT-5.4 Nano, and because you pay per GPU hour your effective rate falls further as traffic grows.

* Blended 3:1: three input tokens per output token. Effective rates at sustained load, derived from prefill and decode throughput on one H100.

1

Train free

2 runs, 10+ examples, a model in under 24 hours.

2

Deploy

Hosted endpoint, or your own hardware.

3

Improve

Production traces feed retraining automatically.

Questions

What happens at low traffic?+

A dedicated GPU makes sense above roughly 30% utilization. Below that, start on the free tier and self-host.

Which models can you distill into?+

Open-source students from 100M to ~9B parameters, including the Qwen3 family. We only use open-source teacher models, so your model is yours, with no licensing strings.

Is the free tier really free?+

Yes. Two complete training runs, downloadable weights, and local deployment for evaluation. You only pay when you go to production.

Can we buy more training runs without a production plan?+

Yes. Training credit packs are available (currently $1,000 for 10 runs).Buy credits →

Train your first model free.

10 examples, one task description, a model in under 24 hours. Pay only when you go to production.