What will this actually cost β and what's the cheapest honest way to run it? Describe a workload and this computes real monthly spend across every priced model in the catalog, shows how much of it the documented free tiers absorb, and reports the true saving between the cheapest model that clears a quality bar and a frontier-grade one. Prices are the catalog's own $/M-token figures (28 chat models). No markup, no vendor maths β the numbers move when you move the sliders.
1 Β· Describe the workload
Quality floor 60 β the cheapest model must still score at least this on the blended capability index.
2 Β· The verdict
3 Β· Free-tier coverage β how much of this workload the documented free allowances absorb
4 Β· Every model, priced for your workload
| Model | Quality | $/mo | $/1k requests | vs cheapest |
|---|
8 models aren't ranked above β they have no per-token list price in the catalog: GPT-OSS 20B, Compound, Compound Mini, NVIDIA Nemotron 3 Ultra 550B-A55B, NVIDIA Nemotron 3 Super 120B-A12B, NVIDIA Nemotron 3 Nano 30B-A3B, NVIDIA Nemotron 3 Nano Omni 30B-A3B (Reasoning), gpt-oss-20b. Most are credit-based (NVIDIA NIM free credits / self-host) or usage-billed, so quoting them at "$0/mo" would flatter the numbers. Their real cost is free-tier allowance then credits β priced differently, not free at scale.
Honesty notes: costs are list $/M-token prices from the catalog Γ your token volumes β real usage varies with caching, retries and system prompts. Only models with a genuine per-token price are ranked (28 of 36 chat models). Free-tier coverage uses the documented daily allowances (Groq per-model caps summed, Cloudflare neuron-derived); one-time credits are excluded. Quality is the blended capability index used by the router, not a benchmark score.