Baseten Model serving platforms
https://www.baseten.co βKnown forProduction-grade dedicated deployments with a performance-obsessed runtime β 99.99% uptime guarantee and forward-deployed engineersPricingper-minute GPU (H100 $0.10833/min, B200 $0.16633/min) + per-token Model APIs (GPT OSS 120B $0.10/$0.50 per 1M in/out); no idle billingNotableTruss packaging; Chains claims 6x better GPU usage; B200/H100/T4 fleet; DeepSeek V4 at $1.74/$3.48 per 1M tokens; Frontier Gateway lets model creators monetize on Baseten's infra.
LLM Switchboard executes on Groq, NVIDIA, and Cloudflare Workers AI today β but those are three points in a much larger open-model inference ecosystem. This directory maps who else serves open models and how (per-token APIs, serving platforms, raw GPU clouds, gateways, edge), so you can see where a model can run, roughly what it costs, and which providers are natural next wires for the router. Every entry URL-verified Β· βwired inβ = the router executes on it today.
The field β 25 providers
| Provider | Category | Pricing | API | Known for |
|---|---|---|---|---|
Groq β wired in The LPU β custom silicon that made instant open-model inference famous | Serverless model APIs | free tier + per-token usage | OpenAI-compat | Deterministic custom-silicon speed: the provider that made near-instant open-model inference a mainstream expectation Customer case on site: 7.41x speedup with 89% cost reduction; clients include Dropbox, Vercel, Canva, Robinhood, and McLaren F1. Current homepage deliberately sβ¦ |
NVIDIA build.nvidia.com (NIM) β wired in Prototype on NVIDIA's hosted endpoints, then self-host the identical container | Serverless model APIs | free dev credits on hosted endpoints; self-host via NVIDIA AI Enterprise licensing | OpenAI-compat | Hosted-to-self-hosted portability: the API you prototype against is the same container you deploy on cloud, data center, or workstation Catalog spans Llama, Nemotron, DeepSeek, and NVIDIA-optimized community models across LLM/retrieval/vision categories; docs emphasize abstracting execution engiβ¦ |
Research-driven serverless API with one of the largest open-model catalogs | Serverless model APIs | per-token (serverless) + per-hour GPU clusters | OpenAI-compat | Catalog breadth backed by its own inference-research stack β new frontier open models at launch Site claims 31% more TPS than the next-fastest OSS engine for production coding-agent workloads; serves frontier open models at launch (e.g. MiniMax-M3 with 1M-β¦ |
Production-grade open-model inference tuned at every layer | Serverless model APIs | per-token (serverless) + per-second dedicated GPU | OpenAI-compat | Enterprise production posture β powers Notion, Cursor, Quora, and Vercel; site claims 40T+ tokens processed per day Explicitly 'OpenAI and Anthropic compatible'; DeepSeek-V4-Flash at $0.14/$0.28 per 1M in/out on the homepage; context windows up to 1,048,576 tokens; Notion latβ¦ |
One of the cheapest per-token menus in the business, from its own US data centers | Serverless model APIs | per-token + per-hour GPU | OpenAI-compat | Aggressive pricing with no long-term contracts and a zero-retention data policy DeepSeek-V4-Flash $0.09/$0.18 per 1M in/out; Nemotron-3-Ultra-550B $0.50/$2.20; on-demand DGX B300 at $4.89/instance-hr; SOC 2 and ISO 27001, own US-based data β¦ |
Wafer-scale chips serving thousands of tokens per second per user | Serverless model APIs | free tier + per-token usage | OpenAI-compat | Raw single-user speed from wafer-scale hardware β site claims up to 15x faster inference than NVIDIA GPUs Over 2,000 tokens/s stated for Llama 4 Scout; OpenAI compatibility 'with just two code changes'; catalog includes Llama 4 Scout, Gemma-4-31B, and GLM 4.7. |
RDU dataflow chips running the largest open models at hundreds of tokens per second | Serverless model APIs | per-token (cloud API); enterprise racks for on-prem | OpenAI-compat | Custom RDU silicon that hot-swaps multiple frontier-scale models on one node Site claims MiniMax M2.7 at 435 output tokens/s ('>3x faster than any other provider'), gpt-oss-120b at over 600 tokens/s, DeepSeek-V3.1 up to 200 tokens/s. |
Enterprise-grade per-token inference on Nebius's own AI cloud | Serverless model APIs | per-token (volume discounts) + dedicated endpoints | OpenAI-compat | Governed production posture: 99.9% uptime SLA on dedicated endpoints, with integrated fine-tuning and distillation-based cost/latency reduction 60+ open models; site claims the platform handles hundreds of millions of tokens per minute; fine-tuned models deploy with transparent $/token pricing. |
200+ multimodal model APIs, serverless GPUs, and agent sandboxes under one roof | Serverless model APIs | per-token + per-second serverless GPU + per-hour instances | OpenAI-compat | One-stop breadth priced 'up to 50% less' than major clouds (site claim) DeepSeek V4 Pro $1.74/1M input, Kimi K2.6 $0.95/1M input, Gemma 4 31B $0.14/1M input; dedicated endpoints via api.novita.ai base URLs. |
Aggregated global GPU network with a serverless OpenAI-compatible API on top | Serverless model APIs | per-token inference + per-hour GPU rentals | OpenAI-compat | Marketplace/aggregator angle β 'Zero Quota Limit' pay-as-you-go GPU access with no lock-in Site claims 250,000+ engineers on the platform; H100, H200, and B200 on demand; logos include Hugging Face, Midjourney, and Qwen. |
1000+ generative media models β image, video, audio, 3D β built for speed | Serverless model APIs | per-output on model APIs + hourly GPU for dedicated compute | own API | Fastest diffusion/media inference β the go-to API for FLUX and frontier video models FLUX 2 and variants, Kling 3.0, Seedance 2.0, Krea 2; partner models from OpenAI, Google, ByteDance, Alibaba; claims 99.99% uptime and zero cold starts on serveβ¦ |
Cloudflare Workers AI β wired in Open models on serverless GPUs across Cloudflare's global edge network | Edge inference | free tier (10,000 neurons/day) + usage at $0.011 per 1,000 neurons, mapped to per-token rates | OpenAI-compat | Inference at the edge with a genuinely free daily tier β no GPU to rent, requests served from Cloudflare's network GPT-OSS 20B at $0.20/$0.30 per 1M in/out tokens, Llama 3.1 70B at $0.293/$2.253, Mistral 7B at $0.11/$0.19; free 10,000 neurons/day resets at 00:00 UTC on everyβ¦ |
One API for any model β 400+ models from 70+ providers behind a single endpoint | Gateways & routers | pass-through per-token via prepaid credits that work across all models | OpenAI-compat | The de-facto meta-router: one bill, automatic provider failover, and live price/latency comparison across the whole ecosystem Site stats: 100 trillion monthly tokens processed, 10M+ users, 250,000+ integrated apps. The closest public analogue to what this router does β worth studying eβ¦ |
Dedicated inference deployments plus pre-optimized per-token Model APIs | Model serving platforms | per-minute GPU (H100 $0.10833/min, B200 $0.16633/min) + per-token Model APIs (GPT OSS 120B $0.10/$0.50 per 1M in/out); no idle billing | partial | Production-grade dedicated deployments with a performance-obsessed runtime β 99.99% uptime guarantee and forward-deployed engineers Truss packaging; Chains claims 6x better GPU usage; B200/H100/T4 fleet; DeepSeek V4 at $1.74/$3.48 per 1M tokens; Frontier Gateway lets model creators monetize β¦ |
Run thousands of community models with one line of code, pay per second | Serverless model APIs | per-second of hardware (T4 $0.000225/s, A100-80GB $0.0014/s); some models per-output/per-token; scale-to-zero | own API | Breadth of the public model library β the default place to try any new open image/video/audio model via API Cog open-source packaging; featured models include FLUX-2 Pro, Seedream 4.5, Seedance 2.0, plus lab models; its own prediction API (cog-packaged models). |
Open-source serving framework plus a managed inference cloud | Model serving platforms | BYOC/on-prem/managed options; homepage doesn't state rates | partial | Framework-first flexibility β one open-source abstraction for any model, plus a unified LLM gateway with one API for all LLMs NVIDIA H100/B200 plus AMD MI300X support; day-one pre-optimized builds of newly released open models; vLLM-backed services commonly expose OpenAI-compatible endβ¦ |
Python-native serverless GPUs with sub-second cold starts | GPU clouds | per-second GPU/CPU/memory (H100 $0.001097/s, T4 $0.000164/s) + $30/mo free credits on Starter | own API | Cold-start speed and developer experience β Python code compiles to cloud infra with instant zero-to-N autoscaling B300/B200/H100/A100/L4/T4; multi-node up to 128 B200s over 3200 Gbps InfiniBand; SOC 2 and HIPAA certified; up to $10k free compute for academic researchers. |
Per-millisecond GPU pods and serverless endpoints with sub-200ms cold starts | GPU clouds | per-second GPU (billed per-millisecond); serverless scale-to-zero with zero idle cost | partial | FlashBoot sub-200ms serverless cold starts plus a cheap Community Cloud tier of consumer GPUs 30+ GPU SKUs incl. B200 $5.89/hr, H200 $4.39/hr, H100 PCIe $1.99/hr Community, RTX 4090 $0.69/hr; 99.9% uptime SLA, SOC 2 Type II; OpenAI-compatible serving via⦠|
Energy-first AI factory: from power plant to OpenAI-compatible inference API | GPU clouds | per-hour GPU (on-demand H100 from $3.90/GPU-hr; reserved discounts) | OpenAI-compat | Builds its own power and data centers, now layered with a full managed open-model inference stack Energy-first AI cloud (stranded/renewable power); on-demand and reserved NVIDIA GPUs; pricing published per GPU-hour. |
Single-tenant NVIDIA superclusters, 1-Click Clusters, and on-demand instances | GPU clouds | per-hour GPU (per-GPU-hour on multi-GPU nodes); volume discounts on clusters | own API | Dedicated single-tenant AI supercomputers ('Your supercomputer. Your rules.') rather than shared multi-tenant serving B200 SXM6 $6.69/GPU/hr, H100 SXM $3.99/GPU/hr; 1-Click H100 clusters from $6.16/hr; SOC 2 Type II. The Inference API wind-down is a notable ecosystem exit from β¦ |
Kubernetes-native AI hyperscaler serving OpenAI, Mistral, and IBM | GPU clouds | per-hour GPU + reserved capacity (up to 60% off on-demand); spot available | partial | Hyperscale enterprise posture β the infrastructure behind OpenAI and other frontier labs, with claimed 96% cluster goodput HGX B200 8-GPU $68.80/hr, GB200 NVL72 (4 GPU) $42.00/hr; dedicated inference tier H100 at $6.16/hr. Customers: OpenAI, Mistral AI, IBM, Jane Street, Cloudflare,β¦ |
GPU marketplace where prices are set by the market, not the cloud | GPU clouds | market-set per-second GPU: on-demand, interruptible (50%+ cheaper), reserved terms up to 50% off | own API | Cheapest route to raw GPUs via marketplace competition 68+ GPU types, 40+ data centers, per-second billing, no lock-in; interruptible tier suits batch workloads; live rates published only in the console. |
The Hub as a universal front door β one API routing to many inference providers | Gateways & routers | pass-through per-token (Providers) + per-hour dedicated (Endpoints) | OpenAI-compat | The Hub network effect: the largest open-model catalog, one account, many serving backends Providers routing includes the same serverless APIs listed here; Endpoints handles autoscaling incl. scale-to-zero. |
Open models per-token inside the AWS compliance envelope | Serverless model APIs | per-token (on-demand) + provisioned throughput | partial | Where regulated enterprises already living on AWS meet open models β procurement and compliance included Model access is region-scoped; Bedrock also fronts guardrails, agents and knowledge bases. |
Microsoft's model catalog with serverless open-model endpoints | Serverless model APIs | per-token (serverless) + managed compute | partial | The Azure enterprise path to open models β same subscription, same governance as the OpenAI service Catalog spans first-party, partner and open models under one deployment surface. |