Inference providers

Who serves open models and how β€” serverless APIs, serving platforms, GPU clouds, gateways, edge. Seeded by Baseten.

Baseten Model serving platforms
https://www.baseten.co β†—
you asked about this

Deploys custom, open-source, and fine-tuned models on inference-optimized infrastructure, packaged via its open-source Truss framework and accelerated by a performance runtime (custom kernels, speculative decoding, Chains for compound AI). Two modes: Dedicated Deployments billed per-minute of GPU time, and per-token Model APIs (DeepSeek V4, Kimi K2.6, GLM 5.1, GPT OSS 120B). Runs managed, in your VPC, or hybrid.

Known forProduction-grade dedicated deployments with a performance-obsessed runtime β€” 99.99% uptime guarantee and forward-deployed engineersPricingper-minute GPU (H100 $0.10833/min, B200 $0.16633/min) + per-token Model APIs (GPT OSS 120B $0.10/$0.50 per 1M in/out); no idle billingNotableTruss packaging; Chains claims 6x better GPU usage; B200/H100/T4 fleet; DeepSeek V4 at $1.74/$3.48 per 1M tokens; Frontier Gateway lets model creators monetize on Baseten's infra.
LLM Switchboard executes on Groq, NVIDIA, and Cloudflare Workers AI today β€” but those are three points in a much larger open-model inference ecosystem. This directory maps who else serves open models and how (per-token APIs, serving platforms, raw GPU clouds, gateways, edge), so you can see where a model can run, roughly what it costs, and which providers are natural next wires for the router. Every entry URL-verified Β· β€œwired in” = the router executes on it today.
The field β€” 25 providers
25 / 25
ProviderCategoryPricingAPIKnown for
Groq β†— wired in
The LPU β€” custom silicon that made instant open-model inference famous
Serverless model APIsfree tier + per-token usageOpenAI-compatDeterministic custom-silicon speed: the provider that made near-instant open-model inference a mainstream expectation
Customer case on site: 7.41x speedup with 89% cost reduction; clients include Dropbox, Vercel, Canva, Robinhood, and McLaren F1. Current homepage deliberately s…
Prototype on NVIDIA's hosted endpoints, then self-host the identical container
Serverless model APIsfree dev credits on hosted endpoints; self-host via NVIDIA AI Enterprise licensingOpenAI-compatHosted-to-self-hosted portability: the API you prototype against is the same container you deploy on cloud, data center, or workstation
Catalog spans Llama, Nemotron, DeepSeek, and NVIDIA-optimized community models across LLM/retrieval/vision categories; docs emphasize abstracting execution engi…
Research-driven serverless API with one of the largest open-model catalogs
Serverless model APIsper-token (serverless) + per-hour GPU clustersOpenAI-compatCatalog breadth backed by its own inference-research stack β€” new frontier open models at launch
Site claims 31% more TPS than the next-fastest OSS engine for production coding-agent workloads; serves frontier open models at launch (e.g. MiniMax-M3 with 1M-…
Production-grade open-model inference tuned at every layer
Serverless model APIsper-token (serverless) + per-second dedicated GPUOpenAI-compatEnterprise production posture β€” powers Notion, Cursor, Quora, and Vercel; site claims 40T+ tokens processed per day
Explicitly 'OpenAI and Anthropic compatible'; DeepSeek-V4-Flash at $0.14/$0.28 per 1M in/out on the homepage; context windows up to 1,048,576 tokens; Notion lat…
One of the cheapest per-token menus in the business, from its own US data centers
Serverless model APIsper-token + per-hour GPUOpenAI-compatAggressive pricing with no long-term contracts and a zero-retention data policy
DeepSeek-V4-Flash $0.09/$0.18 per 1M in/out; Nemotron-3-Ultra-550B $0.50/$2.20; on-demand DGX B300 at $4.89/instance-hr; SOC 2 and ISO 27001, own US-based data …
Wafer-scale chips serving thousands of tokens per second per user
Serverless model APIsfree tier + per-token usageOpenAI-compatRaw single-user speed from wafer-scale hardware β€” site claims up to 15x faster inference than NVIDIA GPUs
Over 2,000 tokens/s stated for Llama 4 Scout; OpenAI compatibility 'with just two code changes'; catalog includes Llama 4 Scout, Gemma-4-31B, and GLM 4.7.
RDU dataflow chips running the largest open models at hundreds of tokens per second
Serverless model APIsper-token (cloud API); enterprise racks for on-premOpenAI-compatCustom RDU silicon that hot-swaps multiple frontier-scale models on one node
Site claims MiniMax M2.7 at 435 output tokens/s ('>3x faster than any other provider'), gpt-oss-120b at over 600 tokens/s, DeepSeek-V3.1 up to 200 tokens/s.
Enterprise-grade per-token inference on Nebius's own AI cloud
Serverless model APIsper-token (volume discounts) + dedicated endpointsOpenAI-compatGoverned production posture: 99.9% uptime SLA on dedicated endpoints, with integrated fine-tuning and distillation-based cost/latency reduction
60+ open models; site claims the platform handles hundreds of millions of tokens per minute; fine-tuned models deploy with transparent $/token pricing.
200+ multimodal model APIs, serverless GPUs, and agent sandboxes under one roof
Serverless model APIsper-token + per-second serverless GPU + per-hour instancesOpenAI-compatOne-stop breadth priced 'up to 50% less' than major clouds (site claim)
DeepSeek V4 Pro $1.74/1M input, Kimi K2.6 $0.95/1M input, Gemma 4 31B $0.14/1M input; dedicated endpoints via api.novita.ai base URLs.
Aggregated global GPU network with a serverless OpenAI-compatible API on top
Serverless model APIsper-token inference + per-hour GPU rentalsOpenAI-compatMarketplace/aggregator angle β€” 'Zero Quota Limit' pay-as-you-go GPU access with no lock-in
Site claims 250,000+ engineers on the platform; H100, H200, and B200 on demand; logos include Hugging Face, Midjourney, and Qwen.
1000+ generative media models β€” image, video, audio, 3D β€” built for speed
Serverless model APIsper-output on model APIs + hourly GPU for dedicated computeown APIFastest diffusion/media inference β€” the go-to API for FLUX and frontier video models
FLUX 2 and variants, Kling 3.0, Seedance 2.0, Krea 2; partner models from OpenAI, Google, ByteDance, Alibaba; claims 99.99% uptime and zero cold starts on serve…
Open models on serverless GPUs across Cloudflare's global edge network
Edge inferencefree tier (10,000 neurons/day) + usage at $0.011 per 1,000 neurons, mapped to per-token ratesOpenAI-compatInference at the edge with a genuinely free daily tier β€” no GPU to rent, requests served from Cloudflare's network
GPT-OSS 20B at $0.20/$0.30 per 1M in/out tokens, Llama 3.1 70B at $0.293/$2.253, Mistral 7B at $0.11/$0.19; free 10,000 neurons/day resets at 00:00 UTC on every…
One API for any model β€” 400+ models from 70+ providers behind a single endpoint
Gateways & routerspass-through per-token via prepaid credits that work across all modelsOpenAI-compatThe de-facto meta-router: one bill, automatic provider failover, and live price/latency comparison across the whole ecosystem
Site stats: 100 trillion monthly tokens processed, 10M+ users, 250,000+ integrated apps. The closest public analogue to what this router does β€” worth studying e…
Dedicated inference deployments plus pre-optimized per-token Model APIs
Model serving platformsper-minute GPU (H100 $0.10833/min, B200 $0.16633/min) + per-token Model APIs (GPT OSS 120B $0.10/$0.50 per 1M in/out); no idle billingpartialProduction-grade dedicated deployments with a performance-obsessed runtime β€” 99.99% uptime guarantee and forward-deployed engineers
Truss packaging; Chains claims 6x better GPU usage; B200/H100/T4 fleet; DeepSeek V4 at $1.74/$3.48 per 1M tokens; Frontier Gateway lets model creators monetize …
Run thousands of community models with one line of code, pay per second
Serverless model APIsper-second of hardware (T4 $0.000225/s, A100-80GB $0.0014/s); some models per-output/per-token; scale-to-zeroown APIBreadth of the public model library β€” the default place to try any new open image/video/audio model via API
Cog open-source packaging; featured models include FLUX-2 Pro, Seedream 4.5, Seedance 2.0, plus lab models; its own prediction API (cog-packaged models).
Open-source serving framework plus a managed inference cloud
Model serving platformsBYOC/on-prem/managed options; homepage doesn't state ratespartialFramework-first flexibility β€” one open-source abstraction for any model, plus a unified LLM gateway with one API for all LLMs
NVIDIA H100/B200 plus AMD MI300X support; day-one pre-optimized builds of newly released open models; vLLM-backed services commonly expose OpenAI-compatible end…
Python-native serverless GPUs with sub-second cold starts
GPU cloudsper-second GPU/CPU/memory (H100 $0.001097/s, T4 $0.000164/s) + $30/mo free credits on Starterown APICold-start speed and developer experience β€” Python code compiles to cloud infra with instant zero-to-N autoscaling
B300/B200/H100/A100/L4/T4; multi-node up to 128 B200s over 3200 Gbps InfiniBand; SOC 2 and HIPAA certified; up to $10k free compute for academic researchers.
Per-millisecond GPU pods and serverless endpoints with sub-200ms cold starts
GPU cloudsper-second GPU (billed per-millisecond); serverless scale-to-zero with zero idle costpartialFlashBoot sub-200ms serverless cold starts plus a cheap Community Cloud tier of consumer GPUs
30+ GPU SKUs incl. B200 $5.89/hr, H200 $4.39/hr, H100 PCIe $1.99/hr Community, RTX 4090 $0.69/hr; 99.9% uptime SLA, SOC 2 Type II; OpenAI-compatible serving via…
Energy-first AI factory: from power plant to OpenAI-compatible inference API
GPU cloudsper-hour GPU (on-demand H100 from $3.90/GPU-hr; reserved discounts)OpenAI-compatBuilds its own power and data centers, now layered with a full managed open-model inference stack
Energy-first AI cloud (stranded/renewable power); on-demand and reserved NVIDIA GPUs; pricing published per GPU-hour.
Single-tenant NVIDIA superclusters, 1-Click Clusters, and on-demand instances
GPU cloudsper-hour GPU (per-GPU-hour on multi-GPU nodes); volume discounts on clustersown APIDedicated single-tenant AI supercomputers ('Your supercomputer. Your rules.') rather than shared multi-tenant serving
B200 SXM6 $6.69/GPU/hr, H100 SXM $3.99/GPU/hr; 1-Click H100 clusters from $6.16/hr; SOC 2 Type II. The Inference API wind-down is a notable ecosystem exit from …
Kubernetes-native AI hyperscaler serving OpenAI, Mistral, and IBM
GPU cloudsper-hour GPU + reserved capacity (up to 60% off on-demand); spot availablepartialHyperscale enterprise posture β€” the infrastructure behind OpenAI and other frontier labs, with claimed 96% cluster goodput
HGX B200 8-GPU $68.80/hr, GB200 NVL72 (4 GPU) $42.00/hr; dedicated inference tier H100 at $6.16/hr. Customers: OpenAI, Mistral AI, IBM, Jane Street, Cloudflare,…
GPU marketplace where prices are set by the market, not the cloud
GPU cloudsmarket-set per-second GPU: on-demand, interruptible (50%+ cheaper), reserved terms up to 50% offown APICheapest route to raw GPUs via marketplace competition
68+ GPU types, 40+ data centers, per-second billing, no lock-in; interruptible tier suits batch workloads; live rates published only in the console.
The Hub as a universal front door β€” one API routing to many inference providers
Gateways & routerspass-through per-token (Providers) + per-hour dedicated (Endpoints)OpenAI-compatThe Hub network effect: the largest open-model catalog, one account, many serving backends
Providers routing includes the same serverless APIs listed here; Endpoints handles autoscaling incl. scale-to-zero.
Open models per-token inside the AWS compliance envelope
Serverless model APIsper-token (on-demand) + provisioned throughputpartialWhere regulated enterprises already living on AWS meet open models β€” procurement and compliance included
Model access is region-scoped; Bedrock also fronts guardrails, agents and knowledge bases.
Microsoft's model catalog with serverless open-model endpoints
Serverless model APIsper-token (serverless) + managed computepartialThe Azure enterprise path to open models β€” same subscription, same governance as the OpenAI service
Catalog spans first-party, partner and open models under one deployment surface.
Sign in to continue

LLM Switchboard is private β€” sign in with Authlee to access the control room.

Sign in with Authlee
← Back to home