New: the AI Gateway — one endpoint for Cursor, Cline, Codex & every OpenAI tool

The right model for every job —
picked automatically.

LLM Switchboard is a router, an OpenAI-compatible gateway and a control room in one zero-dependency app. Point any AI tool at /v1, say model: "auto", and every request is scored, routed and executed across Groq, NVIDIA and Cloudflare — with quota-aware auto-fallback and ≈46M documented free tokens a month when you run on auto/free. Cloud or local. No lock-in. Cut your AI bill 40–80% without your customers noticing.

Sign in with Authlee → Open the Gateway
53cloud models
303+run locally < 25GB
≈46Mfree tokens/mo (documented)
31architectures explained
18routing strategies mapped

Always improving — on its own

New capabilities land every week, and the Trending feed even curates itself daily. What's freshest right now:

Router

Describe the job in plain words; get the right model with a transparent score — and run it.

● updated yesterday

RSI Lab

Watch a model improve its own work under a frozen judge — live, loop by loop, on real models.

● updated yesterday

Benchmarks

Leaderboards you can trust — we grade the benchmarks themselves.

● updated yesterday

Benchmark papers

A live radar of new benchmark research with trend intelligence.

● updated yesterday

Vector databases

Pick the right memory layer for RAG and agents, from pgvector to zvec.

● updated yesterday

Architectures

The LLM architecture field guide — every major design with a diagram, verified references and fresh deep dives.

● updated 4 days ago

See everything that’s new →

One endpoint. Every tool. Never stop coding.

OpenAI-compatible /v1

Point Cursor, Cline, Codex CLI, Continue or any OpenAI SDK at <origin>/v1. Your Groq or NVIDIA key doubles as the API key — auto-detected from the Bearer token. Streaming, tools and JSON mode pass through untouched.

🎯

auto · coding · fast · cheap · smart

Say model: "auto" and the router classifies your request and picks the best model — or pin the policy with auto/coding, auto/fast, auto/cheap, auto/smart.

🆓

≈46M free tokens/month

auto/free routes only across documented free tiers, draining the pool with the most headroom first — Groq's 1.4M free tokens/day + Cloudflare's daily neurons, verified from provider docs. Built for long-horizon agent runs.

Quota-aware fallback

Rate-limited models cool down honoring Retry-After, dead ones get benched, the next-ranked candidate answers — your task keeps running. Watch it live on the Connections page.

🔌

Connections manager

OmniRoute-style control panel: connect keys, live-test them, toggle providers in and out of the chain, and run a real request to see which model answers.

🔀

18 routing strategies, animated

Priority, headroom, cost-optimized, fusion, pipeline and more — every flow animated, with an honest map of which run here today.

Your team can't afford an ML platform group — but the model landscape changes weekly. LLM Switchboard is the opinionated middle layer: it knows the models, grades the benchmarks, and routes the traffic, so you ship instead of researching.

One control room for your whole AI stack

Smart router

Classifies each job, filters by your constraints, and ranks models with a transparent score. Callable as a REST API or an importable module.

AI compliance, dated

EU AI Act Article 50 has applied since 2 August 2026 — disclosure and content marking are features now, not policy PDFs. A timeline of the dates that actually bind, what each means in engineering terms, and how much is enforceable as routing policy. Every date cited to EUR-Lex, the Commission, CEN-CENELEC or a state legislature.

💰

Cost Lab

Describe a workload and see real monthly spend across every priced model, how much the documented free tiers absorb, and the honest saving between the cheapest model that clears your quality bar and a frontier one. The sliders move the numbers — no vendor maths.

Architecture field guide

31 LLM architectures — dense, MoE, MLA, Mamba hybrids, multimodal, diffusion — each with a block diagram, research-verified arXiv references, and an atlas mapping 20 model families version-by-version (fresh deep dives: Kimi K3 & GLM-5.2).

🔥

Self-updating Trending

Every day, the best Groq model reads github.com/trending, curates what matters, and publishes it to the app on its own — every update dated. No human in the loop.

RSI Lab

Recursive self-improvement you can watch: a frozen mental model judges while the artifact evolves — with an algorithm atlas, domain examples and a live multi-loop runner on Groq/NVIDIA.

Run locally

303+ open models under 25GB for reasoning, coding, vision, STT, TTS & embeddings — copy-paste Ollama/Docker commands, picked by benchmark, with a built-in sandbox to test them.

Business recipes

MEDDIC & BANT call analysis, website SDR chat, AI voice SDR, compliance triage — each with the routed model, an example result, and a live test.

Benchmark intelligence

We grade the benchmarks themselves with the Benchmark² framework, so you trust the right signal — not just whoever topped a leaderboard.

Provider directory

25 inference providers — serverless APIs, serving platforms, GPU clouds, gateways, edge — URL-verified, with the wired ones manageable from Connections.

Always current

A freshness pipeline pulls new open-source models from free feeds, and every feature shows exactly when it was added and last improved — history back to day one.

Harnesses & skills

A buyer's guide to LangGraph, CrewAI, DeerFlow and the SDKs, plus the portable SKILL.md ecosystem (NVIDIA-verified, cybersecurity, gstack).

Vector databases

From pgvector to Milvus to Alibaba's zvec — pick the right memory layer for RAG and agents.

How it works

1

Describe the job

Send a prompt or pick a recipe. LLM Switchboard classifies it into one of 18 job types and reads your constraints (budget, latency, context, modality).

2

It picks the model

The engine filters the catalog and scores every candidate on task-fit, cost and speed — then explains the choice with capability scores and the relevant benchmarks.

3

You route to it — or your tools do

Get the decision over REST, execute on Groq/NVIDIA with automatic cross-provider fallback — or skip the integration entirely: point any OpenAI-compatible tool at /v1 and the gateway does all of it per request.

Real jobs, real picks

◎ MEDDIC deal scorer

Turn a sales-call transcript into a MEDDIC scorecard with a deal-health score. → routes to a frontier reasoning model with 128k+ context.

✸ Website SDR chat

Real-time lead qualification on your homepage. → routes to the fastest Groq model — sub-second, nearly free at scale.

☎ AI voice SDR

Outbound voice that books meetings. → a Whisper → LLM → Orpheus pipeline, all on low-latency infra.

⛉ Compliance triage

Classify SOC 2 / ISO evidence at scale. → routes to a cheap fast classifier; pennies for the whole pile.

∞ Long-horizon agent, $0

An overnight agent run that never stops. model: "auto/free" drains the free pools by headroom; 429s cool down, the next pool answers.

⌨ Your IDE, smarter

Cursor / Cline / Codex on router-picked models. → base URL /v1, your Groq key as the API key — done in one settings change.

Sign in to explore the full recipe library →

Stop guessing. Start routing.

Sign in with your iCompaas account to open the control room — the app is private, every page behind Authlee SSO.

Sign in with Authlee →
running build dfc576a · 2026-08-22