LLM Switchboard is a router, an OpenAI-compatible gateway and a control room in one zero-dependency app. Point any AI tool at /v1, say model: "auto", and every request is scored, routed and executed across Groq, NVIDIA and Cloudflare — with quota-aware auto-fallback and ≈46M documented free tokens a month when you run on auto/free. Cloud or local. No lock-in. Cut your AI bill 40–80% without your customers noticing.
New capabilities land every week, and the Trending feed even curates itself daily. What's freshest right now:
Describe the job in plain words; get the right model with a transparent score — and run it.
● updated yesterday
Watch a model improve its own work under a frozen judge — live, loop by loop, on real models.
● updated yesterday
Leaderboards you can trust — we grade the benchmarks themselves.
● updated yesterday
A live radar of new benchmark research with trend intelligence.
● updated yesterday
Pick the right memory layer for RAG and agents, from pgvector to zvec.
● updated yesterday
The LLM architecture field guide — every major design with a diagram, verified references and fresh deep dives.
● updated 4 days ago
Point Cursor, Cline, Codex CLI, Continue or any OpenAI SDK at <origin>/v1. Your Groq or NVIDIA key doubles as the API key — auto-detected from the Bearer token. Streaming, tools and JSON mode pass through untouched.
Say model: "auto" and the router classifies your request and picks the best model — or pin the policy with auto/coding, auto/fast, auto/cheap, auto/smart.
auto/free routes only across documented free tiers, draining the pool with the most headroom first — Groq's 1.4M free tokens/day + Cloudflare's daily neurons, verified from provider docs. Built for long-horizon agent runs.
Rate-limited models cool down honoring Retry-After, dead ones get benched, the next-ranked candidate answers — your task keeps running. Watch it live on the Connections page.
OmniRoute-style control panel: connect keys, live-test them, toggle providers in and out of the chain, and run a real request to see which model answers.
Priority, headroom, cost-optimized, fusion, pipeline and more — every flow animated, with an honest map of which run here today.
Your team can't afford an ML platform group — but the model landscape changes weekly. LLM Switchboard is the opinionated middle layer: it knows the models, grades the benchmarks, and routes the traffic, so you ship instead of researching.
Classifies each job, filters by your constraints, and ranks models with a transparent score. Callable as a REST API or an importable module.
EU AI Act Article 50 has applied since 2 August 2026 — disclosure and content marking are features now, not policy PDFs. A timeline of the dates that actually bind, what each means in engineering terms, and how much is enforceable as routing policy. Every date cited to EUR-Lex, the Commission, CEN-CENELEC or a state legislature.
Describe a workload and see real monthly spend across every priced model, how much the documented free tiers absorb, and the honest saving between the cheapest model that clears your quality bar and a frontier one. The sliders move the numbers — no vendor maths.
31 LLM architectures — dense, MoE, MLA, Mamba hybrids, multimodal, diffusion — each with a block diagram, research-verified arXiv references, and an atlas mapping 20 model families version-by-version (fresh deep dives: Kimi K3 & GLM-5.2).
Every day, the best Groq model reads github.com/trending, curates what matters, and publishes it to the app on its own — every update dated. No human in the loop.
Recursive self-improvement you can watch: a frozen mental model judges while the artifact evolves — with an algorithm atlas, domain examples and a live multi-loop runner on Groq/NVIDIA.
303+ open models under 25GB for reasoning, coding, vision, STT, TTS & embeddings — copy-paste Ollama/Docker commands, picked by benchmark, with a built-in sandbox to test them.
MEDDIC & BANT call analysis, website SDR chat, AI voice SDR, compliance triage — each with the routed model, an example result, and a live test.
We grade the benchmarks themselves with the Benchmark² framework, so you trust the right signal — not just whoever topped a leaderboard.
25 inference providers — serverless APIs, serving platforms, GPU clouds, gateways, edge — URL-verified, with the wired ones manageable from Connections.
A freshness pipeline pulls new open-source models from free feeds, and every feature shows exactly when it was added and last improved — history back to day one.
A buyer's guide to LangGraph, CrewAI, DeerFlow and the SDKs, plus the portable SKILL.md ecosystem (NVIDIA-verified, cybersecurity, gstack).
From pgvector to Milvus to Alibaba's zvec — pick the right memory layer for RAG and agents.
Send a prompt or pick a recipe. LLM Switchboard classifies it into one of 18 job types and reads your constraints (budget, latency, context, modality).
The engine filters the catalog and scores every candidate on task-fit, cost and speed — then explains the choice with capability scores and the relevant benchmarks.
Get the decision over REST, execute on Groq/NVIDIA with automatic cross-provider fallback — or skip the integration entirely: point any OpenAI-compatible tool at /v1 and the gateway does all of it per request.
Turn a sales-call transcript into a MEDDIC scorecard with a deal-health score. → routes to a frontier reasoning model with 128k+ context.
Real-time lead qualification on your homepage. → routes to the fastest Groq model — sub-second, nearly free at scale.
Outbound voice that books meetings. → a Whisper → LLM → Orpheus pipeline, all on low-latency infra.
Classify SOC 2 / ISO evidence at scale. → routes to a cheap fast classifier; pennies for the whole pile.
An overnight agent run that never stops. → model: "auto/free" drains the free pools by headroom; 429s cool down, the next pool answers.
Cursor / Cline / Codex on router-picked models. → base URL /v1, your Groq key as the API key — done in one settings change.
Sign in with your iCompaas account to open the control room — the app is private, every page behind Authlee SSO.
Sign in with Authlee →