route-switch vs LiteLLM vs Portkey: LLM Gateways in 2026
Comparing route-switch, LiteLLM, Portkey and OpenRouter: what each gateway routes, what it records, and where MIPROv2 prompt optimisation fits in.
The question
Should I use route-switch, LiteLLM, Portkey, or OpenRouter in front of my LLM calls?
All four are gateways: they put one API in front of several model providers, so application code stops being welded to a particular SDK. They differ in what they do with the traffic once it passes through. This is the comparison we wish we had when we started route-switch.
The short version: LiteLLM is about breadth, Portkey about observability and guardrails, OpenRouter about not running anything, and route-switch about closing the loop on prompts. route-switch routes by a strategy you configure, records how each prompt performs, and reruns MIPROv2 prompt optimisation against the traffic it captured.
What each option is
route-switch is an OpenAI-compatible LLM gateway written in
Go, self-hosted and MIT-licensed. It holds a registry of
combinations — tuples of prompt template, model, provider,
weight and fallbacks — and routes each request across them with
a configured strategy: round_robin, weighted or
performance_based (which leans on observed success rate and
response time), with fallback when a combination’s success rate
drops below a threshold. Providers are reached through the gollm
adapter: OpenAI, Anthropic, Google, Ollama, Cohere, Mistral and
others. Every call is logged to a per-prompt SQLite dataset and
to a DuckDB analytics store recording success, cost and latency.
On demand or on a schedule, it reruns MIPROv2 — instruction and
few-shot search driven by goptuna’s Bayesian optimiser — against
those captured traces and writes the winning prompt back to the
registry.
LiteLLM is an open-source Python SDK and proxy server that gives an OpenAI-compatible interface to a very large number of providers (over 100, by its own count). Its router supports load balancing and fallbacks, and the proxy adds virtual keys and spend tracking. It is the common default for “call any LLM from any client”.
Portkey is an AI gateway with an open-source core and a hosted product. Its strengths are observability — traces, logs, dashboards, cost breakdowns — plus guardrails, a prompt library and response caching.
OpenRouter is a managed service: one API key and one bill for models from many providers, with provider failover. There is nothing to deploy.
The dimensions
| Dimension | route-switch | LiteLLM | Portkey | OpenRouter |
|---|---|---|---|---|
| Shape | Self-hosted Go binary | Python SDK or self-hosted proxy | Hosted, with open-source gateway | Managed service |
| Routing | Strategy across registered combinations | Router config (load balancing, fallbacks) | Configs (fallbacks, load balancing, conditional) | Provider selection and failover |
| Prompt optimisation | MIPROv2 on captured traces | Not a core feature | Not the headline feature | No |
| Analytics | DuckDB: per-prompt success, cost, latency | Spend tracking, logging integrations | First-class dashboards | Usage and billing dashboard |
| Guardrails | Out of scope | Via integrations | Built in | Not the focus |
| Response caching | No | Yes | Yes | Provider-dependent |
| Provider coverage | OpenAI, Anthropic, Google, Ollama, Cohere, Mistral and others via gollm | Very broad | Broad | Very broad |
| License | MIT | Open-source core; enterprise features commercial | Open-source gateway; managed product commercial | Proprietary service |
| Cost | Free to self-host | Free to self-host | Free self-host; paid cloud | Pay per token |
route-switch publishes no synthetic latency, cost or quality numbers, and neither do we here. The analytics endpoints exist precisely so you can measure your own deployment.
When to use which
Use route-switch when:
- The recurring problem is that prompts drift and nobody re-evaluates them, and you want the optimiser in the same binary as the gateway.
- You want per-prompt success, cost and latency in a local analytics store rather than a hosted dashboard.
- You want a single self-hosted binary with local state files.
- You already have, or will soon have, real traffic to capture.
Use LiteLLM when:
- You need the widest provider coverage behind one OpenAI-compatible API.
- You are working in Python, or want the larger community and ecosystem.
- You are prototyping and want the lowest setup cost.
Use Portkey when:
- Observability and guardrails are at the top of the list: PII detection, content moderation, pre-built spend dashboards.
- Non-engineers need a UI to manage prompts.
- You want semantic or simple response caching at the gateway.
Use OpenRouter when:
- You do not want to run infrastructure at all.
- You want one key and one bill across many providers, and are happy to pay per token for that.
Routing and optimisation are separate
The earlier version of this article described route-switch as a router that learns which model to send each query to. That was wrong, and the distinction matters.
route-switch is not a semantic router. It does not classify
arbitrary user queries and pick a model for each. It routes
across combinations you registered, by the strategy you
configured. The closest thing to adaptation is
performance_based, which favours combinations with better
observed success and response time, and fallback, which moves
traffic away from a combination whose success rate has dropped.
What MIPROv2 changes is the prompt. For a registered template, it bootstraps a calibration sample from captured calls, proposes candidate instructions, searches over instruction and few-shot demonstration sets, scores candidates by replaying captured rows under an evaluation strategy (similarity, exact match, keyword match, or your own Go implementation), and writes the winner back. That is per prompt and per model — not query rewriting, not training, not RLHF.
The “routing triangle” post on the route-switch site frames the result well: you choose which of quality, cost and latency is the constraint, and encode that choice as a strategy, weights and a fallback threshold. Optimisation does not change which corner you aim for; it can make the prompts at that corner better at your task. Whether it lowers your bill depends on your traffic, your evaluator and your configuration — measure it.
Where route-switch loses
- No traffic yet. Optimisation against captured traces is worth nothing until there are traces. For a prototype, use LiteLLM or OpenRouter.
- No quality signal. MIPROv2 scores candidates against an evaluation strategy. If none of the built-in ones fits your task and you will not write one, the optimiser has no target.
- You need guardrails, caching or a polished dashboard. None of those is in scope. Portkey, or a dedicated layer in front of route-switch, is the right shape.
- Provider breadth. LiteLLM covers more providers and has a much larger community.
- Latency budgets.
performance_basedbalances on average response time; you cannot write “p95 under 800 ms” in the config. That enforcement lives upstream today. - Drift. A prompt optimised against last month’s traffic is not necessarily optimal against this month’s, and there is no principled re-optimisation trigger yet beyond a schedule.
Trying it
The route-switch quickstart, unchanged:
# build
git clone https://github.com/Skelf-Research/route-switch && cd route-switch
go build -o route-switch ./...
# run the gateway
./route-switch --gateway --config config.yaml
# send a request with your existing OpenAI client
curl http://localhost:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model":"gpt-4","messages":[{"role":"user","content":"hi"}]}'
Once traffic is flowing, /v1/prompts/{id}/stats and
/v1/system/analytics show per-prompt and aggregate success,
cost and latency. Run the optimiser on demand with
--optimize-prompt, or enable the background optimiser with
gateway.optimization.enabled: true.
What to read next
- Intelligent LLM Routing: Spending Compute Where It Matters — route-switch background
- Formalising Prompts as First-Class Research Objects — the promptel side of prompt engineering
- Quality / cost / latency: the routing triangle
- route-switch repository
- LiteLLM
- Portkey
- OpenRouter