route-switch vs LiteLLM vs Portkey: LLM Gateways in 2026

Comparing route-switch, LiteLLM, Portkey and OpenRouter: what each gateway routes, what it records, and where MIPROv2 prompt optimisation fits in.

The question

Should I use route-switch, LiteLLM, Portkey, or OpenRouter in front of my LLM calls?

All four are gateways: they put one API in front of several model providers, so application code stops being welded to a particular SDK. They differ in what they do with the traffic once it passes through. This is the comparison we wish we had when we started route-switch.

The short version: LiteLLM is about breadth, Portkey about observability and guardrails, OpenRouter about not running anything, and route-switch about closing the loop on prompts. route-switch routes by a strategy you configure, records how each prompt performs, and reruns MIPROv2 prompt optimisation against the traffic it captured.

What each option is

route-switch is an OpenAI-compatible LLM gateway written in Go, self-hosted and MIT-licensed. It holds a registry of combinations — tuples of prompt template, model, provider, weight and fallbacks — and routes each request across them with a configured strategy: round_robin, weighted or performance_based (which leans on observed success rate and response time), with fallback when a combination’s success rate drops below a threshold. Providers are reached through the gollm adapter: OpenAI, Anthropic, Google, Ollama, Cohere, Mistral and others. Every call is logged to a per-prompt SQLite dataset and to a DuckDB analytics store recording success, cost and latency. On demand or on a schedule, it reruns MIPROv2 — instruction and few-shot search driven by goptuna’s Bayesian optimiser — against those captured traces and writes the winning prompt back to the registry.

LiteLLM is an open-source Python SDK and proxy server that gives an OpenAI-compatible interface to a very large number of providers (over 100, by its own count). Its router supports load balancing and fallbacks, and the proxy adds virtual keys and spend tracking. It is the common default for “call any LLM from any client”.

Portkey is an AI gateway with an open-source core and a hosted product. Its strengths are observability — traces, logs, dashboards, cost breakdowns — plus guardrails, a prompt library and response caching.

OpenRouter is a managed service: one API key and one bill for models from many providers, with provider failover. There is nothing to deploy.

The dimensions

Dimensionroute-switchLiteLLMPortkeyOpenRouter
ShapeSelf-hosted Go binaryPython SDK or self-hosted proxyHosted, with open-source gatewayManaged service
RoutingStrategy across registered combinationsRouter config (load balancing, fallbacks)Configs (fallbacks, load balancing, conditional)Provider selection and failover
Prompt optimisationMIPROv2 on captured tracesNot a core featureNot the headline featureNo
AnalyticsDuckDB: per-prompt success, cost, latencySpend tracking, logging integrationsFirst-class dashboardsUsage and billing dashboard
GuardrailsOut of scopeVia integrationsBuilt inNot the focus
Response cachingNoYesYesProvider-dependent
Provider coverageOpenAI, Anthropic, Google, Ollama, Cohere, Mistral and others via gollmVery broadBroadVery broad
LicenseMITOpen-source core; enterprise features commercialOpen-source gateway; managed product commercialProprietary service
CostFree to self-hostFree to self-hostFree self-host; paid cloudPay per token

route-switch publishes no synthetic latency, cost or quality numbers, and neither do we here. The analytics endpoints exist precisely so you can measure your own deployment.

When to use which

Use route-switch when:

  • The recurring problem is that prompts drift and nobody re-evaluates them, and you want the optimiser in the same binary as the gateway.
  • You want per-prompt success, cost and latency in a local analytics store rather than a hosted dashboard.
  • You want a single self-hosted binary with local state files.
  • You already have, or will soon have, real traffic to capture.

Use LiteLLM when:

  • You need the widest provider coverage behind one OpenAI-compatible API.
  • You are working in Python, or want the larger community and ecosystem.
  • You are prototyping and want the lowest setup cost.

Use Portkey when:

  • Observability and guardrails are at the top of the list: PII detection, content moderation, pre-built spend dashboards.
  • Non-engineers need a UI to manage prompts.
  • You want semantic or simple response caching at the gateway.

Use OpenRouter when:

  • You do not want to run infrastructure at all.
  • You want one key and one bill across many providers, and are happy to pay per token for that.

Routing and optimisation are separate

The earlier version of this article described route-switch as a router that learns which model to send each query to. That was wrong, and the distinction matters.

route-switch is not a semantic router. It does not classify arbitrary user queries and pick a model for each. It routes across combinations you registered, by the strategy you configured. The closest thing to adaptation is performance_based, which favours combinations with better observed success and response time, and fallback, which moves traffic away from a combination whose success rate has dropped.

What MIPROv2 changes is the prompt. For a registered template, it bootstraps a calibration sample from captured calls, proposes candidate instructions, searches over instruction and few-shot demonstration sets, scores candidates by replaying captured rows under an evaluation strategy (similarity, exact match, keyword match, or your own Go implementation), and writes the winner back. That is per prompt and per model — not query rewriting, not training, not RLHF.

The “routing triangle” post on the route-switch site frames the result well: you choose which of quality, cost and latency is the constraint, and encode that choice as a strategy, weights and a fallback threshold. Optimisation does not change which corner you aim for; it can make the prompts at that corner better at your task. Whether it lowers your bill depends on your traffic, your evaluator and your configuration — measure it.

Where route-switch loses

  • No traffic yet. Optimisation against captured traces is worth nothing until there are traces. For a prototype, use LiteLLM or OpenRouter.
  • No quality signal. MIPROv2 scores candidates against an evaluation strategy. If none of the built-in ones fits your task and you will not write one, the optimiser has no target.
  • You need guardrails, caching or a polished dashboard. None of those is in scope. Portkey, or a dedicated layer in front of route-switch, is the right shape.
  • Provider breadth. LiteLLM covers more providers and has a much larger community.
  • Latency budgets. performance_based balances on average response time; you cannot write “p95 under 800 ms” in the config. That enforcement lives upstream today.
  • Drift. A prompt optimised against last month’s traffic is not necessarily optimal against this month’s, and there is no principled re-optimisation trigger yet beyond a schedule.

Trying it

The route-switch quickstart, unchanged:

# build
git clone https://github.com/Skelf-Research/route-switch && cd route-switch
go build -o route-switch ./...

# run the gateway
./route-switch --gateway --config config.yaml

# send a request with your existing OpenAI client
curl http://localhost:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model":"gpt-4","messages":[{"role":"user","content":"hi"}]}'

Once traffic is flowing, /v1/prompts/{id}/stats and /v1/system/analytics show per-prompt and aggregate success, cost and latency. Run the optimiser on demand with --optimize-prompt, or enable the background optimiser with gateway.optimization.enabled: true.