self-hosted ai gateway

One endpoint in front of every LLM you run

LLM Router sits between your services and your models — vLLM, Ollama, llama.cpp, LM Studio, OpenAI-compatible APIs and Anthropic — behind a single OpenAI-compatible REST surface. Auth, sliding-window rate limits, PII masking, guardrails, load balancing, provider failover and Prometheus metrics run in your infrastructure, so prompts never take a detour through someone else's cloud.

$pip install "radlab-llm-router[api]"copied
  • python ≥ 3.10
  • image quay.io/radlab/llm-router
  • helm chart included
  • no vendor lock-in
  • zero telemetry
ops@bastion: ~/infra/llm-router
$ pip install "radlab-llm-router[api,metrics]"
✓ installed llm-router 1.1.6 · console script llm-router
 
$ llm-router config discover localhost 10.0.4.21 -o models-config.json
✓ discovered ollama    http://localhost:11434  (3 models)
✓ discovered vllm     http://10.0.4.21:8000 (1 model)
‐ wrote models-config.json — review before prod
 
$ llm-router server start
✓ All good
• Server is running (gunicorn, 4 workers)
 
  Details
    PID                41207
    Server          gunicorn
    Host            0.0.0.0:8080
    Models config   /srv/models-config.json
    Log            ~/.llm-router/server.log
 
$ curl -s localhost:8080/api/ping
{"status": true, "body": "pong"}  HTTP/1.1 200 OK
6
provider types
6
LB strategies
30+
PII rule types
25+
REST endpoints
~95
env vars
100%
self-hosted
01 / REQUEST PATH

Deterministic pipeline, not a proxy pile-up

Every request walks the same ordered path. Each stage is a registry-backed plugin with a tiny apply(), so you can drop a stage, add your own, or reorder the pipeline without touching application code. EndpointAutoLoader discovers EndpointI implementations at startup — new routes appear with no wiring.

7 STAGES · SCROLL →
INBOUND
client
OpenAI SDK, curl, LangChain, LlamaIndex, your service
→
MIDDLEWARE
AuthMiddleware
Bearer key → key store → permission engine → rate limiter
→
STAGE 1
MaskerPipeline
fast_masker regex, then pii_masker ML token classifier
→
STAGE 2
GuardrailPipeline
classifier check → pass or block with 4xx
→
STAGE 3
UtilsPipeline
RAG enrichment, built-in tasks, custom endpoints
→
RESOLVE
load_balancer
strategy pick, provider lease, retry & failover
→
UPSTREAM
model provider
vLLM, Ollama, llama.cpp, LM Studio, OpenAI-compatible, Anthropic
plugin registry — opt in per stage via env
reload — llm-router server reload
every stage — llm_router_pipeline_stage_total{stage,result}
fail closed — guardrail block never reaches a provider
03 / PLUGINS

Pluggable processing pipeline

Every request walks through a configurable pipeline of masker, guardrail and routing plugins. Each plugin is a small, well-defined class that implements apply() and can be chained with others. Plugins live in the separate llm-router-plugins repository and are loaded by the router via its plugin registry.

01 Masker plugins

Strip personally identifiable information before it reaches a model.

  • fast_masker — 30+ regex rules: email, IP, phone, IBAN, JWT, Polish identifiers (PESEL, NIP, KRS, REGON)
  • pii_masker — ML-based token classification (RoBERTa) for context-aware PII detection in Polish and English

Enable via LLM_ROUTER_FORCE_MASKING=1 plus LLM_ROUTER_MASKING_STRATEGY_PIPELINE="fast_masker,pii_masker"

02 Guardrail plugins

Validate requests and responses against policy rules before spend.

  • nask_guard — Polish content safety (HerBERT-PL-Guard model via HTTP service)
  • sojka_guard — Bielik-Guard-0.1B multi-category safety check
  • Fail-closed mode: LLM_ROUTER_FORCE_GUARDRAIL_REQUEST=1 checks every request

Enable via LLM_ROUTER_FORCE_GUARDRAIL_REQUEST=1 plus LLM_ROUTER_GUARDRAIL_STRATEGY_PIPELINE_REQUEST="nask_guard"

03 Codex CLI routing

Route OpenAI Responses-style requests from the Codex coding agent based on task type.

  • Recognises plan, implement, test, debug and review modes from collaboration metadata
  • Routes plan mode to a fast reasoning model, implementation to a larger model
  • Uses auto_codex as a stable model alias — no per-developer config needed
  • Ranks real work above Codex housekeeping — compaction and auxiliary title calls are classified apart from main turns

Enable: LLM_ROUTER_UTILS_PLUGINS_PIPELINE="agentic_routing_codex"

Codex config: model = "auto_codex" in ~/.codex/config.toml

04 Claude Code model swap

Rewrite Anthropic tier model names (Fable, Opus, Sonnet, Haiku, Plan Mode) to the models your router actually serves.

  • Maps claude-fable-*, claude-opus-*, claude-sonnet-*, claude-haiku-* and opusplan to local or custom models
  • Handles Claude Code's internal subagent model selection (CLAUDE_CODE_SUBAGENT_MODEL)
  • One JSON config file on the router — no environment variables on every client host
  • Fail-open: unmatched model names pass through unchanged

Enable: LLM_ROUTER_UTILS_PLUGINS_PIPELINE="agentic_routing_claude_code"

Point Claude Code at router: ANTHROPIC_BASE_URL and ANTHROPIC_AUTH_TOKEN

05 Semantic routing

Embedding-based routing that matches user messages against pre-configured targets.

  • simple_semantic_routing — two-stage heuristic: intent classification + complexity analysis
  • semantic_biencoder_routing — embedding model + FAISS index for similarity-based routing
  • Configure targets with similarity thresholds and fallback models
  • Works with any OpenAI-compatible embedding endpoint

06 RAG enrichment

Augment requests with relevant context before model inference.

  • langchain_rag — LangChain-based retrieval with custom vector stores
  • Index external documentation and knowledge bases
  • CLI tools for indexing and searching: llm-router-rag-langchain-index
02 / FEATURES

The boring infrastructure you needed anyway

Everything below is off by default until you switch it on with an env var or a config file. No sidecar, no external control plane, no phone home.

One OpenAI-compatible surface

Point existing clients at the router and keep them there: /v1/chat/completions, /v1/responses, /v1/embeddings plus native /v1/messages for Anthropic. SSE streaming honors the stream flag with correct cache headers.

openaianthropicvllmollamallama.cpplm studio

Keys, policies, permissions

Issue scoped keys as sk-llmr-live-<base62> against a memory, Redis or Vault backing store. Rotate, disable and delete from the CLI; map keys to endpoint permissions in the permission engine.

memoryredisvaultrotate

Sliding-window rate limiting

Redis sorted-set windows tracked per key and client IP, so bursts get smoothed instead of punished. 429 responses carry Retry-After; X-Forwarded-For is only trusted from LLM_ROUTER_TRUSTED_PROXIES.

redis zsetper key+ipretry-after

PII masking, two tiers

fast_masker applies 30+ regex rule types — email, IP, phone, IBAN-ish amounts, JWTs, and Polish identifiers such as PESEL, NIP, KRS, REGON. pii_masker adds a remote ML token classifier with an in-memory cache.

regexml classifiercached

Guardrails before spend

Run nask_guard (HerBERT-PL-Guard) or sojka_guard (Bielik-Guard-0.1B) in front of the model call, so blocked prompts never burn GPU time or land in an upstream provider's logs. Force it globally with LLM_ROUTER_FORCE_GUARDRAIL_REQUEST=1.

pass | blockherbertbielik

Semantic routing on model:"auto"

Send "model": "auto" and let simple_semantic_routing judge intent and complexity, or use semantic_biencoder_routing with EmbeddingGemma-300m and a FAISS IndexFlatIP index over your model cards.

faissintentcomplexityindex persists

Six load-balancing strategies

From in-memory balanced to latency-aware dynamic_weighted, down to Redis-backed first_available with exclusive leases when a model must be pinned to one worker — or per-provider worker slots with first_available_optim_nworkers. Shared state across replicas where it matters.

balancedweighteddynamicredis lease

Prometheus, already wired

Set LLM_ROUTER_USE_PROMETHEUS=1 and scrape /metrics: provider calls, latency histograms, errors, retries, token counters and per-stage pipeline results. A Grafana dashboard JSON ships in the image.

/metricshistogramsgrafana json

Extend, don't fork

Maskers, guardrails, routers and endpoints are registered plugins. Subclass, implement apply(), ship it in llm-router-plugins, and the loader picks it up. Typed sync & async SDK in llm-router-lib; integration boilerplates for LangChain, LlamaIndex, Haystack and LiteLLM.

plugin registryapache-2.0
03 / CONFIG

One JSON file, one env file.
That's the whole control plane.

Providers live in models-config.json — point LLM_ROUTER_MODELS_CONFIG at it, mount it read-only, and reload the process. Behaviour lives in ~95 LLM_ROUTER_* variables, all documented in ENV_DEFINITIONS.md and all settable from a plain env file, systemd unit or ConfigMap. No database to migrate, no admin SaaS to trust.

/srv/llm-router/models-config.json — groups → model → providers[] / providers_sleep[] / fallback_model
{
  "google_models": {
    "google/gemma-3-27b-it": {
      "providers": [
        {
          "id": "gemma27b-vllm-node-01:8000",
          "api_host": "http://10.0.4.21:8000/",
          "api_token": "",
          "api_type": "vllm",
          "input_size": 56000,
          "weight": 2.0,
          "nworkers": 4,
          "tool_calling": true
        },
        {
          "id": "gemma27b-ollama-lab",
          "api_host": "http://localhost:11434",
          "api_type": "ollama",
          "model_path": "gemma3:27b",
          "keep_alive": "10m",
          "weight": 1.0
        }
      ],
      // standby providers, used only when every primary is busy or down
      "providers_sleep": [
        { "id": "gemma27b-vllm-node-02:8000",
          "api_host": "http://10.0.4.22:8000/",
          "api_type": "vllm" }
      ],
      // last resort when no provider of this model can serve the request
      "fallback_model": "gpt-4o-mini"
    }
  },
  "openai_models": {
    "gpt-4o-mini": {
      "providers": [
        { "id": "openai-eu", "api_host": "https://api.openai.com/v1",
          "api_token": "${OPENAI_API_KEY}", "api_type": "openai" }
      ]
    }
  },
  "active_models": {
    "google_models": ["google/gemma-3-27b-it"],
    "openai_models": ["gpt-4o-mini"]   // the fallback target must be active as well
  }
}
drop-in for any OpenAI-compatible client — stream defaults to true
$ curl -s https://llm.internal:8080/v1/chat/completions \
    -H "Authorization: Bearer sk-llmr-live-8fQ2..." \
    -H "Content-Type: application/json" \
    -d '{
      "model": "auto",
      "messages": [{"role": "user", "content": "Summarise this incident report"}],
      "stream": true
    }'

# data: {"choices":[{"delta":{"content":"The outage"}}]}   ← SSE, chunked
# model "auto"    → semantic routing picks the provider for you
# "gpt-4o-mini"   → resolved via models-config.json + LB strategy
# /api/chat/completions, /chat/completions, /v1/messages,
# /responses, /v1/responses, /embeddings, /api/embed — same handler
/etc/llm-router/router.env — ordered pipeline, opt-in stages
# request pipeline: applied left to right (guardrail = request side)
LLM_ROUTER_MASKING_STRATEGY_PIPELINE="fast_masker,pii_masker"
LLM_ROUTER_GUARDRAIL_STRATEGY_PIPELINE_REQUEST="nask_guard"
LLM_ROUTER_UTILS_PLUGINS_PIPELINE="simple_semantic_routing"
LLM_ROUTER_BALANCE_STRATEGY="balanced"

# force the stages on every request (fail closed)
LLM_ROUTER_FORCE_MASKING=1
LLM_ROUTER_FORCE_GUARDRAIL_REQUEST=1
LLM_ROUTER_MASKING_WITH_AUDIT=1

# process & observability
LLM_ROUTER_MODELS_CONFIG="/srv/llm-router/models-config.json"
LLM_ROUTER_SERVER_TYPE="gunicorn"
LLM_ROUTER_SERVER_WORKERS_COUNT=4
LLM_ROUTER_USE_PROMETHEUS=1
LLM_ROUTER_TRUSTED_PROXIES="10.0.0.0/8"
llm-router auth — issue keys, bind policies, cap the blast radius
$ llm-router auth key generate --policy internal-only --store redis
✓ key-id key-3f9a1c72  token sk-llmr-live-8fQ2d7Xk...
‐ store redis · shown once · hash at rest

$ llm-router auth rate-limit apply key-3f9a1c72 --preset basic
✓ bucket auth:ratelimit:key-3f9a1c72:{ip} → sliding window

$ llm-router auth key list
key-3f9a1c72  payments-svc   enabled   models: gemma-3-27b, gpt-4o-mini
key-91bc0d45  batch-etl      disabled  rotated 2026-08-14

$ llm-router auth key rotate key-91bc0d45 --grace 3600  # grace overlap, no downtime

Auth is LLM_ROUTER_AUTH_ENABLED=false by default for the happy path on a private network. Switch it on and LLM_ROUTER_AUTH_PUBLIC_ENDPOINTS decides what stays key-free — out of the box that is only /health and /metrics. Everything else takes a key, and because matching happens on the raw request path, a probe on a prefixed route needs its own entry: LLM_ROUTER_AUTH_PUBLIC_ENDPOINTS="/metrics,/health,/api/ping".

04 / LOAD BALANCING

Pick a strategy, not a religion

Six strategies ship in the core. The difference is where state lives and how hard it pins a provider — selection counters in Redis so every gunicorn worker and replica agrees (in-process when Redis is absent), up to exclusive leases that hold a GPU-backed replica for one request, or worker-slot leases that let one provider absorb several parallel requests. Switch with one env var.

strategystate backendbehaviouruse it when
balanced default Redis counters, in-process fallback Picks the least-used provider; the counters are shared across workers and replicas and degrade to in-process when Redis is missing. Rotates on any upstream error. Stateless single replica, homogeneous GPUs.
weighted config weight + Redis Static ratios from models-config.json, walked off one global selection sequence so every replica hits the same CDF — send 2× traffic to the 8×GPU node. Known capacity differences between nodes.
dynamic_weighted beta in-process + latency Re-weights providers as latency and error signals drift; slow nodes lose share automatically. Mixed hardware, noisy neighbours, shared clusters.
first_available Redis exclusive lease One worker holds a provider at a time; everyone else moves on. No cold-cache thrash. Huge weights, KV-cache warm-up, one-process-per-GPU.
first_available_optim Redis :last_host / :hosts Same lease plus host bookkeeping (:last_host, :hosts, and host:<host> pinned for LLM_ROUTER_LB_HOST_PIN_TTL_SECONDS) so replicas reuse a warm host without re-scanning. Falls back to first_available. Many replicas, chatty provider sets.
first_available_optim_nworkers Redis :in_use slot leases Worker-slot version of first_available_optim: each provider serves up to nworkers concurrent requests (default 1). A least-loaded-with-a-free-slot step ranks providers by absolute busy-worker count, so they fill in layers instead of saturating the largest one. Leases live under their own fa_optim_nworkers_ prefix and expire, so a crashed router gives its slots back. Engines with real parallelism — vLLM / llama.cpp slots on one GPU.

$ Provider failover, not just retries

Any 4xx/5xx answer — and any transport error — is replayed on another provider of the same model, streaming and non-streaming alike. The providers a request has already failed on travel with it, so every attempt gets a fresh candidate, and the last provider error is what finally reaches the client. Streams fail over before the first byte: the first chunk is awaited, so a provider that rejects the stream is swapped before a truncated answer can start — a failure mid-stream, after bytes were delivered, is never replayed.

  • Tuned per endpoint — RETRY_ON_ANY_ERROR_STATUS is on, so rotation covers every error status; switch it off and only RETRY_WHEN_STATUS (429, 500, 502, 503, 504) retriggers
provider rotationstream-awareper-endpoint policy

↓ fallback_model, the last resort

A model can name another active model that takes over when none of its own providers can serve a request — none configured, none healthy, or every one busy until the selection timeout. The reroute happens before load balancing, so the very same strategy balances over the fallback model’s providers.

  • Chains — a → b → c; an unknown target, a self reference or a cycle aborts startup with a ValueError
  • Standby providers — providers_sleep keeps low-priority providers for the model itself, used when the primaries are busy or down
  • Counted — llm_router_model_fallback_total{model_name, fallback_model}
model-level fallbackvalidated at startupstrategy-agnostic

↻ Retry budget and backoff

Failover is bounded, not endless: MAX_RECONNECTIONS = 10 attempts with exponential backoff from 0.1 s, capped at 2.0 s, plus jitter. llm_router_retry_total{model_name,error_code} and llm_router_retry_exhausted_total{model_name,last_error_code} tell you when a model is genuinely out of capacity rather than briefly grumpy.

capped attemptsbackoff + jitterexhaustion counter

; Redis is optional, not required

Everything runs on memory stores until you need cross-replica state. Flip redis.enabled in the Helm chart (Bitnami subchart) or point at your own instance, and leases, rate-limit buckets and the key store move over. vault is available for key storage too.

memoryredisvault
05 / SECURITY & COMPLIANCE

Redact before you route,
log what you can prove

Two stages sit in front of every upstream call, and both fail closed. Nothing leaves the process that you did not explicitly allow, and the audit trail is encrypted at rest with GPG so a leaked log volume is not a data breach.

01 PII masking

fast_masker runs 30+ regex rule types synchronously; pii_masker adds a remote ML token classifier with a cache for the things regex never catches.

  • Deterministic rules — e-mail, IBAN, credit card, phone, MAC, IPv4/IPv6, JWT
  • PL paperwork — PESEL, NIP, KRS, REGON, ID card, passport, NRB bank account
  • ML pass — names, addresses and free-text identifiers, cached per text
  • Reversible — masking_with_audit keeps the mapping for response un-masking
fast_maskerpii_maskercustom rulesforce_masking

02 Guardrails

Local classifier models screen the request before a provider is ever contacted. A blocked request returns a 4xx and the prompt never leaves the cluster.

  • nask_guard — HerBERT-PL-Guard, Polish + English harmfulness
  • sojka_guard — Bielik-Guard-0.1B, small enough to run on CPU
  • Fail closed — LLM_ROUTER_FORCE_GUARDRAIL_REQUEST=1 checks every request
  • Your own — subclass the guardrail base and register it
herbert-pl-guardbielik-guardblock → 4xx

03 Audit trail, encrypted

Masking and guardrail decisions are written as GPG-encrypted entries under logs/auditor/. Key material is generated and exported by gen_and_export_gpg.sh; review with decrypt_auditor_logs.sh.

  • Who — key id, source IP, timestamp
  • What — stage, result, rule or classifier that fired
  • At rest — asymmetric encryption, no plaintext on disk
logs/auditor/
2026-09-09/audit-03.gpg   # 4.1k · encrypted
$ ./scripts/decrypt_auditor_logs.sh 2026-09-09

04 Boring, auditable defaults

Defence in depth for the parts that usually leak: error bodies, container user, and outbound traffic.

  • sanitize_error_message() strips IPs, hostnames, ports, URLs and stack traces from every response — detail stays in server logs
  • Non-root image — python:3.14-slim-trixie, uid/gid 5000, read-only friendly
  • No telemetry — the router makes no outbound calls except to providers you configure
  • Proxy-aware — LLM_ROUTER_TRUSTED_PROXIES controls who may set X-Forwarded-For
06 / OBSERVABILITY

Scrape one path. Alert on everything.

Turn on LLM_ROUTER_USE_PROMETHEUS=1 and /metrics exposes counters, histograms and per-stage pipeline results in text format. Multiprocess mode is wired for gunicorn via PROMETHEUS_MULTIPROC_DIR (~/.llm-router/metrics/prometheus/multiproc by default), and a Grafana dashboard JSON ships in the repo at resources/configs/prometheus/grafana-llm-router-dashboard-v1.json.

metrictypewhat it tells you
llm_router_provider_calls_totalcounterTraffic split per provider and model — is the LB actually balancing?
llm_router_provider_latency_secondshistogramUpstream time in buckets: p50 for capacity review, p95/p99 for the SLO.
llm_router_provider_error_totalcounterUpstream failures by provider — page on ratio, not on absolute count.
llm_router_lb_strategy_selected_totalcounterSelection counts per strategy and model — is the strategy you configured the one actually running?
llm_router_pipeline_stage_total{stage,result}counterguardrail_request pass|block, masking applied|skipped, request_received, provider_resolved.
llm_router_retry_total
llm_router_retry_exhausted_total
counterFailover pressure vs. hard capacity loss — the difference between flaky and full.
llm_router_model_fallback_totalcounterRequests served by a model’s fallback_model — who is quietly carrying whose traffic.
llm_router_tokens_total{direction}counterPrompt and completion tokens — cost attribution per service or key.
llm_router_response_format_total
llm_router_payload_conversion_total
counterSchema compatibility across openai / vllm / ollama dialects.

P95 Latency by provider

histogram_quantile(0.95, sum by (le, provider) ( rate(llm_router_provider_latency_seconds_bucket[5m) ))

ERR Error ratio

sum(rate(llm_router_provider_error_total[5m])) / sum(rate(llm_router_provider_calls_total[5m])) > 0.02 # page: 2% of traffic failing

GUARD Block rate

sum by (result) ( rate(llm_router_pipeline_stage_total {stage="guardrail_request"}[15m]) )

IO Tokens per minute

sum by (direction) ( rate(llm_router_tokens_total[1m) ) * 60

/metrics requires Redis for the router-level metrics store; the Prometheus exposition itself is served from the process.

07 / OPS SURFACE

One console script, no web console required

llm-router does everything the Config Manager UI does, so it fits in a Run block, an Ansible task or a Friday-night SSH session. The same script supervises the server itself — one default router or many named ones side by side (-i/--instance), each with its own pid file, port, log and config.env. Tab completion included.

deploy@router-01: /srv/llm-router
$ llm-router config discover localhost # auto-discover local providers
 
$ llm-router server start -i dev --port 8081 --save-config # many routers, one CLI
Saved 1 setting(s) to ~/.llm-router/instances/dev/config.env:
  LLM_ROUTER_SERVER_PORT
$ llm-router server start -i prod # own pid, port, log, config.env
$ llm-router server list
NAME STATUS PID PORT SERVER STARTED
dev • running 41207 8081 gunicorn 2026-09-16T19:58:29+0200
prod • running 41212 8080 gunicorn 2026-09-16T19:58:31+0200
batch • stopped - 8082 gunicorn 2026-09-16T18:40:02+0200
$ llm-router server reload -i prod # pick up models-config.json
$ llm-router server log -i dev # follow that instance's log
$ llm-router server stop --all # every instance at once
 
$ llm-router auth policy list
internal-only models: gemma-3-27b masking: required guardrail: required
partner-ex models: gpt-4o-mini masking: required guardrail: request
 
$ cat ./incidents/*.md | llm-router anonymizer run --algorithm fast_masker -o redacted.md
$ llm-router util translate --llm-router-url http://localhost:8080 ...
$ llm-router completion zsh > ~/.zsh/completions/_llm-router
  • auth — key generate | list | delete | disable | enable | rotate, policy, rate-limit
  • config — discover scans for Ollama, vLLM, LM Studio, llama.cpp, KoboldCpp and TabbyAPI and emits a config; merge unions two of them
  • server — start | stop | status | reload | log | list | rm-instance
  • -i, --instance — several routers on one host, state under ~/.llm-router/instances/NAME; status --all / stop --all reach every one of them
  • --save-config — store the flags passed on this command line, so the next start -i NAME needs none
  • config.env — per-instance overrides, precedence CLI > config.env > shell > defaults; a port already taken fails before the daemon starts
  • anonymizer — run the masking pipeline over local text
  • util — translate, genai-classifier, data augmentation via the router
  • completion — bash and zsh
08 / DEPLOY

Container, Helm chart or a bare port —
no control plane to babysit

Everything the router needs is one JSON file and a handful of env vars. The state it shares across replicas lives in Redis, so you can scale the deployment the same way you scale any other stateless Flask service.

docker
python:3.14-slim-trixie, runs as a non-root uid/gid 5000, listens on 8080. Config goes in via LLM_ROUTER_MODELS_CONFIG.
one container, no sidecars
$ docker run -d -p 5555:8080 \
    -v $PWD/models-config.json:/srv/models.json:ro \
    -e LLM_ROUTER_MODELS_CONFIG=/srv/models.json \
    quay.io/radlab/llm-router:v1.1.6

$ curl localhost:5555/api/version
{"version": "1.1.6"}
helm
Chart in-repo (appVersion 1.1.6), Deployment + Service + Ingress, ConfigMap built from files/models-config.json, readiness on /, Bitnami redis 23.2.12 subchart behind redis.enabled.
helm_charts/llm-router
$ helm upgrade --install my-llm-router ./helm_charts/llm-router \
     --namespace my-namespace --create-namespace \
     --set redis.enabled=true

# models ship from files/models-config.json;
# edit values.yaml and upgrade again —
# config lives in source control, not in a UI
bare metal
Prefork gunicorn (4 workers by default) under systemd. Optional extras keep the base install lean: [api], [metrics], [vault].
run-rest-api-gunicorn.sh
$ pip install "radlab-llm-router[api,metrics]"

export LLM_ROUTER_SERVER_TYPE=gunicorn
export LLM_ROUTER_SERVER_WORKERS_COUNT=4
export LLM_ROUTER_MODELS_CONFIG=/srv/llm-router/models.json
# host 0.0.0.0, port 8080
$ ./run-rest-api-gunicorn.sh

[probe] Probes, timeouts, failure budget

/health and /metrics are the only paths public by default (LLM_ROUTER_AUTH_PUBLIC_ENDPOINTS) — they never need an API key, even with LLM_ROUTER_AUTH_ENABLED=true. Probing /api/ping or /api/version? Add the exact prefixed path to that variable.

  • LLM_ROUTER_PROVIDER_MONITOR_PING_TIMEOUT_SECONDS (5.0s) — per-provider ping budget
  • 2 consecutive failures to eject a provider (LLM_ROUTER_PROVIDER_MONITOR_MAX_CONSECUTIVE_FAILURES) — one slow reply never drops a saturated GPU replica, and a single success brings it back
  • LLM_ROUTER_EXTERNAL_TIMEOUT (300s) for upstream calls; LLM_ROUTER_TIMEOUT (0 = off) bounds the router’s own API
  • Dead providers are removed from rotation, then re-admitted by the monitor on their own

[scale] Workers, replicas and Redis

Redis holds API keys, sliding-window rate limits and exclusive provider leases, so every replica sees the same fleet state. No sticky sessions on your LB.

  • PROMETHEUS_MULTIPROC_DIR is set for you — mount an emptyDir/tmpfs so counters survive worker restarts
  • Chart's redis.enabled=false and bring your own instance or cluster
  • Auth, keys and policy definitions are shared state — rotate a key once, every worker agrees
  • /metrics is unauthenticated by design: scrape it, and keep it inside your network policy
09 / ECOSYSTEM

Five repos, one wire format

The gateway is split so you only install what you actually run. Core library, REST surface and SDK ship in this repo; the UIs, plugins and guardrail services are separate deployables that talk to the same endpoints.

← Clients that just work

It speaks the OpenAI wire format, so existing client code changes base URL and key — nothing else. Python adds LLMRouterClient and, for asyncio apps, AsyncLLMRouterClient (httpx) with typed stream events. Runnable scripts live in examples/.

OpenAI SDKLLMRouterClientAsyncLLMRouterClient LangChainLlamaIndex LiteLLMHaystackembeddingsstreaming

→ Backends behind it

One api_type field per provider. Self-hosted and commercial endpoints can sit in the same model group and share a balancing strategy.

vLLMOllamallama.cpp LM StudioOpenAI-compatible Anthropiclocal GGUF
10 / COMPARISON

Local-first, by design

Same job — route, balance and protect LLM traffic — different centres of gravity. LLMRouter compared against the gateways local-AI stacks actually run: LiteLLM, Bifrost and Portkey. Table stakes are left out on purpose: streaming, tool calls and OpenAI-compatible endpoints are something every gateway in this table already has.

strong partial absent / weak ★ strongest in row — more stars, bigger edge

Ratings read from public docs, READMEs and repository trees, checked September 2026: LiteLLM ~59k stars and 136 provider directories, Portkey ~13k and 75, Bifrost ~8k and 23+. Rows an alternative wins stay in the table — the point is where each tool is strongest, not a clean sweep.

capabilityllmrouterlitellmbifrostportkey
local-first · zero external services ★★★
ollama / vllm / lm studio backends ★★★
claude code over local models ★★★
codex / responses api ★★★
model-cache-aware load balancing ★★★
low-overhead load balancing ★★
semantic routing (biencoder + faiss) ★★
agentic work-mode routing ★★★
rag context injection in the gateway ★
built-in task endpoints (translate / polarity / mask) ★★
pii masking in the request path ★★
polish pii masking (pesel / nip / regon) ★★★
ml pii ner model ★
guardrails on your own models (nask / sójka) ★★★
token budgets and rate limits ★★ ★★★
spend tracking / multi-tenant billing ★★★ ★★
admin ui / gui config ★ ★★
named-instance server supervision (cli) ★★★
audit trail + prometheus metrics ★★
audit log encryption (gpg at rest) ★★
mcp gateway / tools server ★★
broad hosted-provider ecosystem ★★★
enterprise governance and maturity ★★★ ★ ★★
10b / PLUGIN SURFACE

What the plugin layer adds

A lot of the interesting surface lives outside the core. llm-router-plugins is a separate Apache-2.0 package — pip install "radlab-llm-router-plugins" — where every masking rule, guardrail and routing strategy hangs off one apply() hook in the pipeline. The alternatives spread the same ground over OSS plugins, built-in features and paid editions, so it is worth comparing separately.

plugin surfacellm-router + pluginslitellmbifrostportkey
plugin model (pip install, apply hook) ★★
masking rules shipped in the oss package ★★★
id checksums + formats (pesel / nip / regon / krs) ★★★
ml pii ner on your own model ★★
guardrails on your own services ★★★
guardrail vendor integrations ★★★ ★★
routing plugins (keyword / biencoder / agentic) ★★★
rag retrieval plugin (langchain + faiss) ★★★
semantic caching ★★ ★
virtual keys / secret management ★★★
all of the above free to self-host ★★★

Self-hosting is the row most teams feel: Bifrost unlocks adaptive load balancing, guardrails and its MCP gateway in a commercial edition, Portkey semantic caching is a hosted-platform feature, and LiteLLM ships an MIT core with a separate enterprise edition. Everything in the first column is Apache-2.0 and runs on your own hardware — the guardrail plugins are plain HTTP clients, so NASK and Sójka are services you deploy and point them at. Secret managers on LiteLLM and Bifrost, and Bifrost's audit logs, require that Enterprise license as well.

→ Pick LLMRouter when

The stack lives on your own hardware and the data should stay there.

  • Polish PII — PESEL, NIP and REGON checksums plus KRS, NRB and Polish IBAN formats, with a PL NER model you run yourself for free text
  • Guardrails on your own hardware — NASK HerBERT-PL-Guard and SpeakLeash Sójka behind services you host, no third-party API in the request path
  • Agentic routing — Codex work modes (plan / implement / test / review / debug), request classes ranked apart from main turns, and Claude Code tier swaps
  • Warm-cache routing — keep talking to the host that already has the weights loaded instead of paying for a 30–100 GB reload elsewhere

← Pick the alternatives when

Honest version — the others are great at different jobs.

  • LiteLLM — 136 provider directories, ~59k stars, dollar spend tracking, budgets and an MCP gateway
  • Bifrost — Go data path with the lowest overhead of the three, plus zero-config npx @maximhq/bifrost and a built-in web UI
  • Portkey — hosted observability, 40+ guardrail validators and multi-tenant control

Point your client at localhost:8080

Apache-2.0, no accounts, no phone-home. Install it, write one JSON file, and your prompts stop taking a detour through someone else's cloud.

$pip install "radlab-llm-router[api]"copied