Getting started
Install the gateway and get a first completion through the router.
operator documentation
Self-hosted AI gateway -- operator documentation
v1.1.6
llm-router-plugins v0.1.1 @ b7b68ce llm-router-services v0.0.4 @ d1b9641
Install the gateway and get a first completion through the router.
Environment variables and the models.json routing configuration.
Load balancing strategies, failover and connection keep-alive.
Authentication, authorization, rate limiting and the GPG audit trail.
Prometheus metrics exposed by the router and how to scrape them.
Endpoint catalogue and the guide for adding your own endpoints.
llm-router command line reference: server lifecycle, models, utilities.
Embed the router in your own code with llm_router_lib.
Helm chart, container image and the runtime knobs that matter in production.
Runnable stacks: vLLM, Ollama, LangChain, LlamaIndex and client SDKs.
What changed in every released version.
The llm-router-plugins package: anonymizers, guardrails, semantic routing and RAG plugins.
Built-in and custom PII anonymizer plugins.
Semantic routing, bi-encoder routing and Codex CLI routing.
RAG support and other pipeline utilities.
The llm-router-services HTTP API: guardrails, maskers and the unified Flask application.
NASK-PIB, Sojka and dynamic guardrail routing endpoints.
Building, running and deploying the services container.