llm-router/docs

LLM Router - Open-Source AI Gateway for Local and Cloud LLM Infrastructure#

LLM Router is a service that can be deployed on‑premises or in the cloud. It adds a layer between any application and the LLM provider. In real time it controls traffic, distributes load among providers of a specific LLM, and enables analysis of outgoing requests from a security perspective (masking, anonymization, prohibited content). It is an open‑source solution (Apache 2.0) that can be launched instantly by running a ready‑made image in your own infrastructure.


🌐 Ecosystem Overview#

The LLM‑Router project is split across five dedicated repositories:

Repository Description
llm-router (this repo) Core gateway — unified REST proxy, Python SDK, and configuration management
llm-router-api (subdirectory) REST proxy that routes requests to any supported LLM backend (OpenAI‑compatible, Ollama, vLLM, LM Studio, Anthropic), with built‑in load‑balancing, health checks, streaming responses and optional Prometheus metrics
llm-router-lib (subdirectory) Python SDK that wraps the API with typed request/response models, automatic retries, token handling and a rich exception hierarchy
llm-router-web Ready‑to‑use Flask UIs — a Config Manager for model/user settings and an Anonymizer UI that masks sensitive data
llm-router-plugins Pluggable anonymizers (maskers), guardrails, semantic routing and RAG plugins
llm-router-services HTTP services that power the plugin ecosystem (NASK‑PIB/Sojka guardrails, PII masker)
llm-router-utils CLI tools, batch translation, GenAI classification and ready‑made deployment configs (Speakleash models)

✨ Key Features#

Feature Description
Unified REST interface One endpoint schema works for OpenAI‑compatible, Ollama, vLLM, LM Studio and Anthropic.
Provider‑agnostic streaming The stream flag (default true) controls whether the proxy forwards chunked responses as they arrive or returns a single aggregated payload. Streaming responses include proper Cache‑Control, Pragma, Expires and Vary headers.
Built‑in prompt library Language‑aware system prompts stored under resources/prompts can be referenced automatically.
Dynamic model configuration JSON file (models-config.json) defines providers, model name, default options and per‑model overrides.
Request validation Pydantic models guarantee correct payloads; errors are returned with clear messages.
Structured logging Configurable log level, filename, and optional JSON formatting.
Health & metadata endpoints /ping (simple 200 OK) and /tags (available model tags/metadata).
Embeddings support Dedicated endpoints for generating text embeddings across all supported providers.
Simple deployment One‑liner run script, Docker image, or Helm chart for Kubernetes.
Extensible conversation formats Basic chat, conversation with system prompt, and extended conversation with richer options (temperature, top‑k, custom system prompt).
Multi‑provider model support Each model can be backed by multiple providers (VLLM, Ollama, OpenAI, Anthropic) defined in models-config.json.
Load‑balanced default strategy LoadBalancedStrategy distributes requests evenly across providers using in‑memory usage counters.
Dynamic model handling ModelHandler loads model definitions at runtime and resolves the appropriate provider per request.
Pluggable endpoint architecture Automatic discovery and registration of all concrete EndpointI implementations via EndpointAutoLoader.
Prometheus metrics integration Optional /metrics endpoint for latency, error counts, and provider usage statistics.
Docker & Kubernetes ready Dockerfile (non‑root user) and Helm charts for containerised deployment.

🧩 Plugin System Architecture#

LLM Router uses a registry-based pipeline pattern. Each plugin implements a tiny, well‑defined apply method and can be composed in an ordered list to form a pipeline. Pipelines are instantiated by the MaskerPipeline, GuardrailPipeline and UtilsPipeline classes and are driven automatically by the endpoint logic in endpoint_i.py.

Data flow#

Request → MaskerPipeline → GuardrailPipeline → UtilsPipeline → Model Provider

Masker Plugins#

Plugin ID Type Description
fast_masker Local Regex‑based PII masker with 30+ rule types (emails, IPs, URLs, phone numbers, PESEL, NIP, KRS, REGON, monetary amounts, dates, credit cards, JWTs, passports and more).
pii_masker HTTP (remote) ML‑based PII masker using a token‑classification model with an in‑memory cache to avoid redundant model calls for identical text inputs.

Guardrail Plugins#

Plugin ID Type Description
nask_guard HTTP (remote) Safety check using the HerBERT‑PL‑Guard model (NASK‑PIB).
sojka_guard HTTP (remote) Safety check using the Bielik‑Guard‑0.1B model from SpeakLeash.

Utility Plugins#

Plugin ID Type Description
langchain_rag Local Retrieves relevant document chunks from a FAISS vector store and injects them into the payload for Retrieval‑Augmented Generation.
simple_semantic_routing Local Two‑stage heuristic model selection: intent classification + complexity analysis. Activated when payload["model"] == "auto".
semantic_biencoder_routing Local Embedding‑based semantic routing using FAISS — matches user messages against pre‑configured target embeddings to select the best model. See Semantic BiEncoder Routing for configuration details.

Pipelines are configured via environment variables:

# Comma-separated list of masker plugins to apply
export LLM_ROUTER_MASKING_STRATEGY_PIPELINE="fast_masker,pii_masker"

# Enable guardrails
export LLM_ROUTER_FORCE_GUARDRAIL_REQUEST=1

# Enable masking entirely
export LLM_ROUTER_FORCE_MASKING=1

# Record masking operations in audit log
export LLM_ROUTER_MASKING_WITH_AUDIT=1

Semantic BiEncoder Routing#

The semantic_biencoder_routing plugin uses a neural embedding model (google/embeddinggemma-300m) to compute semantic embeddings for a set of pre‑configured routing targets. Each target has a name, a model_name (the model to route to), a description, and a list of examples. At query time the user message is embedded and matched against all stored target embeddings using FAISS (IndexFlatIP on L2‑normalised vectors = cosine similarity). The best‑matching target determines the selected model.

How it works#

1. Index building (on first load or when the persist directory is missing):

  • For each routing target, its description and examples are combined into text.
  • The text is split into overlapping token chunks using a sliding window (chunk_size tokens, chunk_overlap tokens overlap).
  • Each chunk is embedded via the BiEncoder model (e.g. google/embeddinggemma-300m).
  • All embedding vectors are L2‑normalised to unit length.
  • Vectors are inserted into a faiss.IndexFlatIP index (inner product).
  • A docstore maps each FAISS document ID to its target name (for reverse lookup).

2. Routing (query):

  • The user message is embedded and L2‑normalised.
  • FAISS performs a nearest‑neighbor search returning the top_k closest chunks.
  • Scores are aggregated per target: the mean cosine similarity of all chunks belonging to the same target is computed.
  • The target with the highest mean similarity wins and its model_name is returned.

3. Persistence:

The FAISS index and docstore are saved to disk (files index.faiss and docstore.pkl) under the configured persist directory. On subsequent starts the index is loaded from disk — embeddings are not recomputed. If the embedding model changes (different output dimension) the index is automatically rebuilt.

Full example JSON config with routing_targets and their examples is available in the plugins repo: llm_router_plugins/resources/routing/semantic_biencoder.json. Detailed variable descriptions and usage examples: Plugin Routing README.

Example routing targets:

Target name Model routed to Description
code-generation qwen3.6:35b Code‑related tasks: writing, debugging, refactoring.
math-analysis qwen3.6:35b Mathematical computations, statistical analysis, quantitative reasoning.
creative-writing gpt-oss:120b Creative and generative writing: stories, poems, marketing content.
general-assistant gpt-oss:120b Everyday questions, explanations, research, conversation.
data-science qwen3.6:35b Data analysis, visualization, ML pipelines, reporting.
system-admin gpt-oss:120b System administration, DevOps, infrastructure, technical ops.

📦 Quick Start#

1️⃣ Create & activate a virtual environment#

python3 -m venv .venv
source .venv/bin/activate

# Only the core library (llm-router-lib).
pip install .

# Core library + API wrapper (llm-router-api).
pip install .[api]

# Core library + API wrapper + Prometheus metrics.
pip install .[api,metrics]

Note: When Prometheus metrics are enabled, LLM_ROUTER_USE_PROMETHEUS=1 must be set and Redis is required (used for provider availability state).
The multiproc directory defaults to $HOME/.llm-router/metrics/prometheus/multiproc — override via the PROMETHEUS_MULTIPROC_DIR environment variable if needed.

Then start the application with the environment variable set:

export LLM_ROUTER_USE_PROMETHEUS=1

When LLM_ROUTER_USE_PROMETHEUS is enabled, the router automatically registers a /metrics endpoint (under the API prefix, e.g. /api/metrics). This endpoint exposes Prometheus‑compatible metrics such as request counts, latencies, and any custom counters defined by the application.

2️⃣ Run the REST API#

./run-rest-api.sh
# or
LLM_ROUTER_MINIMUM=1 python3 -m llm_router_api.rest_api

3️⃣ Quick‑start guides for local models#

  • Gemma 3 12B‑IT – README
  • Bielik 11B‑v2.3‑Instruct – README

4️⃣ Integration boilerplates#

Integration examples for popular LLM libraries (LlamaIndex, LangChain, OpenAI, LiteLLM, Haystack) are in the examples/ directory. See examples README for details.


🔐 Auditing#

The router can record request‑level events (guard‑rail checks, payload masking, custom logs) in a tamper‑evident, encrypted form. All audit entries are written by the auditor module and stored under logs/auditor/ as GPG‑encrypted files.

For a complete guide — including key generation, encryption workflow, and decryption utilities — see:

➡️ Auditing subsystem documentation

Utility scripts:

  • scripts/gen_and_export_gpg.sh — generate and export GPG keys
  • scripts/decrypt_auditor_logs.sh — decrypt encrypted audit logs

🔐 Authentication#

The router supports API-key-based authentication with per-endpoint policies, rate limiting, and audit trail.

➡️ Authentication documentation


⏱️ Rate Limiting#

Sliding-window rate limiting backed by Redis sorted sets. Each API key + IP gets a configurable number of requests per minute. Returns Retry-After on 429 responses and exposes Prometheus metrics.

➡️ Rate Limiting documentation


🖥️ CLI Reference#

The llm-router package provides a command-line tool for managing API keys, policies, rate-limit presets, and anonymizing text. A full command reference is available here:

➡️ CLI Command Reference


🔒 Security#

🔍 Error message sanitization#

All error messages returned to API callers are sanitized to prevent leakage of internal infrastructure details (IP addresses, hostnames, URLs, ports, connection strings).

How it works:

  • sanitize_error_message() in llm_router_api/core/errors.py strips URLs, IP addresses, ports, hostnames, and urllib3/requests exception internals from error strings.
  • Applied at every output choke point:
    • HTTP provider errors (httprequest.py)
    • Streaming error chunks (stream_handler.py)
    • return_response_not_ok() — the central error builder for all non-streaming errors
    • Parameter validation errors in register.py
  • Server-side logs still receive the full, unsanitized exception — debugging remains fully possible.

What you will see as a caller:

  • ✅ "ConnectTimeout: The read operation timed out"
  • ✅ "A connection error occurred"

What you will NOT see:

  • ❌ 192.168.x.x, 10.0.x.x — internal IPs
  • ❌ http://..., https://... — internal URLs
  • ❌ port=8080, host='...' — connection details
  • ❌ Stack traces or internal provider addresses

This protection applies to all error responses regardless of whether they originate from HTTP provider calls, streaming endpoints, or request validation.


📦 Docker#

Run the container with the default configuration:

docker run -p 5555:8080 quay.io/radlab/llm-router:rc1

For more advanced usage you can use a custom launch script:

#!/bin/bash

PWD=$(pwd)

docker run \
  -p 5555:8080 \
  -e LLM_ROUTER_TIMEOUT=500 \
  -e LLM_ROUTER_IN_DEBUG=1 \
  -e LLM_ROUTER_MINIMUM=1 \
  -e LLM_ROUTER_EP_PREFIX="/api" \
  -e LLM_ROUTER_SERVER_TYPE=gunicorn \
  -e LLM_ROUTER_SERVER_PORT=8080 \
  -e LLM_ROUTER_SERVER_WORKERS_COUNT=4 \
  -e LLM_ROUTER_DEFAULT_EP_LANGUAGE="pl" \
  -e LLM_ROUTER_LOG_FILENAME="llm-proxy-rest.log" \
  -e LLM_ROUTER_EXTERNAL_TIMEOUT=300 \
  -e LLM_ROUTER_BALANCE_STRATEGY=balanced \
  -e LLM_ROUTER_REDIS_HOST="192.168.100.67" \
  -e LLM_ROUTER_REDIS_PORT=6379 \
  -e LLM_ROUTER_MODELS_CONFIG=/srv/cfg.json \
  -e LLM_ROUTER_PROMPTS_DIR="/srv/prompts" \
  -v "${PWD}/resources/configs/models-config.json":/srv/cfg.json \
  -v "${PWD}/resources/prompts":/srv/prompts \
  quay.io/radlab/llm-router:rc1

Kubernetes (Helm)#

Helm charts for Kubernetes deployment are available in the helm_charts/ directory.


🛠️ Configuration (via environment)#

A full list of environment variables is available at: API README

Core variables#

Variable Default Description
LLM_ROUTER_PROMPTS_DIR resources/prompts Directory containing predefined system prompts.
LLM_ROUTER_MODELS_CONFIG resources/configs/models-config.json Path to the models configuration JSON file.
LLM_ROUTER_DEFAULT_EP_LANGUAGE pl Default language for endpoint prompts.
LLM_ROUTER_TIMEOUT 0 Timeout (seconds) for llm-router API calls.
LLM_ROUTER_EXTERNAL_TIMEOUT 300 Timeout (seconds) for external model API calls.
LLM_ROUTER_MAX_REQUEST_BODY_SIZE 10485760 (10 MB) Maximum allowed request body size in bytes. Larger payloads are rejected with HTTP 413 to prevent memory exhaustion.
LLM_ROUTER_LOG_FILENAME llm-router.log Name of the log file.
LLM_ROUTER_LOG_LEVEL INFO Logging level (e.g., INFO, DEBUG).
LLM_ROUTER_EP_PREFIX /api Prefix for all API endpoints.
LLM_ROUTER_MINIMUM 1 Run service in proxy‑only mode.
LLM_ROUTER_IN_DEBUG 1 Run server in debug mode.
LLM_ROUTER_BALANCE_STRATEGY first_available Load‑balancing strategy: balanced, weighted, dynamic_weighted, first_available, first_available_optim.
LLM_ROUTER_SERVER_TYPE flask Server implementation: flask, gunicorn, waitress.
LLM_ROUTER_SERVER_PORT 8080 Port on which the server listens.
LLM_ROUTER_SERVER_HOST 0.0.0.0 Host address for the server.
LLM_ROUTER_SERVER_WORKERS_COUNT 4 Number of workers.
LLM_ROUTER_SERVER_THREADS_COUNT 16 Number of worker threads.
LLM_ROUTER_SERVER_WORKER_CLASS None Worker class for servers that support it.
LLM_ROUTER_USE_PROMETHEUS 1 Enable Prometheus metrics (/metrics endpoint).
PROMETHEUS_MULTIPROC_DIR $HOME/.llm-router/metrics/prometheus/multiproc Directory where prometheus multiprocess worker data files are stored. Overrides are allowed but the default works in most deployments.

Masking & guardrail variables#

Variable Default Description
LLM_ROUTER_FORCE_MASKING False Enable force‑masking of every endpoint's payload.
LLM_ROUTER_MASKING_STRATEGY_PIPELINE ["fast_masker"] Ordered list of masker plugins (e.g. fast_masker,pii_masker).
LLM_ROUTER_MASKING_WITH_AUDIT False Record each masking operation in the audit log.
LLM_ROUTER_FORCE_GUARDRAIL_REQUEST False Force guardrail evaluation on every request.
LLM_ROUTER_MASKER_PII_HOST — Host URL for the PII masker service.
LLM_ROUTER_GUARDRAIL_SOJKA_GUARD_HOST — Host URL for the Sojka guardrail service.

Redis variables#

Variable Default Description
LLM_ROUTER_REDIS_HOST (empty) Redis host for load‑balancing across multi‑provider models.
LLM_ROUTER_REDIS_PORT 6379 Redis port.
LLM_ROUTER_REDIS_PASSWORD (not set) Redis password.
LLM_ROUTER_REDIS_DB 0 Redis database number.

Note: When LLM_ROUTER_REDIS_HOST is set, the router uses Redis for load‑balancing state and provider availability tracking.

Semantic BiEncoder Routing variables#

Variable Default Description
LLM_ROUTER_ROUTING_SEMANTIC_BIENCODER_CONFIG (empty) Config source of truth — either a raw JSON string (starts with {) or a file path. When unset, falls back to the bundled semantic_biencoder.json. Individual env vars below override values from the loaded config only when explicitly set.
LLM_ROUTER_ROUTING_SEMANTIC_BIENCODER_MODEL (empty) Override the embedding model name or local path.
LLM_ROUTER_ROUTING_SEMANTIC_BIENCODER_TARGETS (empty) Pipe‑separated list of target names (overrides all targets).
LLM_ROUTER_ROUTING_SEMANTIC_BIENCODER_CHUNK_SIZE (empty) Token chunk size for embedding.
LLM_ROUTER_ROUTING_SEMANTIC_BIENCODER_CHUNK_OVERLAP (empty) Token overlap between consecutive chunks.
LLM_ROUTER_ROUTING_SEMANTIC_BIENCODER_PERSIST_DIR (empty) Directory for FAISS index + docstore persistence (index.faiss, docstore.pkl).

LangChainRAG variables#

Variable Default Description
LLM_ROUTER_LANGCHAIN_RAG_COLLECTION (empty) Vector store collection name.
LLM_ROUTER_LANGCHAIN_RAG_EMBEDDER (empty) Path to the embedding model (e.g. /mnt/data2/llms/models/community/google/embeddinggemma-300m).
LLM_ROUTER_LANGCHAIN_RAG_DEVICE cpu Compute device for the embedding model (cpu, cuda:0, …).
LLM_ROUTER_LANGCHAIN_RAG_CHUNK_SIZE 1024 Chunk size for document splitting.
LLM_ROUTER_LANGCHAIN_RAG_CHUNK_OVERLAP 100 Overlap between consecutive chunks.
LLM_ROUTER_LANGCHAIN_RAG_PERSIST_DIR (empty) Directory for LangChainRAG index persistence.

Note: LangChainRAG requires the llm-router-plugins package. When LLM_ROUTER_UTILS_PLUGINS_PIPELINE includes langchain_rag, these variables configure the RAG plugin behavior.

Plugin Pipeline variables#

Variable Default Description
LLM_ROUTER_UTILS_PLUGINS_PIPELINE semantic_biencoder_routing Comma‑separated list of utility plugins to apply (e.g. simple_semantic_routing,langchain_rag).

Note: Utility plugins run per‑request in the pipeline between endpoint processing and model provider dispatch. Available plugins: simple_semantic_routing, semantic_biencoder_routing, langchain_rag. Each plugin has additional configuration documented above.

Authentication variables#

Variable Default Description
LLM_ROUTER_AUTH_ENABLED false Master switch — "true" enables all authentication. Default is false (no auth).
LLM_ROUTER_AUTH_KEY_STORE memory Key store backend: vault, redis, or memory.
LLM_ROUTER_AUTH_VAULT_ADDR (empty) HashiCorp Vault server URL (e.g. https://vault.example.com).
LLM_ROUTER_AUTH_VAULT_PATH secret/data/llm-router/api-keys KV v2 mount path for key storage.
LLM_ROUTER_AUTH_VAULT_AUTH_METHOD kubernetes Auth method: kubernetes, approle, or token.
LLM_ROUTER_AUTH_VAULT_ROLE_ID (empty) AppRole role ID (or K8s SA token for K8s auth).
LLM_ROUTER_AUTH_VAULT_SECRET_ID (empty) AppRole secret ID.
LLM_ROUTER_AUTH_VAULT_TOKEN (empty) Vault token (for token auth).
LLM_ROUTER_AUTH_KEY_CACHE_TTL 300 Key cache TTL in seconds.
LLM_ROUTER_AUTH_KEY_CACHE_JITTER 60 Random jitter to prevent cache stampede.
Auth Redis (separate from general REDIS)
LLM_ROUTER_AUTH_REDIS_HOST (empty) Auth Redis host for key store and rate limiting.
LLM_ROUTER_AUTH_REDIS_PORT 6379 Auth Redis port.
LLM_ROUTER_AUTH_REDIS_DB 0 Auth Redis database number.
LLM_ROUTER_AUTH_REDIS_PASSWORD (not set) Auth Redis password.
LLM_ROUTER_AUTH_DEFAULT_RATE_LIMIT 60 Default rate limit (requests per minute).
LLM_ROUTER_AUTH_PUBLIC_ENDPOINTS /ping,/version,/models,/ Comma-separated paths that bypass authentication.
LLM_ROUTER_AUTH_KEY_PREFIX sk-litm Key prefix (like LiteLLM/OpenAI format).
LLM_ROUTER_AUTH_KEY_LENGTH 48 Entropy bytes for key generation (produces 64-char key).
LLM_ROUTER_AUTH_ROTATION_GRACE_PERIOD 3600 Old keys remain valid for this many seconds after rotation.
LLM_ROUTER_AUTH_AUDIT (empty) Record auth events in the audit log.

Note: Rate limiting is always applied when authentication is enabled — there is no separate toggle for it. Auth Redis (LLM_ROUTER_AUTH_REDIS_*) is independent from general Redis (LLM_ROUTER_REDIS_*).

See full authentication docs: llm_router_api/AUTHENTICATION.md


⚖️ Load Balancing Strategies#

The current list of available strategies, the interface description, and an example extension can be found at: Load‑Balancing Strategies

Strategies: balanced, weighted, dynamic_weighted, first_available, first_available_optim.


🛣️ Endpoints Overview#

The list of endpoints — categorized into built‑in, provider‑dependent, and utility endpoints — and a description of the streaming mechanisms can be found at: Endpoints Overview

Highlights#

Endpoint Method Auth (when LLM_ROUTER_AUTH_ENABLED=true) Description
/ping GET ✅ Public Health‑check
/version GET ✅ Public Return router version
/ GET ✅ Public Ollama health endpoint
/models GET ✅ Public List OpenAI‑compatible models
/v1/models GET ❌ Requires chat permission List OpenAI‑compatible models (v1)
/tags GET ✅ Public List Ollama model tags
/api/v0/models GET ❌ Requires chat permission List LM Studio models
/metrics GET ✅ Public Prometheus metrics (requires Redis)
/chat/completions POST ❌ Requires chat permission OpenAI‑style chat completion
/api/chat/completions POST ❌ Requires chat permission OpenAI‑style chat completion (with prefix)
/v1/chat/completions POST ❌ Requires chat permission vLLM‑like chat completion
/v1/messages POST ❌ Requires anthropic permission Anthropic‑compatible messages endpoint (Claude)
/responses POST ❌ Requires chat permission OpenAI‑like responses endpoint
/v1/responses POST ❌ Requires chat permission OpenAI‑like responses endpoint (v1)
/embeddings POST ❌ Requires embedding permission Standard embeddings
/api/embeddings POST ❌ Requires embedding permission Standard embeddings (with prefix)
/v1/embeddings POST ❌ Requires embedding permission OpenAI‑compatible embeddings endpoint
/api/embed POST ❌ Requires embedding permission Ollama‑native embeddings endpoint
/api/chat POST ❌ Requires ollama permission Ollama‑style chat completion
/api/conversation_with_model POST ❌ Requires builtin permission Built‑in standard chat
/api/extended_conversation_with_model POST ❌ Requires builtin permission Built‑in chat with extended fields
/api/generative_answer POST ❌ Requires builtin permission Answer a question using provided context
/api/translate POST ❌ Requires builtin permission Translate texts
/api/generate_questions POST ❌ Requires builtin permission Generate questions from texts
/api/simplify_text POST ❌ Requires builtin permission Simplify input texts

Note: By default LLM_ROUTER_AUTH_ENABLED=false, so all endpoints are accessible without authentication. Set it to "true" to enforce auth. The _public list (default /ping,/version,/models,/) can be customized via LLM_ROUTER_AUTH_PUBLIC_ENDPOINTS.


🌐 Web Applications#

Config Manager (port 8081)#

Full web UI for managing LLM Router model configurations:

  • Multi‑user with authentication and role‑based access (admin/user)
  • Projects — group configurations by project
  • Model configuration — create, edit, import/export JSON configs; manage providers across families (Google, OpenAI, Qwen)
  • Version control — snapshot history with restore capability
  • Active model selection — choose which models to activate per config
  • Drag‑and‑drop provider reordering (HTMX)
  • Light/dark themes (Alpine.js)
  • 26+ API endpoints under /configs

Run: ./run-configs-manager.sh

Anonymizer (port 8082)#

Web UI for text anonymization and interactive chat:

  • 3 anonymization algorithms: fast (regex), pii_masking (ML model), fast+pii (hybrid)
  • Interactive chat with streaming SSE responses and session persistence
  • Dynamic model selection from the router
  • i18n — Polish and English translations (122 keys)
  • Privacy warnings when anonymization is disabled
  • Privacy policy & terms pages

Run: ./run-anonymizer.sh


🧰 llm-router-utils#

The llm-router-utils repository provides CLI tools and ready‑made deployment configs:

CLI Tools#

Tool Description
translate-texts Batch translate texts in JSON/JSONL datasets via LLM Router
genai-classifier Classify dataset texts using LLM prompts with multi‑threading and XLSX export

Speakleash Deployment Configs#

The resources/llm-router-speakleash/ directory contains ready‑made configs for deploying Speakleash models:

  • speakleash-models.json — configures Bielik-11B-v2.3-Instruct across 8 vLLM providers on 3 hosts
  • run-bielik-*.sh — vLLM launch scripts for each GPU (cuda:0, cuda:1, cuda:2)
  • run-rest-api-gunicorn.sh — full LLM Router server with masking, guardrails, Redis balancing, and Prometheus metrics
  • run-sojka-guardrail.sh — guardrail service with Bielik‑Guard model

⚙️ Configuration Details#

Config File / Variable Meaning
resources/configs/models-config.json JSON map of provider → model → default options (e.g., keep_alive, options.num_ctx).
LLM_ROUTER_PROMPTS_DIR Directory containing prompt templates (*.prompt). Sub‑folders are language‑specific (en/, pl/).
LLM_ROUTER_DEFAULT_EP_LANGUAGE Language code used when a prompt does not explicitly specify one.
LLM_ROUTER_TIMEOUT Upper bound for any request to an upstream LLM (seconds).
LLM_ROUTER_LOG_FILENAME / LLM_ROUTER_LOG_LEVEL Logging destinations and verbosity.
LLM_ROUTER_IN_DEBUG When set, enables DEBUG‑level logs and more verbose error payloads.

🔧 Development#

  • Python 3.10+ (project is tested on 3.10.6)
  • All dependencies are listed in requirements.txt. Install them inside the virtualenv.
  • To add a new provider, create a class in llm_router_api/core/api_types that implements the BaseProvider interface and register it in llm_router_api/register/__init__.py.

📚 Changelog#

See the CHANGELOG for a complete history of changes.

📜 License#

See the LICENSE file.

llm-router · docs are generated from the repository by tools/build_docs.py 0.6.5 @ d9c38ae