{"version":"1.1.5","pages":[{"k":"overview.html","t":"Overview","s":"Getting started","x":"# LLM Router - Open-Source AI Gateway for Local and Cloud LLM Infrastructure [**LLM Router**](https://llm-router.cloud) is a service that can be deployed on‑premises or in the cloud. It adds a layer between any application and the LLM prov…","h":["🌐 Ecosystem Overview","✨ Key Features","🧩 Plugin System Architecture","Data flow","Masker Plugins","Guardrail Plugins","Utility Plugins","Semantic BiEncoder Routing","📦 Quick Start","1️⃣ Create &amp; activate a virtual environment","2️⃣ Install (recommended: from PyPI)","3️⃣ Run the REST API","4️⃣ Quick‑start guides for local models","5️⃣ Integration boilerplates","🔐 Auditing","🔐 Authentication","⏱️ Rate Limiting","🖥️ CLI Reference","🔒 Security","🔍 Error message sanitization","📦 Docker","Kubernetes (Helm)","Configuration","⚖️ Load Balancing Strategies","🛣️ Endpoints Overview","Highlights","🌐 Web Applications","Config Manager (port 8081)","Anonymizer (port 8082)","🧰 llm-router-utils","CLI Tools","Speakleash Deployment Configs","⚙️ Configuration Details","🔧 Development","📚 Changelog","📜 License"],"b":"LLM Router - Open-Source AI Gateway for Local and Cloud LLM Infrastructure # LLM Router is a service that can be deployed on‑premises or in the cloud. It adds a layer between any application and the LLM provider. In real time it controls traffic, distributes load among providers of a specific LLM, and enables analysis of outgoing requests from a security perspective (masking, anonymization, prohibited content). It is an open‑source solution (Apache 2.0) that can be launched instantly by running a ready‑made image in your own infrastructure. 🌐 Ecosystem Overview # The LLM‑Router project is split across five dedicated repositories: Repository Description llm-router (this repo) Core gateway — unified REST proxy, Python SDK, and configuration management llm-router-api (subdirectory) REST proxy that routes requests to any supported LLM backend (OpenAI‑compatible, Ollama, vLLM, LM Studio, Anthropic), with built‑in load‑balancing, health checks, streaming responses and optional Prometheus metrics llm-router-lib (subdirectory) Python SDK that wraps the API with typed request/response models, automatic retries, token handling, a rich exception hierarchy, and sync ( LLMRouterClient ) + async ( AsyncLLMRouterClient , streaming) clients llm-router-web Ready‑to‑use Flask UIs — a Config Manager for model/user settings and an Anonymizer UI that masks sensitive data llm-router-plugins Pluggable anonymizers (maskers), guardrails, semantic routing and RAG plugins llm-router-services HTTP services that power the plugin ecosystem (NASK‑PIB/Sojka guardrails, PII masker) llm-router-utils CLI tools, batch translation, GenAI classification and ready‑made deployment configs (Speakleash models) ✨ Key Features # Feature Description Unified REST interface One endpoint schema works for OpenAI‑compatible, Ollama, vLLM, LM Studio and Anthropic. Provider‑agnostic streaming The stream flag (default true ) controls whether the proxy forwards chunked responses as they arrive or returns a single aggregated payload. Streaming responses include proper Cache‑Control, Pragma, Expires and Vary headers. Built‑in prompt library Language‑aware system prompts stored under resources/prompts can be referenced automatically. Dynamic model configuration JSON file ( models-config.json ) defines providers, model name, default options and per‑model overrides. Request validation Pydantic models guarantee correct payloads; errors are returned with clear messages. Structured logging Configurable log level, filename, and optional JSON formatting. Health &amp; metadata endpoints /health and /api/ping (liveness), /api/version (build), /models , /v1/models and /api/tags (metadata), / for the Ollama probe. Built‑in article generation Two builtin endpoints were added: /api/generate_article_from_texts — generate a short (~A4) Polish article summarising a list of texts; and /api/create_full_article_from_texts — create a fuller article framed by user_query . Embeddings support Dedicated endpoints for generating text embeddings across all supported providers. Simple deployment One‑liner run script, Docker image, or Helm chart for Kubernetes. Extensible conversation formats Basic chat, conversation with system prompt, and extended conversation with richer options (temperature, top‑k, custom system prompt). Multi‑provider model support Each model can be backed by multiple providers (VLLM, Ollama, OpenAI, Anthropic) defined in models-config.json . Model‑level fallback A model can declare fallback_model : when none of its own providers can serve a request, the router reroutes it — before load balancing — to the fallback model, whose providers are balanced by the same strategy ( Models configuration ). Provider failover Any 4xx/5xx answer (or an unreachable provider) re-issues the request on the next provider of the same model — streaming included; only after every provider of the model was tried does fallback_model take over. Load‑balanced default strategy LoadBalancedStrategy picks the least‑…"},{"k":"llm-router-api/docs/installation.html","t":"Installation","s":"Getting started","x":"# Installation Three ways to get **LLM Router** running: from **PyPI** (recommended), from the **GitHub** source repository, or as a pre-built **Docker image on Quay**. Requires **Python ≥ 3.10**. --- ## 1. PIP (recommended) The package is…","h":["1. PIP (recommended)","Create &amp; activate a virtual environment","Install","Extras","2. GitHub (from source)","Option A — clone &amp; install locally","Option B — install directly from the git URL (no clone)","3. Quay (Docker image)","Quick start","Advanced usage","Kubernetes (Helm)"],"b":"Installation # Three ways to get LLM Router running: from PyPI (recommended), from the GitHub source repository, or as a pre-built Docker image on Quay . Requires Python ≥ 3.10 . 1. PIP (recommended) # The package is published on PyPI: radlab-llm-router . Create &amp; activate a virtual environment # python3 -m venv .venv source .venv/bin/activate Install # # Only the core library (llm_router_lib). pip install radlab-llm-router # Core library + API wrapper (llm_router_api). pip install radlab-llm-router [ api ] # Core library + API wrapper + Prometheus metrics. pip install radlab-llm-router [ api,metrics ] Extras # Extra Adds api The REST API wrapper ( llm_router_api ) metrics Prometheus metrics ( prometheus-client ) vault HashiCorp Vault integration ( hvac , bcrypt ) Note: When Prometheus metrics are enabled, LLM_ROUTER_USE_PROMETHEUS=1 must be set and Redis is required (used for provider availability state). The multiproc directory defaults to $HOME/.llm-router/metrics/prometheus/multiproc — override via the PROMETHEUS_MULTIPROC_DIR environment variable if needed. 2. GitHub (from source) # Installing from source pulls the latest code from the llm-router repository — prefer PyPI for production deployments. Option A — clone &amp; install locally # git clone https://github.com/radlab-dev-group/llm-router.git cd llm-router python3 -m venv .venv source .venv/bin/activate # Only the core library (llm_router_lib). pip install . # Core library + API wrapper (llm_router_api). pip install . [ api ] # Core library + API wrapper + Prometheus metrics. pip install . [ api,metrics ] Option B — install directly from the git URL (no clone) # # Pin a release tag (e.g. v1.0.5): pip install \"git+https://github.com/radlab-dev-group/llm-router@v1.0.5\" # With extras: pip install \"git+https://github.com/radlab-dev-group/llm-router@v1.0.5#[api]\" Note: @main tracks the repository HEAD and may contain unreleased, unstable changes — always pin a release tag for reproducible installs. 3. Quay (Docker image) # A pre-built container image is available on Quay : quay.io/radlab/llm-router . Quick start # docker pull quay.io/radlab/llm-router:rc1 docker run -p 5555 :8080 quay.io/radlab/llm-router:rc1 Advanced usage # All runtime settings are configured via LLM_ROUTER_* environment variables — see the Docker section of the README for the full custom launch-script example (server type, timeouts, balance strategy, Redis, models config mount, …), and ENV_DEFINITIONS.md for the complete variable reference. Kubernetes (Helm) # Helm charts for Kubernetes deployment are available in the helm_charts/ directory."},{"k":"llm-router-api/docs/env-definitions.html","t":"Environment Variables","s":"Configuration","x":"# Environment Variables All environment variables share the `LLM_ROUTER_` prefix. They are loaded from `os.environ` in [ `llm_router_api/base/constants.py`](llm_router_api/base/constants.py) at import time and validated by `_StartAppVerifi…","h":["Core variables","Redis variables","Masking &amp; Guardrail variables","Payload masking","Request guardrails","Response guardrails","Semantic BiEncoder Routing variables","LangChainRAG variables","Utils plugins variables","Authentication variables","Core switch","Memory store","Custom policies","Vault settings","Redis cache for keys","Auth Redis (separate from general REDIS)","Rate limiting","Public endpoints","Hardening","Key generation","Key rotation","Audit","Monitoring intervals","Programmatic source"],"b":"Environment Variables # All environment variables share the LLM_ROUTER_ prefix. They are loaded from os.environ in llm_router_api/base/constants.py at import time and validated by _StartAppVerificator . Core variables # Variable Default Description LLM_ROUTER_PROMPTS_DIR resources/prompts Directory containing predefined system prompts. LLM_ROUTER_MODELS_CONFIG resources/configs/models-config.json Path to the models configuration JSON file. LLM_ROUTER_DEFAULT_EP_LANGUAGE pl Default language for endpoint prompts (fallback). LLM_ROUTER_TIMEOUT 0 Timeout (seconds) for llm-router API calls. LLM_ROUTER_EXTERNAL_TIMEOUT 300 Timeout (seconds) for external model API calls. LLM_ROUTER_MAX_REQUEST_BODY_SIZE 10485760 (10 MB) Maximum request body size in bytes; oversized payloads get HTTP 413. LLM_ROUTER_LOG_FILENAME llm-router.log Name of the log file. LLM_ROUTER_LOG_LEVEL INFO Logging level (e.g. INFO, DEBUG). LLM_ROUTER_LOG_TO_FILE false Also write logs to the log file (in addition to console). LLM_ROUTER_LOG_MAX_BYTES 52428800 (50 MB) Rotate the log file once it reaches this size in bytes (applied when LLM_ROUTER_LOG_TO_FILE is set). LLM_ROUTER_LOG_BACKUP_COUNT 5 Maximum number of rotated log files to keep, &lt;name&gt;.1 … &lt;name&gt;.N . LLM_ROUTER_EP_PREFIX /api Prefix for all API endpoints. LLM_ROUTER_MINIMUM False Run service in proxy-only mode. LLM_ROUTER_IN_DEBUG False Run server in debug mode; also forces log level to DEBUG. LLM_ROUTER_VERBOSE False Log RAW, unmasked request params (PII!). Startup logs a warning and waits 3 s. Never use in production. LLM_ROUTER_BALANCE_STRATEGY balanced Load-balancing strategy: balanced , weighted , dynamic_weighted , first_available , first_available_optim , first_available_optim_nworkers . LLM_ROUTER_LB_SLOT_LEASE_SECONDS 120 Lifetime of one held worker slot in first_available_optim_nworkers . Renewed while the router process lives, so it bounds how long a crashed process keeps occupying a provider. Must exceed the longest expected request only if the KeepAlive monitor cannot keep up. LLM_ROUTER_LB_SLOT_MAX_AGE_SECONDS 1800 How long an unreleased worker slot keeps being renewed, counted from the acquisition. Covers a request that never released its slot (an abandoned stream): it is left to expire instead of occupying the provider until the router restarts. 0 or negative disables the cap. LLM_ROUTER_LB_HOST_PIN_TTL_SECONDS 3600 Lifetime of the “this host serves this model” pin used by first_available_optim / first_available_optim_nworkers (Redis key host:&lt;host&gt; ). Refreshed on every selection, so only a host nothing has selected for this long becomes available to another model. LLM_ROUTER_SERVER_TYPE flask Server implementation: flask, gunicorn, waitress. LLM_ROUTER_SERVER_PORT 8080 Port on which the server listens. LLM_ROUTER_SERVER_HOST localhost Host address for the server. LLM_ROUTER_SERVER_WORKERS_COUNT 2 Number of workers (for servers that support them). LLM_ROUTER_SERVER_THREADS_COUNT 8 Number of worker threads (for servers that support them). LLM_ROUTER_SERVER_WORKER_CLASS None Worker class for servers that support it (e.g. gevent). LLM_ROUTER_USE_PROMETHEUS False Enable Prometheus metrics collection ( /metrics endpoint). See also PROMETHEUS_MULTIPROC_DIR for the directory where Prometheus multiprocess worker data files are stored. LLM_ROUTER_VERBOSE is a debugging aid, not a logging level. With this variable on, SecureEndpointI stores it as self._verbose_mode and run_ep writes the raw, unmasked request payload (JSON) to the log before masking, so the log contains exactly the PII that masking would otherwise strip. On top of that, rest_api.main() emits a WARNING and sleeps 3 s ( VERBOSE_STARTUP_DELAY_SECONDS ) before the server binds its port, so an accidental production start can still be aborted. Enable it only for local troubleshooting ( LLM_ROUTER_VERBOSE=1 , llm-router server start --verbose ) and never in production or on shared log storage. Redis variables # Variable De…"},{"k":"llm-router-api/docs/models-config.html","t":"Models configuration description (JSON)","s":"Configuration","x":"# Models configuration description (JSON) ## 📄 Purpose This document explains the **model configuration** used by the LLM Router. It describes the JSON schema that drives **`ModelHandler`** and **`ApiModelConfig`**, clarifies each field, a…","h":["📄 Purpose","🏗️ High‑level structure","🔎 Detailed field description","Provider dictionary (items in providers / providers_sleep)","🛟 fallback_model (model level)","🔁 Failover order (providers first, then fallback_model)","active_models section","🧩 How ModelHandler uses the config","📦 Sample configuration (models-config.json)","🎉 Summary"],"b":"Models configuration description (JSON) # 📄 Purpose # This document explains the model configuration used by the LLM Router. It describes the JSON schema that drives ModelHandler and ApiModelConfig , clarifies each field, and provides a ready‑to‑use example ( models-config.json ). Having a single source of truth for model definitions makes it easy to: Add or remove providers for a given model. Switch between cloud (OpenAI, Google) and local (vLLM, Ollama) back‑ends. Control load‑balancing, keep‑alive, and tool‑calling options per provider. Activate only the models you want to expose through the router. 🏗️ High‑level structure # { \"&lt;model_type&gt;\": { # e.g. \"google_models\", \"openai_models\", \"qwen_models\" \"&lt;model_name&gt;\": { # full identifier used by the router, e.g. \"google/gemma-3-12b-it\" \"providers\": [ … ], # primary providers (used for normal traffic) \"providers_sleep\": [ … ] # optional low‑priority providers (used when others are busy) \"fallback_model\": \"…\" # optional model used when no provider can serve this one }, … }, \"active_models\": { # **required** – tells the router which models are enabled \"&lt;model_type&gt;\": [ \"&lt;model_name&gt;\", … ], … } } Model type – a top‑level key grouping models that share the same provider‑type logic. Model name – the identifier that appears in API calls ( model field). providers – a list of dictionaries, each describing a concrete endpoint. providers_sleep (optional) – “sleeping” providers that are only used when all primary providers are unavailable or overloaded. fallback_model (optional) – name of another active model that takes over when no provider of this model can serve the request. See Fallback model . active_models – the only place where a model is marked as active . If a model is missing here, the router will ignore it even if it is present in the rest of the file. 🔎 Detailed field description # Provider dictionary (items in providers / providers_sleep ) # Field Type Description Example id str Unique identifier for the provider instance (used for logging &amp; selection). \"gemma3_12b-vllm-71:7000\" api_host str Base URL of the provider API (must include protocol, may contain trailing slash). \"http://192.168.100.71:7000/\" api_token str Authentication token; empty string if not required. \"\" api_type str Type of the backend – determines which concrete BaseProvider class is used ( openai , vllm , ollama , …). \"vllm\" input_size int (or numeric string) Maximum context length the provider accepts. The ApiModel.from_config helper converts it to int . 4096 model_path str Path or name of the model on the provider side (used by Ollama, vLLM, etc.). May be empty for providers that infer it from the URL. \"gpt-3.5-turbo-0125\" weight float Relative weight for weighted‑random load‑balancing strategies. Default 1.0 . 0.1 nworkers int (or numeric string) Maximum number of concurrent requests allowed on the provider. Used only by the first_available_optim_nworkers load‑balancing strategy; a missing, invalid or non‑positive value falls back to 1 . Read from the live configuration on every selection, so changing it takes effect on the next request — the health monitor's registration copy ( monitor:providers:&lt;model&gt; ) is not used for the limit. 1 keep_alive str Optional keep‑alive duration (e.g. \"35m\" ). Empty or null means the provider is not kept alive. \"35m\" tool_calling bool Whether the provider supports tool‑calling (function calling). true is_embedding bool Whether the model is an embedding model (determines use of embedding endpoints). true 🛟 fallback_model (model level) # A model may name another active model that takes over when none of its own providers can serve a request. The switch happens before load balancing : the fallback model is handed to the load-balancing strategy, which then picks one of its providers exactly like for any other request. { \"qwen_models\" : { \"qwen/Qwen3.8-Flash-Next\" : { \"fallback_model\" : \"qwen/qwen3-coder:30b\" , \"providers\" : [ { \"id\" : \"qwen3.8…"},{"k":"llm-router-api/docs/keepalive.html","t":"Keep‑Alive Utility Overview","s":"Routing & resilience","x":"# Keep‑Alive Utility Overview The **keep‑alive** subsystem is responsible for periodically “pinging” model endpoints so that they stay warm and ready to serve requests with minimal latency. It consists of two main components: | Component |…","h":["How It Works","Integration Points","Configuring Keep‑Alive","Example Usage","Logging"],"b":"Keep‑Alive Utility Overview # The keep‑alive subsystem is responsible for periodically “pinging” model endpoints so that they stay warm and ready to serve requests with minimal latency. It consists of two main components: Component Purpose Key Methods KeepAlive Sends a single HTTP request to a model provider. It resolves the correct provider configuration (API type, host, token, model name) and builds the request payload. send(model_name, host, prompt=None) – performs the HTTP call. KeepAliveMonitor Schedules repeated keep‑alive calls for each (model_name, host) pair. It stores scheduling data in Redis, checks host availability, and triggers KeepAlive.send when a provider is due. record_usage(model_name, host, keep_alive) – registers a provider for periodic pinging. start() / stop() – control the background thread. How It Works # Provider discovery – KeepAlive._find_provider looks up the provider configuration for a given model name and host inside the global models_configs dictionary. Endpoint resolution – Depending on the provider’s api_type ( vllm , openai , ollama ), _endpoint_for builds the correct URL ( /v1/chat/completions or /api/chat ). HTTP request – A JSON payload containing a short “keep‑alive” prompt (default: “Send an empty message.” ) is posted to the endpoint. Scheduling – KeepAliveMonitor stores metadata in Redis: A hash key ( keepalive:provider:&lt;model&gt;:&lt;host&gt; ) with keep_alive_seconds . A sorted‑set ( keepalive:providers:next_wakeup ) that orders providers by the next scheduled wake‑up timestamp. A sorted‑set ( keepalive:model:&lt;model_name&gt;:hosts ) that tracks all hosts where a model is currently loaded, used for host reuse in optimised strategies. Background loop – The monitor thread wakes up every check_interval seconds, fetches due providers, verifies that the host is free (via the optional is_host_free_callback ), and invokes KeepAlive.send . After a successful ping, the next wake‑up time is recomputed. Integration Points # Strategy implementations (e.g., FirstAvailableOptimStrategy ) create a KeepAlive instance and pass it to a KeepAliveMonitor . When a provider is selected, the strategy calls keep_alive_monitor.record_usage(model_name, host, keep_alive) so the monitor knows to ping that endpoint. The monitor runs automatically in the background once start() is called (typically during strategy initialization). Configuring Keep‑Alive # Provider configurations live in the global models_configs JSON (see resources/configs/models-config*.json ). To enable keep‑alive for a specific provider, add a keep_alive field with a duration string: { \"model_name\" : \"gpt‑4\" , \"providers\" : [ { \"api_type\" : \"openai\" , \"api_host\" : \"http://localhost:8000\" , \"api_token\" : \"YOUR_TOKEN\" , \"keep_alive\" : \"2m\" // ping every 2 minutes } ] } Supported duration units: Unit Suffix Meaning seconds s e.g., \"30s\" minutes m e.g., \"5m\" hours h e.g., \"1h\" If the keep_alive field is omitted or falsy, the provider will not be scheduled for periodic pings. Example Usage # from llm_router_api.core.monitor.keep_alive import KeepAlive from llm_router_api.core.monitor.keep_alive_monitor import KeepAliveMonitor # Assume `models_configs` has been loaded from the JSON config files. keep_alive = KeepAlive ( models_configs = models_configs ) monitor = KeepAliveMonitor ( redis_client = redis_client , keep_alive = keep_alive , check_interval = 10.0 , # check every 10 seconds is_host_free_callback = my_is_host_free , clear_buffers = True , # clean old keys on start ) monitor . start () # When a provider is selected somewhere in the routing logic: monitor . record_usage ( model_name = \"gpt‑4\" , host = \"http://localhost:8000\" , keep_alive = \"2m\" ) The monitor will now ping the gpt‑4 endpoint every two minutes, provided the host is not busy with another model. Logging # Both KeepAlive and KeepAliveMonitor emit detailed logs at the DEBUG and INFO levels, prefixed with [keep-alive] . Adjust your logger configuration to capture these messa…"},{"k":"llm-router-api/docs/lb-strategies.html","t":"Load Balancing Strategies","s":"Routing & resilience","x":"## Load Balancing Strategies The `llm-router` supports various strategies for selecting the most suitable provider when multiple options exist for a given model. This ensures efficient and reliable routing of requests. The available strate…","h":["Load Balancing Strategies","1. balanced (Default)","2. weighted","3. dynamic_weighted (beta)","4. first_available","5. first_available_optim","6. first_available_optim_nworkers","7. AdaptiveStrategy (beta)","Environments and Redis installation","Extending with Custom Strategies"],"b":"Load Balancing Strategies # The llm-router supports various strategies for selecting the most suitable provider when multiple options exist for a given model. This ensures efficient and reliable routing of requests. The available strategies are: Model‑level fallback_model runs before the strategy. A strategy is always asked to serve the model named by the client. Only when that model cannot be served at all — no providers, no healthy provider, or every provider busy until the selection timeout — ModelHandler reroutes the request to the configured fallback_model and the very same strategy then balances over the providers of that model . Strategies themselves are unaffected; the option is documented in MODELS_CONFIG.md . Provider errors rotate first.— A provider that answers with an error (any 4xx/5xx) or is unreachable makes the dispatcher retry the request on another provider of the same model ; the model's fallback_model is reached only once none of its providers is left untried. 1. balanced (Default) # Description: This is the default strategy. It aims to distribute requests evenly across available providers by keeping track of how many times each provider has been used for a specific model. It selects the provider that has been used the least. When to use: Ideal for scenarios where all providers are considered equal in terms of capacity and performance. It provides a simple and effective way to balance the load. Implementation: Implemented in llm_router_api.core.lb.balanced.LoadBalancedStrategy . 2. weighted # Description: This strategy allows you to assign static weights to providers. Providers with higher weights are more likely to be selected. The selection is deterministic, ensuring that over time, the request distribution closely matches the configured weights. When to use: Useful when you have providers with different capacities or performance characteristics, and you want to prioritize certain providers without needing dynamic adjustments. Implementation: Implemented in llm_router_api.core.lb.weighted.WeightedStrategy . 3. dynamic_weighted (beta) # Description: An extension of the weighted strategy. It not only uses weights but also tracks the latency between successive selections of the same provider. This allows for more adaptive routing, as providers with consistently high latency might be de-prioritized over time. You can also dynamically update provider weights. When to use: Recommended for dynamic environments where provider performance can fluctuate. It offers more sophisticated load balancing by considering both configured weights and real-time performance metrics (latency). Implementation: Implemented in llm_router_api.core.lb.weighted.DynamicWeightedStrategy . 4. first_available # Description: This strategy selects the very first provider that is available. It uses Redis to coordinate across multiple workers, ensuring that only one worker can use a specific provider at a time. When to use: Suitable for critical applications where you need the fastest possible response and want to ensure that a request is immediately handled by any available provider, without complex load distribution logic. It guarantees that a provider, once taken, is exclusive until released. Implementation: Implemented in llm_router_api.core.lb.first_available.FirstAvailableStrategy . When using the first_available load balancing strategy, a Redis server is required for coordinating provider availability across multiple workers. 5. first_available_optim # What it is first_available_optim is an enhanced version of the plain first‑available load‑balancing strategy. It uses Redis to coordinate across multiple workers and tries to reuse a host that has already been used for the requested model before falling back to the classic “pick the first free provider” logic. How it works Step Purpose Behaviour 1️⃣ Re‑use the last host If the model was previously run on a specific host and that host is currently free, select it. The host identifier is…"},{"k":"llm-router-api/core/auditor/index.html","t":"Auditing subsystem – llm-router","s":"Security & auditing","x":"# Auditing subsystem – `llm-router` The **auditor** package provides a pluggable, tamper‑evident audit‑log system for the LLM‑router. All audit entries are written as JSON, encrypted with GPG and stored under `logs/auditor`. The subsystem …","h":["📁 Directory layout","🛠️ How the auditor works","🔐 GPG key management","1️⃣ Generate a new key pair","2️⃣ Place the public key where the router expects it","3️⃣ Decrypt audit logs","📚 Example: Auditing a request guard‑rail decision","🧩 Extending the auditor","📖 Further reading"],"b":"Auditing subsystem – llm-router # The auditor package provides a pluggable, tamper‑evident audit‑log system for the LLM‑router. All audit entries are written as JSON, encrypted with GPG and stored under logs/auditor . The subsystem is used by the router to record: request guard‑rail decisions payload masking operations custom audit events emitted by the application (e.g. business‑logic logs) The implementation is deliberately lightweight so it can be swapped out for a different storage backend (database, cloud bucket, …) without touching the rest of the code base. 📁 Directory layout # llm_router_api/ └─ core/ └─ auditor/ ├─ __init__.py # package marker ├─ auditor.py # public API – AnyRequestAuditor └─ log_storage/ ├─ __init__.py ├─ log_storage_interface.py # abstract storage contract └─ gpg.py # GPG‑backed storage implementation auditor.py – high‑level helper that forwards audit records to a storage backend. The default backend is GPGAuditorLogStorage . log_storage_interface.py – defines the AuditorLogStorageInterface protocol ( store_log(audit_log, audit_type) ). gpg.py – concrete implementation that encrypts each log entry with the public GPG key located at resources/keys/llm-router-auditor-pub.asc and writes the encrypted payload to a timestamped file logs/auditor/&lt;audit_type&gt;__&lt;timestamp&gt;.audit . 🛠️ How the auditor works # Endpoint code (e.g. endpoint_i.py ) creates an AnyRequestAuditor instance with the router’s logger. When an auditable event occurs, the endpoint builds a dictionary that contains at least the keys audit_type and payload . AnyRequestAuditor.add_log() forwards the dictionary to the configured storage backend. GPGAuditorLogStorage.store_log() JSON‑serialises the dictionary (pretty‑printed). Encrypts the JSON string with the imported public key. Writes the encrypted ASCII‑armored data to logs/auditor/ . The resulting files have the extension .audit . They are confidential and tamper‑evident – any modification breaks the GPG decryption. 🔐 GPG key management # The repository ships two helper scripts under scripts/ : Script Purpose gen_and_export_gpg.sh Generates a 4096‑bit RSA key pair (no interactive prompts) and exports the public ( *.asc ) and private ( *-priv.asc ) keys. decrypt_auditor_logs.sh Decrypts all *.audit files in logs/auditor/ and writes the resulting JSON to *.json . 1️⃣ Generate a new key pair # cd scripts ./gen_and_export_gpg.sh The script will: Prompt for an email address (used as the GPG user ID). Prompt for a passphrase (protects the private key). Create a key pair in the local GPG keyring. Export the public key to llm-router-auditor-pub.asc . Export the private key to llm-router-auditor-priv.asc . Important: Keep the private key ( *-priv.asc ) and its passphrase safe. Only the public key is required by the router at runtime. 2️⃣ Place the public key where the router expects it # mkdir -p resources/keys cp llm-router-auditor-pub.asc resources/keys/ The GPGAuditorLogStorage class automatically imports the key from this location when the application starts. 3️⃣ Decrypt audit logs # cd scripts ./decrypt_auditor_logs.sh For each file logs/auditor/&lt;type&gt;__&lt;timestamp&gt;.audit the script produces a human‑readable *.json file next to it: logs/auditor/request__20231129_123456.789012.audit → request__20231129_123456.789012.json You will be prompted for the passphrase of the private key if it is encrypted. 📚 Example: Auditing a request guard‑rail decision # from llm_router_api.core.auditor.auditor import AnyRequestAuditor import logging logger = logging . getLogger ( \"router\" ) auditor = AnyRequestAuditor ( logger ) # Somewhere inside an endpoint, after a guard‑rail check: audit_record = { \"audit_type\" : \"guardrail_request\" , \"payload\" : { \"user_id\" : \"12345\" , \"input\" : \"…\" , \"decision\" : \"blocked\" , \"reason\" : \"PII detected\" } } auditor . add_log ( audit_record ) The record is encrypted and persisted as e.g.: logs/auditor/guardrail_request__20231129_141530.123456.audit 🧩 Exte…"},{"k":"llm-router-api/docs/authentication.html","t":"Authentication & Authorization","s":"Security & auditing","x":"# Authentication & Authorization API-key-based authentication with per-endpoint policies, rate limiting, audit trail, and Prometheus metrics. **Enabled by default: `LLM_ROUTER_AUTH_ENABLED=false`** — set to `\"true\"` to enforce authenticati…","h":["Architecture","Components","Environment Variables","CLI Commands","Key Management","Policy Management","Rate Limit","Seed File (Memory Store)","Seed File Format","How It Works","Default Location","Permission Engine","Builtin Policies","Prometheus Metrics","Key Format","Deployment Options","1️⃣ In-Memory Store — Development / Quick Start","2️⃣ Redis Store — Multi-Process / Stateful Single-Node","3️⃣ HashiCorp Vault — Production / Multi-Cluster / Enterprise","Comparison Matrix","See Also"],"b":"Authentication &amp; Authorization # API-key-based authentication with per-endpoint policies, rate limiting, audit trail, and Prometheus metrics. Enabled by default: LLM_ROUTER_AUTH_ENABLED=false — set to \"true\" to enforce authentication. Architecture # Client Request → AuthMiddleware → Key Store Lookup → Permission Engine → Rate Limiter → Endpoint → Audit Bridge → AnyRequestAuditor → AuthMetrics → Prometheus Components # Component Module Purpose Key Store core/auth/key_store/ Vault, Redis, or in-memory key storage Permission Engine core/auth/policies/engine.py Resolve key → policy → endpoint permissions Rate Limiter core/auth/rate_limiter.py Redis-backed sliding window rate limiter Key Generator core/auth/key_generator.py Generate keys in sk-llmr-live- format Audit Bridge core/auth/audit.py Bridge auth events → AnyRequestAuditor Metrics core/auth/metrics.py Prometheus counters &amp; histograms for auth Middleware core/auth/middleware.py Flask before_request hook Environment Variables # All environment variables are documented in ENV_DEFINITIONS.md . Auth-specific vars start with LLM_ROUTER_AUTH_* . CLI Commands # Key Management # # Generate a new key (persists to seed file when --store memory) llm-router auth key generate --policy developer --store memory # List all keys llm-router auth key list --store memory # Delete a key llm-router auth key delete key-id # Disable a key (revokes access immediately) llm-router auth key disable key-id [ --store memory ] # Enable a previously disabled key llm-router auth key enable key-id [ --store memory ] # Rotate a key (old key stays valid for grace_period) llm-router auth key rotate key-id --grace 3600 Policy Management # # List builtin policies (custom ones are marked with \"(custom)\") llm-router auth policy list # Create a new policy — inline JSON, from a file, or from stdin (-) llm-router auth policy create my-team '{\"can_access\": true, \"rate_limit\": 120}' llm-router auth policy create my-team --file my-team.json cat my-team.json | llm-router auth policy create my-team --file - Custom policies are persisted to $LLM_ROUTER_AUTH_CUSTOM_POLICIES_FILE (default: ~/.llm-router/configs/auth/custom-policies.json ) and resolved by the server without a restart. Rate Limit # # List available rate-limit presets llm-router auth rate-limit list # Apply a rate-limit preset to an existing key llm-router auth rate-limit apply key-id --preset pro # Remove rate-limit override from a key (revert to global default) llm-router auth rate-limit remove key-id Seed File (Memory Store) # When using --store memory , keys are stored in process memory — they are lost on restart . To persist keys across restarts (and between the CLI and router processes), use a seed file. Seed File Format # The seed file is a JSON array. Each record must carry a verifiable credential: a key_hash (+ key_index ) pair, or — for legacy files only — a key_plain which is hashed at load time and never written back . Plaintext keys are never persisted. Field Type Description key_id str Unique identifier for this key key_hash str bcrypt hash of the plaintext key (the verifiable credential) key_index str SHA-256 of the plaintext key — O(1) lookup index (locator only) key_prefix str First 12 characters of the plaintext key (shows the full sk-llmr-live prefix) policy_name str Name of the default policy to apply policy_override dict Inline policy override (takes precedence over the named policy) is_active bool Whether the key is currently valid expires_at float Expiry timestamp ( null = no expiry) created_at float Unix timestamp of key creation (auto-generated if omitted) last_used_at float Last successful authentication time rotate_at float Scheduled rotation time grace_until float Key remains valid until this time after rotation metadata dict Arbitrary metadata (team, cost_center, etc.) [ { \"key_id\" : \"manual-key-001\" , \"key_hash\" : \"$2b$12$Kx7Q2m...bcrypt-hash.../\" , \"key_index\" : \"9f2c...sha256-hex...\" , \"key_prefix\" : \"sk-llmr-live\" , \"pol…"},{"k":"llm-router-api/docs/rate-limiting.html","t":"Rate Limiting","s":"Security & auditing","x":"# Rate Limiting Sliding-window rate limiting backed by Redis sorted sets. Prevents API abuse, controls load on downstream LLM providers, and protects against brute-force key enumeration. --- ## How It Works ### Sliding Window Algorithm The…","h":["How It Works","Sliding Window Algorithm","Key + IP Binning","Bucket Naming","Client IP Resolution (anti-spoofing)","Failed-Authentication Lockout","Configuration","Environment Variables","Enabling Rate Limiting","Response Behavior","Architecture","RateLimitResult Dataclass","RedisRateLimiter Class","Prometheus Metrics","Rate Limit Strategy by Use Case","Per-User Quotas (Default)","Tiered Rate Limiting (Policy-Based)","Redis Requirements","Redis Persistence","Comparison with Fixed-Window","Why Sliding Window?","Best Practices","1. Always Enable Rate Limiting in Production","2. Set up auth Redis connection","3. Monitor Rate Limit Events","4. Set Graceful Limits","5. Use Public Endpoints for Health Checks","Migration from No-Auth / Default Rate","Troubleshooting","\"Why am I getting 429s immediately?\"","\"Rate limit is too strict/loose\"","\"Clients aren't respecting Retry-After\"","Rate Limit Presets","See Also"],"b":"Rate Limiting # Sliding-window rate limiting backed by Redis sorted sets. Prevents API abuse, controls load on downstream LLM providers, and protects against brute-force key enumeration. How It Works # Sliding Window Algorithm # The rate limiter uses a sliding window approach — unlike fixed-window counters, there are no boundary spikes. Each request timestamp is stored as a member score in a Redis sorted set, and entries older than the window are purged on every check. Time → ──────────────────────────────────────────▶ │&lt;────── WINDOW (60s) ──────▶│ ▼ ▼ │ ◉ ◉ ◉ ◉ ◉ ◉ ◉ ◉ │ │ ◉ ◉ ◉ ◉ ◉ ◉ ◉ ◉ ◉ ◉ │ │ │ ←─ remaining slots ────────→ Precision: per-request, no fixed-boundary artifacts Memory: O (requests per key+IP) — old entries are automatically purged Durability: Redis persistence (RDB/AOF) protects against server restart Key + IP Binning # Rate limits are enforced per API key + client IP . This means: Each API key gets its own independent quota Multiple IPs sharing the same key share that key's quota A malicious IP hitting a leaked key can exhaust its quota (mitigated by IP-level monitoring) Bucket Naming # auth:ratelimit:{key_id}:{ip} Example: auth:ratelimit:dev-a1b2c3d3:192.168.1.100 Client IP Resolution (anti-spoofing) # The client IP is the direct peer ( remote_addr ) by default. The X-Forwarded-For header is attacker-controlled and is honoured only when the direct peer is a configured trusted proxy ( LLM_ROUTER_TRUSTED_PROXIES , IP or CIDR); in that case the right-most XFF entry is used (proxies append, not prepend). Without this gate, a client could rotate X-Forwarded-For to obtain a fresh rate-limit bucket per request. export LLM_ROUTER_TRUSTED_PROXIES = \"10.0.0.1,192.168.0.0/24\" # your LB / reverse proxy Failed-Authentication Lockout # Failed attempts (missing or invalid key) never reach the key-scoped limiter, so they are throttled separately in a dedicated per-IP bucket ( auth:ratelimit:__auth_fail__:{ip} ) with the LLM_ROUTER_AUTH_FAILURE_LIMIT budget (default 20 /window; 0 disables). Once exceeded, the client receives 429 — protecting the key-lookup path from brute-force / DoS. Configuration # Environment Variables # All rate-limiting environment variables are documented in ENV_DEFINITIONS.md → Authentication section (Rate limiting subsection) . Enabling Rate Limiting # Rate limiting is applied automatically when authentication is enabled — there is no separate toggle. Just enable auth and configure the default rate limit: export LLM_ROUTER_AUTH_ENABLED = true export LLM_ROUTER_AUTH_DEFAULT_RATE_LIMIT = 60 # 60 requests per minute per key (default) python -m llm_router_api.rest_api Response Behavior # When a request exceeds the rate limit, the router returns: HTTP Status Header Value 429 Too Many Requests Retry-After Seconds until the oldest request in the window expires Example: HTTP/1.1 429 Too Many Requests Retry-After: 35 {\"error\": {\"message\": \"Rate limit exceeded. Please retry later.\", \"type\": \"rate_limit_error\", \"code\": 429, \"retry_after\": 35}} The Retry-After value is calculated from the oldest entry still in the window — this tells the client exactly when it can retry. Architecture # Client Request → AuthMiddleware → get_auth_result() → RedisRateLimiter.is_allowed(key_id, ip, limit) → → Redis sorted set operations (zremrangebyscore, zcard, zadd, expire) → RateLimitResult(allowed, remaining, retry_after) → If denied: HTTP 429 → If allowed: continue to endpoint RateLimitResult Dataclass # @dataclass class RateLimitResult : allowed : bool # Whether the request is within the limit remaining : int # Remaining requests in the current window retry_after : int # Seconds until the oldest request expires (0 if allowed) RedisRateLimiter Class # class RedisRateLimiter : PREFIX = \"auth:ratelimit\" WINDOW = 60 # seconds def __init__ ( self , redis_client : Optional [ redis . Redis ] = None , redis_host : Optional [ str ] = None , redis_port : int = 6379 , redis_db : int = 0 , redis_password : Optional [ str ] = None , …"},{"k":"llm-router-api/docs/routing-metrics.html","t":"Router Prometheus Metrics","s":"Observability","x":"# Router Prometheus Metrics Additional Prometheus metrics for the LLM router core (routing, provider lifecycle, pipeline funnel). Enabling these metrics requires `LLM_ROUTER_USE_PROMETHEUS=true` and the `metrics` extra (`pip install .[metr…","h":["Overview","A. Routing &amp; Provider Metrics","llm_router_provider_calls_total (Counter)","llm_router_provider_latency_seconds (Histogram)","llm_router_provider_error_total (Counter)","llm_router_lb_strategy_selected_total (Counter)","B. Pipeline / Request Funnel Metrics","llm_router_pipeline_stage_total (Counter)","llm_router_retry_total (Counter)","llm_router_retry_exhausted_total (Counter)","C. Token Usage Metric","llm_router_tokens_total (Counter)","D. Streaming &amp; Response Format Metrics","llm_router_response_format_total (Counter)","llm_router_payload_conversion_total (Counter)","Grafana Dashboard Snippet","Implementation Notes"],"b":"Router Prometheus Metrics # Additional Prometheus metrics for the LLM router core (routing, provider lifecycle, pipeline funnel). Enabling these metrics requires LLM_ROUTER_USE_PROMETHEUS=true and the metrics extra ( pip install .[metrics] ). Overview # The router tracks 10 new metrics across four categories: Category Metrics Purpose Routing &amp; Provider 4 Visibility into provider selection, latency, errors, and LB strategy Pipeline Funnel 3 Track how many requests pass through each pipeline stage Token Usage 1 Track input/output token consumption per model Streaming &amp; Format 2 Track streaming vs non-streaming distribution and payload conversions A. Routing &amp; Provider Metrics # llm_router_provider_calls_total (Counter) # Total number of successful calls to each provider type, grouped by model name. Labels: provider_type , model_name Example output: llm_router_provider_calls_total{provider_type=\"openai\", model_name=\"google/gemma-3-12b-it\"} 42 llm_router_provider_calls_total{provider_type=\"ollama\", model_name=\"google/gemma-3-12b-it\"} 18 Use case: Track provider utilization distribution across models. llm_router_provider_latency_seconds (Histogram) # Latency of outbound calls to specific providers, independent of end-to-end client latency. Labels: provider_type , model_name Buckets: 10ms → 30s (0.01, 0.025, 0.05, 0.1, 0.25, 0.5, 1, 2.5, 5, 10, 30) Example output: llm_router_provider_latency_seconds_count{provider_type=\"vllm\", model_name=\"google/gemma-3-12b-it\"} 100 llm_router_provider_latency_seconds_sum{provider_type=\"vllm\", model_name=\"google/gemma-3-12b-it\"} 45.2 Use case: Identify slow/degraded providers independently of client-side factors. Query example: histogram_quantile ( 0.95 , rate ( llm_router_provider_latency_seconds_bucket [ 5m ] )) llm_router_provider_error_total (Counter) # Provider errors by type and HTTP error code (for retriable status codes) or classification (timeout, connection_error). Labels: provider_type , model_name , error_code Example output: llm_router_provider_error_total{provider_type=\"ollama\", model_name=\"gpt-oss:120b\", error_code=\"429\"} 5 llm_router_provider_error_total{provider_type=\"openai\", model_name=\"google/gemma-3-12b-it\", error_code=\"timeout\"} 2 Use case: Alert on specific provider error spikes. Query example: sum by ( provider_type ) ( rate ( llm_router_provider_error_total [ 5m ] )) llm_router_lb_strategy_selected_total (Counter) # Load balancing strategy selections per model, tracking which strategies are used for each model's providers. Labels: strategy , model_name Example output: llm_router_lb_strategy_selected_total{strategy=\"balanced\", model_name=\"google/gemma-3-12b-it\"} 80 llm_router_lb_strategy_selected_total{strategy=\"weighted\", model_name=\"openai/gpt-4\"} 35 Use case: Verify load balancing strategy distribution. Useful when debugging or tuning LB strategies. B. Pipeline / Request Funnel Metrics # llm_router_pipeline_stage_total (Counter) # Request counts at each pipeline stage with pass/fail result. Stages tracked: Stage Possible Results provider_resolved success , failure request_received total (always) guardrail_request pass , block masking applied , skipped Labels: stage , result Example output: llm_router_pipeline_stage_total{stage=\"provider_resolved\", result=\"success\"} 1200 llm_router_pipeline_stage_total{stage=\"guardrail_request\", result=\"block\"} 45 Use case: Build a request funnel chart to see where requests drop off. Query example: # Guardrail block rate sum ( rate ( llm_router_pipeline_stage_total { stage = \" guardrail_request \", result = \" block \"}[ 5m ] )) / sum ( rate ( llm_router_pipeline_stage_total { stage = \" guardrail_request \"}[ 5m ] )) llm_router_retry_total (Counter) # Retry attempts per model and HTTP error code that triggered the retry. Every 4xx/5xx answer moves the request to another provider of the model (and, once all of them were tried, to its fallback_model ), so any error status can appear here. Labels: model_name , error_code Example outpu…"},{"k":"llm-router-api/index.html","t":"REST API reference","s":"REST API","x":"# llm‑router‑api **llm‑router‑api** is a lightweight Python library that provides a flexible, extensible proxy for Large Language Model (LLM) back‑ends. It abstracts the details of multiple model providers (OpenAI‑compatible, Ollama, vLLM,…","h":["Features","Installation","Running the Server","REST API Overview","Load‑Balancing Strategies","Keep‑Alive Mechanism","Extending the Router","Adding a New Provider Type","Adding a New Endpoint","Prompt Files","Monitoring &amp; Metrics","License"],"b":"llm‑router‑api # llm‑router‑api is a lightweight Python library that provides a flexible, extensible proxy for Large Language Model (LLM) back‑ends. It abstracts the details of multiple model providers (OpenAI‑compatible, Ollama, vLLM, LM Studio, etc.) and offers a unified REST interface with built‑in load‑balancing, health‑checking, and monitoring. Repository: https://github.com/radlab-dev-group/llm-router Features # Unified API – One REST surface ( /api/... ) that proxies calls to any supported LLM back‑end. Provider Selection – Choose a provider per request using pluggable strategies (balanced, weighted, adaptive, first‑available). Prompt Management – System prompts are stored as files and can be dynamically injected with placeholder substitution. Streaming Support – Transparent streaming for both OpenAI‑compatible and Ollama endpoints. Health Checks – Built‑in ping endpoint and Redis‑based provider health monitoring. Prometheus Metrics – Optional instrumentation for request counts, latencies, and error rates. Auto‑Discovery – Endpoints are automatically discovered and instantiated at startup. Extensible – Add new providers, strategies, or custom endpoints with minimal boilerplate. Installation # The project uses Python 3.10.6 and a virtualenv ‑based workflow. ```shell script Clone the repository # git clone https://github.com/radlab-dev-group/llm-router.git cd llm-router Create a virtual environment # python3 -m venv venv source venv/bin/activate Install the package (including optional extras) # pip install -e .[metrics] # installs Prometheus support All required third‑party libraries are listed in `requirements.txt` (e.g., Flask, requests, redis, rdl‑ml‑utils, etc.). --- ## Configuration Configuration is driven primarily by environment variables and a JSON model‑config file. ### Environment Variables All environment variables are documented in **[ENV_DEFINITIONS.md](./docs/ENV_DEFINITIONS.md)**. Key categories: **Core** · **Redis** · **Masking &amp; Guardrail** · **Semantic BiEncoder Routing** · **LangChainRAG** · * *Utils Plugins** · **Authentication**. --- ### Authentication variables Auth-specific environment variables are documented in **[ENV_DEFINITIONS.md](./docs/ENV_DEFINITIONS.md) → Authentication section**. &gt; See full authentication docs: **[AUTHENTICATION.md](docs/AUTHENTICATION.md)** ### Model Configuration `models-config.json` follows the schema: ```json { \"active_models\": { \"openai_models\": [ \"gpt-4\", \"gpt-3.5-turbo\" ], \"ollama_models\": [ \"llama2\" ] }, \"openai_models\": { \"gpt-4\": { \"providers\": [ { \"id\": \"openai-gpt4-1\", \"api_host\": \"https://api.openai.com/v1\", \"api_token\": \"sk-...\", \"api_type\": \"openai\", \"input_size\": 8192, \"model_path\": \"\" } ] } }, ... } Only the fields required by the router are needed: id , api_host , api_token (optional), api_type , input_size , and optionally model_path . Configuration Details – see the full schema and a ready‑made example in MODELS_CONFIG.md . Running the Server # The entry point is llm_router_api.rest_api . Choose a server backend via the LLM_ROUTER_SERVER_TYPE variable or command‑line flags. ```shell script Using the built‑in Flask development server (default) # python -m llm_router_api.rest_api Production‑grade with Gunicorn (streaming supported) # python -m llm_router_api.rest_api --gunicorn Windows‑friendly Waitress server # python -m llm_router_api.rest_api --waitress ``` The server starts on the host/port defined by LLM_ROUTER_SERVER_HOST and LLM_ROUTER_SERVER_PORT (default 0.0.0.0:8080 ). Note: The service must be launched with LLM_ROUTER_MINIMUM=1 (or any truthy value) because it operates in “proxy‑only” mode. REST API Overview # All routes are prefixed by LLM_ROUTER_EP_PREFIX (default /api ). The list of endpoints—categorized into built‑in, provider‑dependent, and extended endpoints—and a description of the streaming mechanisms can be found at the link: load endpoints overview Creating a new endpoint? See the Endpoint Development Guide — class hierarchy, …"},{"k":"llm-router-api/endpoints/index.html","t":"Endpoints Overview","s":"REST API","x":"## Endpoints Overview > **Creating new endpoints?** See the [Endpoint Development Guide](../docs/ENDPOINT_DEV.md) > (`llm_router_api/docs/ENDPOINT_DEV.md`). All endpoints are exposed under the REST API service. Unless stated otherwise, met…","h":["Endpoints Overview","Authentication","Health &amp; Info","Auth‑required Endpoints","Streaming vs. Non‑Streaming Responses","Payload format"],"b":"Endpoints Overview # Creating new endpoints? See the Endpoint Development Guide ( llm_router_api/docs/ENDPOINT_DEV.md ). All endpoints are exposed under the REST API service. Unless stated otherwise, methods are POST and consume/produce JSON. The default API prefix is /api (configurable via LLM_ROUTER_EP_PREFIX ). Endpoints registered with dont_add_api_prefix=True appear without this prefix (e.g. /models instead of /api/models ). Authentication # When LLM_ROUTER_AUTH_ENABLED=true , endpoints are divided into public and auth‑required : Scope Description Env var Public Bypass all auth checks — always accessible LLM_ROUTER_AUTH_PUBLIC_ENDPOINTS Auth‑required Return 401 Unauthorized if no valid API key is provided LLM_ROUTER_AUTH_ENABLED=true Public endpoints: LLM_ROUTER_AUTH_PUBLIC_ENDPOINTS , default /metrics,/health — and for every entry, /v1{entry} as well. The list is compared against the full request path, so a bare entry never matches a prefixed route: endpoints built with dont_add_api_prefix=False live under LLM_ROUTER_EP_PREFIX ( /api by default) and need the prefixed entry ( LLM_ROUTER_AUTH_PUBLIC_ENDPOINTS=\"/metrics,/health,/api/ping\" ). Everything else requires a valid API key with the appropriate policy permission: Permission type What it grants access to chat Chat completions, model listing, responses embedding Embeddings endpoints anthropic Anthropic Messages API ( /v1/messages ) ollama Ollama‑style chat completion builtin Built‑in utility endpoints (translate, generate, etc.) API keys are checked in order of priority: Authorization: Bearer &lt;key&gt; header x-api-key header Query parameters api_key / api-key are rejected (logged as a warning) — use one of the headers Health &amp; Info # Public by default ( LLM_ROUTER_AUTH_PUBLIC_ENDPOINTS=\"/metrics,/health\" ), reachable without a key whatever LLM_ROUTER_AUTH_ENABLED says: GET /health – Router health check → {\"status\": true, \"body\": \"healthy\", \"stream\": false} . GET /metrics – Prometheus metrics (requires LLM_ROUTER_USE_PROMETHEUS=1 ). Registered straight on the app, so the path is literal /metrics — LLM_ROUTER_EP_PREFIX does not apply. Everything else below needs a key once auth is on; the permission in brackets comes from _ENDPOINT_PERMISSION_MAP : GET /models – List OpenAI‑compatible models ( chat ). GET /v1/models – List OpenAI‑compatible models, v1 ( chat ). GET / – Ollama health endpoint → Ollama is running ( chat ). GET /api/ping – Simple health‑check → {\"status\": true, \"body\": \"pong\", \"stream\": false} ( builtin ). GET /api/version – Return the router version → {\"version\": \"&lt;semver&gt;\", \"stream\": false} ( builtin ); this is the path llm_router_lib clients call. GET /api/tags – List available Ollama model tags ( chat ). GET /api/v0/models – List LM Studio models ( chat ). Auth‑required Endpoints # Chat completions # POST /chat/completions — OpenAI‑style chat completion (requires chat permission). POST /api/chat/completions — OpenAI‑style chat completion with prefix (requires chat permission). POST /v1/chat/completions — vLLM‑like chat completion (requires chat permission). POST /api/chat — Ollama‑style chat completion (requires ollama permission). Responses # POST /responses — OpenAI‑like responses endpoint (requires chat permission). POST /v1/responses — OpenAI‑like responses endpoint v1 (requires chat permission). Embeddings # POST /embeddings — Standard embeddings (requires embedding permission). POST /api/embeddings — Standard embeddings with prefix (requires embedding permission). POST /v1/embeddings — OpenAI‑compatible embeddings endpoint (requires embedding permission). POST /api/embed — Ollama‑native embeddings endpoint (requires embedding permission). Anthropic # POST /v1/messages — Anthropic Messages API compatible endpoint (requires anthropic permission). Chat &amp; Completions (Built‑in, requires builtin permission) # POST /api/conversation_with_model — Standard chat endpoint (OpenAI‑compatible payload). POST /api/extended_conversation_with_m…"},{"k":"llm-router-api/docs/endpoint-dev.html","t":"Endpoint Development Guide","s":"REST API","x":"# Endpoint Development Guide This guide explains how to create, configure and extend REST **endpoints (EP)** in `llm_router_api`. > **Related pages:** endpoint catalog — [`../endpoints/README.md`](../endpoints/README.md) · > models config …","h":["1. Class hierarchy and base variants","EndpointI (abstract base)","EndpointWithHttpRequestI (proxy base)","PassthroughI (passthrough base)","2. Execution cycle – what run_ep does in EndpointWithHttpRequestI","3. System prompt and content overrides – how the fields work","SYSTEM_PROMPT_NAME","How overrides are passed – payload keys (current mechanism)","4. Hooks – _prepare_response_function","5. Parameter validation – REQUIRED_ARGS and OPTIONAL_ARGS","6. Endpoint constructor parameters – what they set","7. How to add a new endpoint – step by step","8. Most common usage patterns","8.1 Simple OpenAI‑compatible proxy","8.2 Endpoint with a built‑in system prompt","8.3 Local / static endpoint","9. Quick field reference","10. Practical notes","Worked example: BatchFileSummaries (per‑file summaries)","See Also"],"b":"Endpoint Development Guide # This guide explains how to create, configure and extend REST endpoints (EP) in llm_router_api . Related pages: endpoint catalog — ../endpoints/README.md · models config — MODELS_CONFIG.md · load balancing — LB_STRATEGIES.md · environment variables — ENV_DEFINITIONS.md 1. Class hierarchy and base variants # All endpoint code lives in llm_router_api/endpoints/endpoint_i.py (plus passthrough.py ). The hierarchy is: SecureEndpointI – security scaffolding (masking, guardrails, audit, metrics) └── EndpointI (ABC) – abstract base: API surface + validation, no run_ep └── EndpointWithHttpRequestI (ABC) – full proxy implementation (run_ep, HTTP, streaming) └── PassthroughI (ABC) – \"forward as-is\" base for OpenAI‑compatible endpoints EndpointI (abstract base) # Defines the general API and argument validation, but does not implement run_ep . Abstract methods that every concrete endpoint must implement: run_ep(params) – execute the endpoint logic for a request. prepare_payload(params) – convert raw request parameters into the payload understood by the downstream backend (or the final response body). Class attributes (defaults): METHODS = [\"GET\", \"POST\"] – supported HTTP verbs. REQUIRED_ARGS = [] – parameter names that must be present (see §5). OPTIONAL_ARGS = [] – accepted but optional parameter names (see §5). SYSTEM_PROMPT_NAME = {\"pl\": None, \"en\": None} – per‑language system‑prompt ids (see §3). Useful helpers: _check_required_params(params) – raises ValueError (→ HTTP 400) when a required key is missing. _resolve_prompt_name(params, map_prompt, prompt_str_force, prompt_str_postfix) – builds the final system prompt (see §3). _get_choices_from_response(response) – parses a requests.Response into (json_body, choices, assistant_text) ; handles both OpenAI‑style ( choices ) and Ollama‑style ( message ) bodies. return_response_ok(body) / return_response_not_ok(body) – standardized JSON response envelopes. The constructor validates the endpoint definition at startup: api_types must be non‑empty and intersect the global API_TYPES list ( \"builtin\" , \"openai\" , \"ollama\" , \"lmstudio\" , \"vllm\" , \"anthropic\" — defined in llm_router_api/core/api_types/dispatcher.py ), otherwise RuntimeError . method must be one of METHODS , otherwise ValueError . EndpointWithHttpRequestI (proxy base) # Extends EndpointI with the full outbound‑HTTP implementation: Complete run_ep implementation (see §2 for the cycle). Outbound HTTP via HttpRequestExecutor ( endpoints/httprequest.py ) and retry orchestration via HttpDispatch ( endpoints/http_dispatch.py ). timeout – seconds after which outbound HTTP calls are aborted. Defaults to EXTERNAL_API_TIMEOUT (300 s, env LLM_ROUTER_EXTERNAL_TIMEOUT ). System‑prompt injection into messages and streaming (NDJSON) support. PassthroughI (passthrough base) # Sets REQUIRED_ARGS = None , OPTIONAL_ARGS = None , SYSTEM_PROMPT_NAME = None . prepare_payload(params) returns params or {} unchanged (decorated with @EP.response_time ), so the request is forwarded verbatim. Useful for OpenAI‑compatible endpoints where you simply forward what arrives. Why do the endpoints in endpoints/builtin/openai.py subclass PassthroughI ? OpenAI‑compatible endpoints need minimal logic – just forward the incoming request. PassthroughI removes the boilerplate (no required arguments, no system prompt, ready‑made proxy run_ep ). The concrete classes ( OpenAICompletionHandler , OpenAIResponsesHandler , OpenAIModelsHandler , …) only add a prepare_response_function that normalizes non‑OpenAI responses (Ollama / Anthropic) into the OpenAI shape. 2. Execution cycle – what run_ep does in EndpointWithHttpRequestI # In short, run_ep(params) performs the following steps (implementation: EndpointWithHttpRequestI.run_ep in endpoints/endpoint_i.py ): Start timer – self._start_time = time.time() . Prepare the payload – params = self._prepare_incoming_payload(params) : calls your prepare_payload(params) (endpoint logic) and then runs the config…"},{"k":"llm-router-cli/index.html","t":"llm-router CLI — Command Reference","s":"CLI","x":"# llm-router CLI — Command Reference **Package:** `llm-router` **Entry points:** - `llm-router` — main CLI tool (auth, anonymizer, config, util, server, completion) --- ## Quick Start ```bash pip install llm-router[api] llm-router --help l…","h":["Quick Start","Top-Level Commands","llm-router auth — API Key &amp; Authentication Management","Command Tree","Shared Flags (store-backed subcommands)","Key Management: llm-router auth key &lt;command&gt;","Policy Management: llm-router auth policy &lt;command&gt;","Rate Limit Overrides: llm-router auth rate-limit &lt;command&gt;","llm-router config — Provider Discovery &amp; Config Merging","Command Tree","discover — Scan hosts for local LLM servers","merge — Merge multiple models-config.json files","llm-router anonymizer run — Text Anonymization","llm-router util — Utility Apps (translate / genai-classifier / genai-data-augmentation)","Command Tree","translate — Translate texts in JSON/JSONL datasets","genai-classifier — Classify translated datasets (JSONL only)","genai-data-augmentation — Augment a local JSONL dataset","llm-router server — REST API Server Lifecycle","Reloading: server reload","Instances — running several servers side by side","list — every instance at a glance","stop --all / status --all / rm-instance","status — colored status card","log — follow the server log","llm-router completion — Shell Tab-Completion (bash / zsh)","Seed File (Memory Store)","Seed File Format (ApiKeyRecord fields)","Key Format","See Also"],"b":"llm-router CLI — Command Reference # Package: llm-router Entry points: llm-router — main CLI tool (auth, anonymizer, config, util, server, completion) Quick Start # pip install llm-router [ api ] llm-router --help llm-router --version Top-Level Commands # Command Description auth Manage API keys, policies, and rate limiting config Auto-discover local providers &amp; merge configs anonymizer run Anonymize text using a selectable algorithm util Utility apps: translate , genai-classifier , genai-data-augmentation server Manage the REST API server ( start / stop / status / list ), many instances completion Generate / install shell tab-completion ( bash / zsh ) llm-router auth — API Key &amp; Authentication Management # Command Tree # llm-router auth key &lt;command&gt; # API key lifecycle llm-router auth policy &lt;command&gt; # Policy management llm-router auth rate-limit &lt;command&gt; # Per-key rate limit overrides Shared Flags (store-backed subcommands) # Flag Default Description --store &lt;backend&gt; memory Key store: memory , redis , or vault --auth-redis-host (empty) Auth Redis host --auth-redis-port 6379 Auth Redis port --auth-redis-db 0 Auth Redis database number --auth-redis-password — Auth Redis password --auth-redis-protocol 2 Auth Redis protocol: 2 (RESP2) or 3 (RESP3) --verbose false Enable verbose (DEBUG) logging of internal operations Note: These flags are shared by all store-backed subcommands — key generate , key list , key delete , key disable , key enable , key rotate , policy create , rate-limit apply , and rate-limit remove . They do not apply to the read-only policy list and rate-limit list subcommands. Note: Each --auth-redis-* flag falls back to a matching LLM_ROUTER_AUTH_REDIS_&lt;HOST|PORT|DB|PASSWORD|PROTOCOL&gt; environment variable before the built-in default is used. These auth-specific Redis flags are separate from the general LLM_ROUTER_REDIS_* env vars. Key Management: llm-router auth key &lt;command&gt; # generate — Create a new API key # llm-router auth key generate \\ --policy developer \\ --expires 1750000000 \\ --store memory Flag Default Description --policy developer Policy name to assign --expires None Expiry (Unix timestamp or None ) --output (stdout) Output file path (created with 0600 permissions) Output: sk-llmr-live-&lt;base62&gt; key (plaintext shown once at creation) plus the generated Key ID: (e.g. key-fe8fc388 ) — use the ID with list / delete / disable / enable / rotate . list — List all API keys # llm-router auth key list --store memory [ --json ] Flag Default Description --json false Output in JSON format delete &lt;key-id&gt; — Delete a key permanently # llm-router auth key delete &lt;key-id&gt; --store memory disable &lt;key-id&gt; — Deactivate without deleting # llm-router auth key disable &lt;key-id&gt; [ --store memory ] enable &lt;key-id&gt; — Re-activate a disabled key # llm-router auth key enable &lt;key-id&gt; [ --store memory ] rotate &lt;key-id&gt; — Generate a replacement key # llm-router auth key rotate &lt;key-id&gt; --grace 3600 [ --store memory ] Flag Default Description --grace 3600 Grace period in seconds (old key stays valid) Policy Management: llm-router auth policy &lt;command&gt; # list — List builtin policies # llm-router auth policy list Builtin policies: Policy Access Description developer All Full access to all endpoints admin All Admin access chat Chat Chat completion endpoints embedding Embedding Embedding endpoints anthropic Anthropic Anthropic messages endpoint ollama Ollama Ollama endpoints builtin All builtin Built-in endpoints (translate, etc.) create &lt;name&gt; [&lt;json-policy&gt;] — Create a custom policy # # inline JSON llm-router auth policy create my-team '{ \"can_access\": true, \"rate_limit\": 120, \"model_whitelist\": [\"gpt-4\", \"llama-3\"] }' --store memory # from a file, or from stdin (-) — avoids leaking policy JSON into shell history llm-router auth policy create my-team --file my-team.json cat my-team.json | llm-router auth policy creat…"},{"k":"llm-router-lib/index.html","t":"llmrouterlib","s":"Python library","x":"# llm_router_lib ## Overview `llm_router_lib` bundles **Pydantic data‑model definitions** **and** a **thin, opinionated client wrapper** for the LLM‑Router service. - The **data models** live in `llm_router_lib/data_models` and describe ev…","h":["Overview","Installation","Quick start","Data models","Conversation models","Utility models (selected examples)","Services (low‑level wrappers)","Example: using a service directly","Thin client wrapper (LLMRouterClient)","Async client (AsyncLLMRouterClient)","Streaming","Utilities","Development &amp; testing"],"b":"llm_router_lib # Overview # llm_router_lib bundles Pydantic data‑model definitions and a thin, opinionated client wrapper for the LLM‑Router service. The data models live in llm_router_lib/data_models and describe every request payload the router accepts. The client ( LLMRouterClient ) offers a high‑level, Pythonic API that hides HTTP details, retries, and error handling. The async client ( AsyncLLMRouterClient , httpx ‑based) offers the same API for asyncio applications (e.g. FastAPI), plus streaming conversation methods. Low‑level service classes ( ConversationWithModelService , ExtendedConversationWithModelService , TranslateService , GenerativeAnswerService , health services) perform the actual HTTP calls and can be used directly when finer‑grained control is required. HttpRequester (in utils/http.py ) is a small wrapper around requests that adds logging, configurable retries, and unified error translation. AsyncHttpRequester (in utils/http_async.py ) is its httpx ‑based asynchronous counterpart with the same retry and error‑translation contract (plus a stream() helper for SSE responses). A dedicated exception hierarchy ( exceptions.py ) maps HTTP errors to meaningful Python exceptions. In short, llm_router_lib provides both the contract (the “schema”) and a convenient client to consume the router service. Installation # The library targets Python 3.10.6 and uses a virtualenv . Install it in editable mode for development: # Clone the repository (if you haven't already) git clone https://github.com/radlab-dev-group/llm-router.git cd llm-router/llm_router_lib # Create and activate a virtual environment python3 -m venv .venv source .venv/bin/activate # Install the package and its dependencies pip install -e . All runtime dependencies ( requests , pydantic , plus the packages listed in requirements.txt ) are declared in the project’s requirements.txt . Quick start # from llm_router_lib import LLMRouterClient # Initialise the client – point it at the router’s host (do **not** include the `/api` prefix) client = LLMRouterClient ( api = \"http://localhost:8080\" , # router host URL token = \"YOUR_ROUTER_TOKEN\" , # optional, if router requires auth ) # Call the standard conversation endpoint with named keyword arguments – # the client builds the Pydantic request model for you. response = client . conversation_with_model ( user_last_statement = \"Hello, how are you?\" , model = \"google/gemma-3-12b-it\" , temperature = 0.7 , max_new_tokens = 128 , ) # response is a typed `ConversationResponse` model: print ( response . response ) # → the assistant's reply text print ( response . generation_time ) # → seconds taken by the server You can also pass a ready‑made pydantic request model via the payload keyword (all endpoint parameters are keyword‑only – there are no positional arguments): from llm_router_lib.data_models.builtin_chat import ConversationWithModelRequest payload = ConversationWithModelRequest ( model_name = \"google/gemma-3-12b-it\" , user_last_statement = \"Hello, how are you?\" , temperature = 0.7 , max_new_tokens = 128 , ) response = client . conversation_with_model ( payload = payload ) Note: raw dict payloads are no longer accepted – passing a dict as payload raises TypeError . Build the matching Pydantic model explicitly (e.g. ConversationWithModelRequest(**dict_payload) ) or use the named keyword arguments. Calling a generation method with neither payload nor enough named arguments raises NoArgsAndNoPayloadError . Data models # All request payloads are defined in llm_router_lib/data_models . A common base class supplies shared options: class BaseModelOptions ( BaseModel ): \"\"\"Options shared across many endpoint models.\"\"\" mask_payload : bool = False masker_pipeline : Optional [ List [ str ]] = None Conversation models # Model Required fields Optional / extra fields ConversationWithModelRequest model_name , user_last_statement temperature , max_new_tokens , historical_messages , … ExtendedConversationWithModelRequest All of the…"},{"k":"llm-router-lib/response-models.html","t":"Response models","s":"Python library","x":"# Response models ## Introduction The router speaks JSON, and every endpoint emits a *different* shape. Before these models, `LLMRouterClient` handed that JSON straight back as a raw `dict`, so the knowledge of each response schema lived o…","h":["Introduction","📦 Where the models live and how to import them","🗺️ Client method → response model","🧱 Class hierarchy","📚 Model reference","Health / meta","Conversation","Per‑text (list) utilities","Single‑output utilities","Article utilities","🎨 Design decisions","🚀 Usage examples","Access fields (typed)","Serialize / export","Round‑trip (validate a stored/status‑wrapped body)","🔁 Backward compatibility","➕ Adding a new endpoint model","✅ Testing notes","🧾 File map"],"b":"Response models # Introduction # The router speaks JSON, and every endpoint emits a different shape. Before these models, LLMRouterClient handed that JSON straight back as a raw dict , so the knowledge of each response schema lived on the caller side: you had to remember the exact key names, guard against missing fields, and you got no compile‑time or IDE help. The response models move that contract into the library. Every client method returns a typed Pydantic model that mirrors the JSON body the router emits for that endpoint, and the models are defined in one place — data_models/response.py — next to the existing request models ( TranslateModel , Polarity3cModel , …). Together, request response describe the full contract of an endpoint. Concretely, the typing is there to do three things: Fix the contract in one place. response.py is the single definition of what each endpoint returns. Change the router → change the model → type checkers and tests flag every consumer that still relied on the old shape. Validate at the boundary. Pydantic checks the body as it comes in, so a malformed or partial response fails there with a ValidationError instead of as a KeyError a few frames deep in your code. Expose static types. Because the return annotation is a class rather than Dict[str, Any] , mypy and IDEs can resolve resp.response[0].translated . The model also exports a JSON Schema via model_json_schema() for docs, codegen, or contract tests. # resp is a typed model, not a dict resp = client . translate ( texts = [ \"Hello world\" ], model = \"speakleash/Bielik-11B-v2.3-Instruct\" ) resp . response [ 0 ] . translated # str resp . generation_time # Optional[float] resp . model_dump () # → plain dict, if you still need one 📦 Where the models live and how to import them # All response models are exported from the package root, so either import style works: # Option 1 – from the package (recommended) from llm_router_lib.data_models import ( TranslateResponse , Polarity3cResponse , ConversationResponse , ModelsListResponse , # …see the full list in data_models/__init__.py ) # Option 2 – from the module directly from llm_router_lib.data_models.response import TranslateResponse They sit alongside (but are independent of) the request models such as TranslateModel , Polarity3cModel , etc. 🗺️ Client method → response model # The table below maps every LLMRouterClient method to the model it now returns. response is the endpoint‑specific payload; generation_time (seconds) is present on all generation endpoints. Client method Endpoint Returns response payload ping() GET /api/ping PingResponse status: bool , body: str version() GET /api/version VersionResponse version: str models() GET /v1/models ModelsListResponse object: str , data: List[ModelInfo] conversation_with_model(payload) POST /api/conversation_with_model ConversationResponse str (assistant reply) extended_conversation_with_model(payload) POST /api/extended_conversation_with_model ExtendedConversationResponse str (assistant reply) polarity_3c(...) POST /api/polarity_3c Polarity3cResponse List[Polarity3cItem] translate(...) POST /api/translate TranslateResponse List[TranslateItem] simplify_text(...) POST /api/simplify_text SimplifyTextResponse List[str] generative_answer(...) POST /api/generative_answer GenerativeAnswerResponse str (answer) generate_article_from_text(...) POST /api/generate_article_from_text GenerateArticleFromTextResponse ArticleText generate_article_from_texts(...) POST /api/generate_article_from_texts GenerateArticleFromTextsResponse ArticleText create_full_article_from_texts(...) POST /api/create_full_article_from_texts CreateFullArticleFromTextsResponse ArticleText generate_questions(...) POST /api/generate_questions GenerateQuestionsResponse List[TextQuestions] generate_label(...) POST /api/generate_label GenerateLabelResponse str (label) 🧱 Class hierarchy # BaseResponse # extra keys ignored (pydantic default) ├── PingResponse ├── VersionResponse ├── ModelInfo ├── Mod…"},{"k":"helm-charts/index.html","t":"LLM‑Router Helm Chart","s":"Deployment","x":"# LLM‑Router Helm Chart Quick summary – This Helm chart deploys the LLM‑Router application with its optional Redis dependency. It works on any Kubernetes 1.19+ cluster and can be customized via values.yaml, environment variables, or the --…","h":["Table of Contents","Prerequisites","Chart layout","Dependencies","Installing the chart","Customising the deployment","1️⃣ Using a custom values file"],"b":"LLM‑Router Helm Chart # Quick summary – This Helm chart deploys the LLM‑Router application with its optional Redis dependency. It works on any Kubernetes 1.19+ cluster and can be customized via values.yaml, environment variables, or the --set flag. Table of Contents # Prerequisites Chart layout Dependencies Installing the chart Customising the deployment Prerequisites # Requirement Why we need it Kubernetes (v1.19 or newer) The chart creates Deployments, Services, Ingresses, ConfigMaps, … Helm (v3.x) Used to render and apply the chart Access to a container registry (e.g., Docker Hub, Quay, your private registry) The chart pulls the image defined in llm-router``values.yaml (Optional) cert‑manager If you enable TLS for the Ingress, cert‑manager will provision certificates Chart layout # helm_charts/ └─ llm-router/ ├─ Chart.yaml # Chart metadata ├─ values.yaml # Default values ├─ values-dev.yaml # Development‑specific overrides ├─ templates/ │ ├─ _helpers.tpl # Helper functions (name, labels, etc.) │ ├─ deployment.yaml # Deployment definition │ ├─ service.yaml # Service definition │ ├─ ingress.yaml # Ingress definition (optional) │ ├─ configmap.yaml # ConfigMap for runtime env vars │ └─ configmap-models.yaml # ConfigMap for `models-config.json` └─ charts/ └─ redis-23.2.12.tgz # Bitnami Redis sub‑chart (dependency) Dependencies # The chart depends on Redis. The dependency is declared in Chart.yaml and pulled automatically when you run helm dependency update. # Chart.yaml dependencies : - name : redis version : \"23.2.12\" repository : \"oci://registry-1.docker.io/bitnamicharts\" condition : redis.enabled Adding / updating dependencies # From the chart root (helm_charts/llm-router) helm dependency update . You can disable the Redis sub‑chart with: --set redis.enabled = false Installing the chart # helm upgrade --install my-llm-router ./helm_charts/llm-router \\ --namespace my-namespace \\ --create-namespace Customising the deployment # You can tailor the LLM‑Router Helm chart to your environment in three different ways. Pick the approach that best fits the task at hand. Method When to use Edit values.yaml and run helm upgrade Long‑term, version‑controlled configuration that lives in source control. Pass --set flags on the command line One‑off tweaks, CI pipelines, quick experiments, or when you need to override just a handful of values. Provide a custom values file ( -f my-values.yaml ) Complex overrides, reusable profiles, or when you prefer to keep the changes in a separate, shareable file. Below are concrete examples for the two most common scenarios: using a custom values file and using --set flags. 1️⃣ Using a custom values file # Create a file (e.g. my-values.yaml ) with the settings you want to override: # my-values.yaml ingress : enabled : true className : traefik hosts : - host : my-custom.cluster.local paths : - path : / pathType : Prefix tls : - hosts : - my-custom.cluster.local secretName : llm-router-tls image : tag : latest # pull the latest container image web_host : my-custom.cluster.local # the hostname used by the app &amp; Ingress Deploy (or upgrade) the chart with this file: ```shell script helm upgrade --install llm-router ./helm_charts/llm-router \\ -f my-values.yaml --- ### 2️⃣ Using `--set` flags If you only need to change a few values, the `--set` syntax is handy: ```shell script helm upgrade --install llm-router ./helm_charts/llm-router \\ --set ingress.enabled=true \\ --set web_host=llm2.k3s.radlab.dev \\ --set image.tag=latest Flag Effect ingress.enabled=true Turns on the Ingress resource (otherwise it’s omitted). llm-1.cluster.local Sets the host name used both in the Ingress rule and the application’s configuration ( LLM_ROUTER_WEB_HOST ). image.tag=latest Pulls the latest image tag instead of the default chart‑version tag."},{"k":"examples/index.html","t":"Integration Examples with LLM Router","s":"Examples","x":"# Integration Examples with LLM Router This directory contains example boilerplates that demonstrate how easy it is to integrate popular LLM libraries with the router by simply switching the host. --- ## Available Examples - **[LlamaIndex]…","h":["Available Examples","Core Principle","Quick Start","Example Structure","Full Stack with Local Models","Additional Information"],"b":"Integration Examples with LLM Router # This directory contains example boilerplates that demonstrate how easy it is to integrate popular LLM libraries with the router by simply switching the host. Available Examples # LlamaIndex – Integration with LlamaIndex (GPT Index) Using LlamaIndex with Local Models – Guide for mapping OpenAI model names to local models via the router. LangChain – Integration with LangChain OpenAI SDK – Direct integration with the OpenAI Python SDK LiteLLM – Integration with LiteLLM Haystack – Integration with Haystack Core Principle # All examples work on the same principle: just change base_url / api_base to the address of your router , and the router will automatically: ✅ Distribute traffic among available providers ✅ Perform load balancing ✅ Provide health checking ✅ Supply monitoring and metrics ✅ Handle streaming and non‑streaming responses Quick Start # Each example can be run directly: ```shell script LlamaIndex # python examples/llamaindex_example.py LangChain # python examples/langchain_example.py OpenAI SDK # python examples/openai_example.py LiteLLM # python examples/litellm_example.py Haystack # python examples/haystack_example.py ``` Example Structure # Each example includes: Basic configuration – how to point the library at the router Streaming – handling streaming responses Non‑streaming – handling full responses Error handling – managing errors Full Stack with Local Models # The quick‑start guides for running the full stack with local models are included in the repository: Gemma 3 12B‑IT – README Bielik 11B‑v2.3‑Instruct – README These guides walk you through: Installing vLLM and the respective model. Setting up LLM‑Router with the provided models-config.json . Testing the end‑to‑end flow (router → vLLM). Follow the linked README files for step‑by‑step instructions to launch a complete stack locally. Additional Information # Learn more about the router: Main README API Documentation Endpoints Overview Load‑Balancing Strategies"},{"k":"examples/quickstart/index.html","t":"Full Stack with Local Models","s":"Examples","x":"## Full Stack with Local Models The quick‑start guides for running the full stack with **local models** are included in the repository: - **Gemma 3 12B‑IT** – [README](google-gemma3-12b-it/README.md) - **Bielik 11B‑v2.3‑Instruct** – [READM…","h":["Full Stack with Local Models"],"b":"Full Stack with Local Models # The quick‑start guides for running the full stack with local models are included in the repository: Gemma 3 12B‑IT – README Bielik 11B‑v2.3‑Instruct – README"},{"k":"examples/readme-llamaindex.html","t":"Using LlamaIndex with Local Models via the LLM‑Router","s":"Examples","x":"## Using LlamaIndex with Local Models via the LLM‑Router LlamaIndex’s `OpenAI` wrapper expects **OpenAI‑style model names** (e.g. `gpt-3.5-turbo`, `gpt-4`). When you want to run *local* models (Gemma, Ollama, vLLM, etc.) behind an LLM‑Rout…","h":["Using LlamaIndex with Local Models via the LLM‑Router","Why the mapping is required","Router configuration","How it works end‑to‑end","Checklist for a working setup","TL;DR"],"b":"Using LlamaIndex with Local Models via the LLM‑Router # LlamaIndex’s OpenAI wrapper expects OpenAI‑style model names (e.g. gpt-3.5-turbo , gpt-4 ). When you want to run local models (Gemma, Ollama, vLLM, etc.) behind an LLM‑Router, the router must translate those OpenAI names to the actual model identifiers used by the local providers. Why the mapping is required # LlamaIndex validates the model name against a known list of OpenAI models to decide whether the call is a chat request, to obtain context‑window size, etc. If the wrapper receives a name it does not recognise (e.g. google/gemma-3-12b-it ), it raises a ValueError like: ValueError: Unknown model 'google/gemma-3-12b-it'. Please provide a valid OpenAI model name … By keeping the OpenAI name in the request and letting the router forward the request to the appropriate backend, you get the best of both worlds: LlamaIndex works unchanged, and the traffic is routed to your local model. Router configuration # The router’s configuration must contain a section that lists OpenAI model names ( openai_models ). Each entry defines one or more providers that actually serve the model. The key points are: Field Meaning id Arbitrary identifier for the provider instance. api_host URL where the provider’s inference server is reachable. api_type The protocol the provider uses ( vllm , ollama , …). model_path The local model identifier (e.g. google/gemma-3-12b-it ). weight Relative load‑balancing weight when several providers are listed. Example snippet # \"openai_models\" : { (...) \"gpt-3.5-turbo\" : { \"providers\" : [ { \"id\" : \"gpt_35_turbo-gemma3_12b-vllm-71:7000\" , \"api_host\" : \"http://192.168.100.71:7000/\" , \"api_token\" : \"\" , \"api_type\" : \"vllm\" , \"input_size\" : 4096 , \"model_path\" : \"google/gemma-3-12b-it\" , \"weight\" : 1.0 }, { \"id\" : \"gpt_35_turbo-gemma3_12b-vllm-71:7001\" , \"api_host\" : \"http://192.168.100.71:7001/\" , \"api_token\" : \"\" , \"api_type\" : \"vllm\" , \"input_size\" : 4096 , \"model_path\" : \"google/gemma-3-12b-it\" , \"weight\" : 1.0 } ] }, \"gpt-4\" : { \"providers\" : [ { \"id\" : \"gpt-4-gpt-oss-20b-ollama-66:11434\" , \"api_host\" : \"http://192.168.100.66:11434\" , \"api_token\" : \"\" , \"api_type\" : \"ollama\" , \"input_size\" : 256000 , \"model_path\" : \"gpt-oss:120b\" } ] } }, ... \"active_models\" : { \"openai_models\" : [ \"gpt-3.5-turbo\" , \"gpt-4\" ... ] } How it works end‑to‑end # LlamaIndex code creates an OpenAI client with a model name that the wrapper knows, e.g.: llm = OpenAI ( model = \"gpt-3.5-turbo\" , api_base = \"http://localhost:8080\" , api_key = \"not-needed\" ) The request is sent to the router ( api_base ). The router looks up \"gpt-3.5-turbo\" in openai_models , selects a provider, and forwards the request to the provider’s api_host . The provider receives the request with model_path=\"google/gemma-3-12b-it\" (or \"gpt-oss:120b\" for gpt-4 ) and runs the local model. The response travels back through the router to LlamaIndex, which treats it exactly like an OpenAI response. Checklist for a working setup # Router is running and reachable at the URL you pass to api_base . openai_models section contains all OpenAI names you intend to use from LlamaIndex. Each OpenAI name maps to at least one provider with the correct model_path . The active_models.openai_models list includes the names you want to expose (otherwise the router will ignore them). Local model servers (vLLM, Ollama, etc.) are up and listening on the api_host URLs specified. TL;DR # *Keep the model names you give to LlamaIndex identical to the OpenAI names defined in the router configuration. The router will translate those names to the real local model identifiers ( model_path ). This mapping is the only thing needed for LlamaIndex to work seamlessly with any self�"},{"k":"examples/quickstart/google-gemma3-12b-it/vllm.html","t":"vLLM + google/gemma‑3‑12b‑it – Quick‑Start Guide (Ubuntu)","s":"Examples","x":"# vLLM + `google/gemma‑3‑12b‑it` – Quick‑Start Guide (Ubuntu) > **Prerequisites** > - Ubuntu 20.04 or newer > - Python 3.10 (our project uses 3.10.6) > - `virtualenv` (installed) > - CUDA 11.8 + GPU **or** a CPU‑only setup --- ## 1️⃣ Creat…","h":["1️⃣ Create &amp; activate a virtual environment","2️⃣ Install vLLM","Verify the installation","3️⃣ Obtain the model google/gemma-3-12b-it","4️⃣ Run the vLLM server","5️⃣ Test the endpoint","6️⃣ Handy tips","🎉 All set!"],"b":"vLLM + google/gemma‑3‑12b‑it – Quick‑Start Guide (Ubuntu) # Prerequisites - Ubuntu 20.04 or newer - Python 3.10 (our project uses 3.10.6) - virtualenv (installed) - CUDA 11.8 + GPU or a CPU‑only setup 1️⃣ Create &amp; activate a virtual environment # mkdir -p ~/vllm-gemma &amp;&amp; cd ~/vllm-gemma python3 -m venv .venv source .venv/bin/activate This creates an optional project directory, sets up a Python virtual environment in .venv , and activates it (you’ll see (.venv) in the prompt). 2️⃣ Install vLLM # pip install --upgrade pip pip install \"vllm[cuda]\" The above installs the latest pip and then installs vLLM with GPU support (the appropriate CUDA libraries are detected automatically). If you have no GPU, install the CPU‑only version instead: pip install vllm[cpu] . Verify the installation # python -c \"import vllm; print(vllm.__version__)\" You should see a version string such as 0.11.2 . 3️⃣ Obtain the model google/gemma-3-12b-it # mkdir -p ./google/gemma-3-12b-it pip install huggingface_hub hf download google/gemma-3-12b-it \\ --local-dir ./google/gemma-3-12b-it This creates a folder for the model, installs the Hugging Face CLI, and downloads the model files into ./google/gemma-3-12b-it . The files are cached under ~/.cache/huggingface/hub by default; you can keep a local copy to avoid re‑downloads. 4️⃣ Run the vLLM server # Copy the ready‑to‑use Bash script ( llm-router/examples/quickstart/google-gemma3-12b-it/run-gemma-3-12b-it-vllm.sh ) to directory wgen the vLLM will be started with the Gemma 3 model. cp path/to/llm-router/examples/quickstart/google-gemma3-12b-it/run-gemma-3-12b-it-vllm.sh . bash run-gemma-3-12b-it-vllm.sh Tip: Run the server inside a tmux or screen session so it stays alive even if you disconnect from the terminal. 5️⃣ Test the endpoint # INFO : curl and jq are system utilities. curl http://localhost:7000/v1/chat/completions \\ -H \"Content-Type: application/json\" \\ -d '{ \"model\": \"google/gemma-3-12b-it\", \"messages\": [{\"role\": \"user\", \"content\": \"Hello, how are you?\"}], \"max_tokens\": 100 }' | jq You should receive a JSON response containing the model’s generated text, for example: { \"id\" : \"chatcmpl-e30bed0db9f9440a8aec14bd287ca63d\" , \"object\" : \"chat.completion\" , \"created\" : 1764516430 , \"model\" : \"google/gemma-3-12b-it\" , \"choices\" : [ { \"index\" : 0 , \"message\" : { \"role\" : \"assistant\" , \"content\" : \"Hello! I'm doing well, thank you for asking! As an AI, I don't experience feelings like humans do, but everything is running smoothly and I'm ready to chat. 😊\\n\\nHow are *you* doing today?\" , \"refusal\" : null , \"annotations\" : null , \"audio\" : null , \"function_call\" : null , \"tool_calls\" : [], \"reasoning\" : null , \"reasoning_content\" : null }, \"logprobs\" : null , \"finish_reason\" : \"stop\" , \"stop_reason\" : 106 , \"token_ids\" : null } ], \"service_tier\" : null , \"system_fingerprint\" : null , \"usage\" : { \"prompt_tokens\" : 15 , \"total_tokens\" : 66 , \"completion_tokens\" : 51 , \"prompt_tokens_details\" : null }, \"prompt_logprobs\" : null , \"prompt_token_ids\" : null , \"kv_transfer_params\" : null } 6️⃣ Handy tips # Topic Recommendation Memory google/gemma‑3‑12b‑it needs ~24GB VRAM. Use --cpu-offload (if supported) for larger models or when GPU memory is limited. Cache location Set HF_HOME=$PWD/.cache/huggingface to keep all model files inside the project directory. Parallelism Export TOKENIZERS_PARALLELISM=false to silence tokenizer warnings. GPU selection export CUDA_VISIBLE_DEVICES=0 (or another index) when multiple GPUs are present. Update pip install -U vllm refreshes the library; the next server start will pull newer model files if available. Deactivate When done, simply run deactivate to leave the virtual environment. 🎉 All set! # You now have a fully functional OpenAI‑compatible API powered by vLLM and the google/gemma‑3‑12b‑it model."},{"k":"examples/quickstart/speakleash-bielik-11b-v2-3-instruct/vllm.html","t":"vLLM + speakleash/Bielik-11B-v2.3-Instruct – Przewodnik Szybkiego Startu (Ubuntu)","s":"Examples","x":"# vLLM + `speakleash/Bielik-11B-v2.3-Instruct` – Przewodnik Szybkiego Startu (Ubuntu) > **Wymagania wstępne** > - Ubuntu 20.04 lub nowszy > - Python 3.10 (w projekcie używamy 3.10.6) > - `virtualenv` (zainstalowany) > - CUDA 11.8 + GPU **l…","h":["1️⃣ Utwórz i aktywuj wirtualne środowisko","2️⃣ Zainstaluj vLLM","Sprawdź instalację","4️⃣ Przygotuj środowisko do pobierania modelu","6️⃣ Pobierz model speakleash/Bielik-11B-v2.3-Instruct","(Opcjonalnie) Ustaw własny katalog cache","7️⃣ Uruchom serwer vLLM","8️⃣ Przetestuj endpoint","9️⃣ Przydatne wskazówki","🎉 Gotowe!"],"b":"vLLM + speakleash/Bielik-11B-v2.3-Instruct – Przewodnik Szybkiego Startu (Ubuntu) # Wymagania wstępne - Ubuntu 20.04 lub nowszy - Python 3.10 (w projekcie używamy 3.10.6) - virtualenv (zainstalowany) - CUDA 11.8 + GPU lub środowisko tylko CPU 1️⃣ Utwórz i aktywuj wirtualne środowisko # mkdir -p ~/bielik &amp;&amp; cd ~/bielik python3 -m venv .venv source .venv/bin/activate Powoduje to utworzenie katalogu projektu, przygotowanie wirtualnego środowiska w folderze .venv oraz jego aktywację (w promptcie pojawi się (.venv) ). 2️⃣ Zainstaluj vLLM # pip install --upgrade pip pip install \"vllm[cuda]\" Instalacja najnowszej wersji vLLM z obsługą GPU (CUDA zostanie wykryte automatycznie). Jeśli nie masz GPU, użyj wersji CPU: pip install vllm[cpu] . Sprawdź instalację # python -c \"import vllm; print(vllm.__version__)\" Powinieneś zobaczyć wersję, np. 0.11.2 . 4️⃣ Przygotuj środowisko do pobierania modelu # pip install huggingface_hub 6️⃣ Pobierz model speakleash/Bielik-11B-v2.3-Instruct # mkdir -p ./speakleash/Bielik-11B-v2.3-Instruct hf download speakleash/Bielik-11B-v2.3-Instruct \\ --local-dir ./speakleash/Bielik-11B-v2.3-Instruct Model zostanie pobrany do wskazanego katalogu. Pliki będą także buforowane domyślnie w ~/.cache/huggingface/hub . (Opcjonalnie) Ustaw własny katalog cache # Jeśli chcesz, aby wszystkie modele były przechowywane wewnątrz projektu, ustaw zmienną przed pobraniem: export HF_HOME=$PWD/.cache/huggingface # np. ./bielik/.cache/huggingface 7️⃣ Uruchom serwer vLLM # Skopiuj gotowy skrypt Bash (przykładowa ścieżka – dostosuj do swojego projektu): cp path/to/llm-router/examples/quickstart/speakleash-bielik-11b-v2_3-Instruct/run-bielik-11b-v2_3-vllm.sh . bash run-bielik-11b-v2_3-vllm.sh Wskazówka: uruchom serwer w sesji tmux lub screen , aby pozostawał aktywny po rozłączeniu się z terminalem. 8️⃣ Przetestuj endpoint # INFO : curl i jq to narzędzia systemowe. curl http://localhost:7000/v1/chat/completions \\ -H \"Content-Type: application/json\" \\ -d '{ \"model\": \"speakleash/Bielik-11B-v2.3-Instruct\", \"messages\": [{\"role\": \"user\", \"content\": \"Cześć, jak się masz?\"}], \"max_tokens\": 100 }' | jq Powinieneś otrzymać odpowiedź w formacie JSON, np.: { \"id\" : \"chatcmpl-xxxx\" , \"object\" : \"chat.completion\" , \"created\" : 1764516430 , \"model\" : \"speakleash/Bielik-11B-v2.3-Instruct\" , \"choices\" : [ { \"index\" : 0 , \"message\" : { \"role\" : \"assistant\" , \"content\" : \"Cześć! Jestem w pełni sprawny i gotowy do rozmowy. Jak mogę Ci pomóc?\" }, \"finish_reason\" : \"stop\" } ], \"usage\" : { \"prompt_tokens\" : 15 , \"total_tokens\" : 66 , \"completion_tokens\" : 51 } } 9️⃣ Przydatne wskazówki # Temat Rekomendacja Pamięć speakleash/Bielik-11B-v2.3-Instruct potrzebuje ok. 24GB VRAM. Użyj --cpu-offload (jeśli wspierane) przy ograniczonej pamięci GPU. Lokalizacja cache Ustaw HF_HOME=$PWD/.cache/huggingface , aby wszystkie pliki modelu znajdowały się w katalogu projektu. Równoległość tokenizera export TOKENIZERS_PARALLELISM=false wyciszy ostrzeżenia tokenizera. Wybór GPU export CUDA_VISIBLE_DEVICES=0 (lub inny indeks) przy wielu kartach GPU. Aktualizacja pip install -U vllm odświeża bibliotekę; przy następnym uruchomieniu serwera zostaną pobrane nowsze pliki modelu, jeśli są dostępne. Dezaktywacja Po zakończeniu pracy wystarczy wpisać deactivate , aby opuścić wirtualne środowisko. 🎉 Gotowe! # Masz już w pełni działające API kompatybilne z OpenAI, oparte na vLLM i modelu speakleash/Bielik-11B-v2.3-Instruct ."},{"k":"examples/quickstart/speakleash-bielik-11b-v2-3-instruct/index.html","t":"🚀 Przewodnik Szybkiego Startu dla speakleash/Bielik-11B-v2.3-Instruct z vLLM & LLM‑Router","s":"Examples","x":"# 🚀 **Przewodnik Szybkiego Startu** dla `speakleash/Bielik-11B-v2.3-Instruct` z **vLLM** & **LLM‑Router** Ten przewodnik prowadzi Cię krok po kroku przez: 1. **Instalację vLLM** i modelu `speakleash/Bielik-11B-v2.3-Instruct`. 2. **Instalac…","h":["📋 Wymagania wstępne","1️⃣ Utworzenie i aktywacja wirtualnego środowiska","6️⃣ Przygotowanie konfiguracji routera","7️⃣ Test pełnego stosu (router → vLLM)","🎉 Co dalej?"],"b":"🚀 Przewodnik Szybkiego Startu dla speakleash/Bielik-11B-v2.3-Instruct z vLLM &amp; LLM‑Router # Ten przewodnik prowadzi Cię krok po kroku przez: Instalację vLLM i modelu speakleash/Bielik-11B-v2.3-Instruct . Instalację LLM‑Router (bramki API). Uruchomienie routera z konfiguracją modeli dostarczoną w models-config.json . Wszystkie polecenia zakładają, że pracujesz na systemie Unix‑like (Linux/macOS) z Python 3.10.6 , virtualenv oraz ( opcjonalnie) kartą GPU obsługującą CUDA 11.8. 📋 Wymagania wstępne # Wymaganie Szczegóły OS Ubuntu 20.04 + (lub dowolna nowsza dystrybucja Linux/macOS) Python 3.10.6 (domyślna wersja projektu) GPU CUDA 11.8 + (minimum 12 GB VRAM) lub środowisko CPU‑only Narzędzia git , curl , jq (opcjonalnie, przydatne do testowania) Sieć Dostęp do PyPI oraz Hugging Face w celu pobrania modelu 1️⃣ Utworzenie i aktywacja wirtualnego środowiska # ```shell script (opcjonalnie) utwórz katalog demo i przejdź do niego # mkdir -p ~/bielik-demo &amp;&amp; cd $_ Inicjalizacja venv # python3 -m venv .venv source .venv/bin/activate Aktualizacja pip (zawsze dobry pomysł) # pip install --upgrade pip --- ## 2️⃣ Instalacja **vLLM** oraz pobranie modelu Bielik &gt; Pełną instrukcję znajdziesz w pliku [`VLLM.md`](./VLLM.md). --- ## 3️⃣ **Uruchomienie serwera vLLM** Skopiuj do bieżącego katalogu dostarczony skrypt Bash (dostosuj ścieżkę, jeśli potrzebujesz) i uruchom go: ```shell script cp path/to/llm-router/examples/quickstart/speakleash-bielik-11b-v2_3-Instruct/run-bielik-11b-v2_3-vllm.sh . chmod +x run-bielik-11b-v2_3-vllm.sh # Uruchom (warto w tmux/screen) ./run-bielik-11b-v2_3-vllm.sh Serwer nasłuchuje na http://0.0.0.0:7000 i udostępnia endpoint zgodny z OpenAI pod /v1/chat/completions . Możesz szybko go przetestować: ```shell script curl http://localhost:7000/v1/chat/completions \\ -H \"Content-Type: application/json\" \\ -d '{ \"model\": \"speakleash/Bielik-11B-v2.3-Instruct\", \"messages\": [{\"role\": \"user\", \"content\": \"Cześć, jak się masz?\"}], \"max_tokens\": 100 }' | jq Powinieneś otrzymać odpowiedź w formacie JSON. --- ## 4️⃣ Instalacja **LLM‑Router** ```shell script # Sklonuj repozytorium (jeśli jeszcze go nie masz) git clone https://github.com/radlab-dev-group/llm-router.git cd llm-router # Instalacja core + API (w tym samym venv) pip install .[api] # (Opcjonalnie) wsparcie dla Prometheus pip install .[api,metrics] Uwaga: Router używa tego samego wirtualnego środowiska, które utworzyłeś wcześniej, więc wszystkie zależności pozostają odizolowane. 6️⃣ Przygotowanie konfiguracji routera # Plik models-config.json znajdujący się w katalogu speakleash‑bielik już zawiera definicję naszego modelu: { \"speakleash_models\" : { \"speakleash/Bielik-11B-v2.3-Instruct\" : { \"providers\" : [ { \"id\" : \"bielik-11B_v2_3-vllm-local:7000\" , \"api_host\" : \"http://localhost:7000/\" , \"api_type\" : \"vllm\" , \"input_size\" : 56000 , \"weight\" : 1.0 } ] } }, \"active_models\" : { \"speakleash_models\" : [ \"speakleash/Bielik-11B-v2.3-Instruct\" ] } } Skopiuj go (lub przenieś) do katalogu resources/configs/ routera: ```shell script mkdir -p resources/configs cp path/to/speakleash-bielik/models-config.json resources/configs/ --- ## 6️⃣ Uruchomienie **LLM‑Router** ### Lokalny Gunicorn W repozytorium znajduje się pomocniczy skrypt `run-rest-api-gunicorn.sh`. Upewnij się, że jest wykonywalny, a następnie go uruchom: ```shell script chmod +x run-rest-api-gunicorn.sh ./run-rest-api-gunicorn.sh Domyślne zmienne środowiskowe (można zmienić w skrypcie): Zmienna Domyślna wartość Opis LLM_ROUTER_SERVER_TYPE gunicorn Backend serwera LLM_ROUTER_SERVER_PORT 8080 Port nasłuchiwania routera LLM_ROUTER_MODELS_CONFIG resources/configs/models-config.json Ścieżka do pliku konfiguracyjnego LLM_ROUTER_USE_PROMETHEUS 1 (jeśli zainstalowano metrics ) Włącza endpoint /api/metrics Router będzie dostępny pod http://0.0.0.0:8080/api . Pełna lista dostępnych zmiennych środowiskowych znajduje się w opisie zmiennych środowiskowych 7️⃣ Test pełnego stosu (router → vLLM) # ```shell script curl http://loc…"},{"k":"examples/quickstart/google-gemma3-12b-it/index.html","t":"🚀 Quick‑Start Guide for google/gemma-3-12b‑it with vLLM & LLM‑Router","s":"Examples","x":"# 🚀 Quick‑Start Guide for `google/gemma-3-12b‑it` with **vLLM** & **LLM‑Router** This guide walks you through: 1. **Installing vLLM** and the `google/gemma‑3‑12b‑it` model. 2. **Installing LLM‑Router** (the API gateway). 3. **Running the r…","h":["📋 Prerequisites","1️⃣ Set up a virtual environment","5️⃣ Prepare the router configuration","7️⃣ Test the full stack (router → vLLM)","🎉 What’s next?"],"b":"🚀 Quick‑Start Guide for google/gemma-3-12b‑it with vLLM &amp; LLM‑Router # This guide walks you through: Installing vLLM and the google/gemma‑3‑12b‑it model. Installing LLM‑Router (the API gateway). Running the router with the model configuration provided in models-config.json . All commands assume you are working on a Unix‑like system (Linux/macOS) with Python 3.10.6 and virtualenv available. 📋 Prerequisites # Requirement Details OS Ubuntu 20.04 + (or any recent Linux/macOS) Python 3.10.6 (project’s default) GPU CUDA 11.8 + (≥ 24 GB VRAM) or CPU‑only setup Tools git , curl , jq (optional but handy for testing) Network Ability to pull Docker images / PyPI packages and download the model from Hugging Face 1️⃣ Set up a virtual environment # ```shell script Create a directory for the whole demo (optional) # mkdir -p ~/gemma3-demo &amp;&amp; cd $_ Initialise the venv # python3 -m venv .venv source .venv/bin/activate Upgrade pip (always a good idea) # pip install --upgrade pip --- ## 2️⃣ Install **vLLM** and download the Gemma 3 model &gt; **See the full step‑by‑step instructions in** [`VLLM.md`](./VLLM.md). --- ## 3️⃣ Run the **vLLM** server Copy the helper script (or run the command manually) inside the demo directory: ```shell script # If you have the script `run-gemma-3-12b-it-vllm.sh` in the repo: cp path/to/llm-router/examples/quickstart/google-gemma3-12b-it/run-gemma-3-12b-it-vllm.sh . chmod +x run-gemma-3-12b-it-vllm.sh # Start the server (you may want to use tmux/screen) ./run-gemma-3-12b-it-vllm.sh The server will listen on http://0.0.0.0:7000 and expose an OpenAI‑compatible endpoint at /v1/chat/completions . You can quickly test it: ```shell script curl http://localhost:7000/v1/chat/completions \\ -H \"Content-Type: application/json\" \\ -d '{ \"model\": \"google/gemma-3-12b-it\", \"messages\": [{\"role\": \"user\", \"content\": \"Hello, how are you?\"}], \"max_tokens\": 100 }' | jq You should receive a JSON payload with the model’s generated text. --- ## 4️⃣ Install **LLM‑Router** ### Local install ```shell script # Clone the router repository (if you haven’t already) git clone https://github.com/radlab-dev-group/llm-router.git cd llm-router # Install the core library + API wrapper (includes the REST server) pip install .[api] # (Optional) Install Prometheus metrics support pip install .[api,metrics] Note: The router uses the same virtual environment you created earlier, so all dependencies stay isolated. 5️⃣ Prepare the router configuration # The example repository already ships a models-config.json that points to the locally running vLLM instance: { \"google_models\" : { \"google/gemma-3-12b-it\" : { \"providers\" : [ { \"id\" : \"gemma3_12b-vllm-local:7000\" , \"api_host\" : \"http://localhost:7000/\" , \"api_type\" : \"vllm\" , \"input_size\" : 56000 , \"weight\" : 1.0 } ] } }, \"active_models\" : { \"google_models\" : [ \"google/gemma-3-12b-it\" ] } } Copy it (or edit the path) to the router’s resources/configs/ directory: ```shell script mkdir -p resources/configs cp path/to/google-gemma3-12b-it/models-config.json resources/configs/ --- ## 6️⃣ Run the **LLM‑Router** ### Local Gunicorn The helper script `run-rest-api-gunicorn.sh` sets a sensible default environment. You can use it directly or export the variables yourself. ```shell script # Make the script executable (if needed) chmod +x path/to/run-rest-api-gunicorn.sh # Run the router ./run-rest-api-gunicorn.sh Key environment variables (already defined in the script) you may want to adjust: Variable Default Meaning LLM_ROUTER_SERVER_TYPE gunicorn Server backend (gunicorn, flask, waitress) LLM_ROUTER_SERVER_PORT 8080 Port on which the router listens LLM_ROUTER_MODELS_CONFIG resources/configs/models-config.json Path to the JSON file above LLM_ROUTER_PROMPTS_DIR resources/prompts Prompt‑template directory (optional) LLM_ROUTER_BALANCE_STRATEGY first_available Load‑balancing strategy LLM_ROUTER_USE_PROMETHEUS 1 (if you installed metrics) Enable /api/metrics endpoint After the script starts, the router will be re…"},{"k":"changelog.html","t":"Changelog","s":"Release notes","x":"## Changelog | Version | Changelog | |-----------|---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------…","h":["Changelog"],"b":"Changelog # Version Changelog 0.0.1 Initialization, License, setup, interface for each endpoint and sample ping EP. Autoloader of builtin endpoints and for the future implementations. 0.0.2 Add base models for api call (module llm_proxy_rest.data_models with error.py handling. Decorators to check required params and to measure the response time. 0.0.3 Proper AutoLoading for each found endpoint. Implementation of ApiTypesDispatcher , ApiModelConfig , ModelHandler . Ollama endpoints: / , tags . Added endpoint to full proxy with params. Streaming in case when external api provides stream. 0.0.4 All llama-service endpoints are refactored to llm-proxy-api . Refactoring base ep_run method. Proper handling system message, prompt name, model etc. 0.1.0 Repository name changed from llm-proxy-api to llm-router . Added class HttpRequestExecutor to handle http requests from EndpointWithHttpRequestI . Handled routing between any models: openai -&gt; ollama and ollama -&gt; openai 0.1.1 Prometheus metrics logging. Workers/Threads/Workers class is able to set by environments. Streaming fixes. Multi-providers for single model with default-balanced strategy. 0.2.0 Add balancing strategies: balanced , weighted , dynamic_weighted and first_available which works for streaming and non streaming requests. Included Prometheus metrics logging via /metrics endpoint. First stage of llm_router_lib library, to simply usage of llm-router-api . 0.2.1 Fix stream: OpenAI-&gt;Ollama, Ollama-&gt;OpenAI. Add Redis caching of availability of model providers (when using first_available strategy). Add llm_router_web module with simple flask-based frontend to manage llm-router config files. 0.2.2 Update dockerfile and requirements. Fix routing with vLLM. 0.2.3 New web configurator: Handling projects, configs for each user separately. First Available strategy is more powerful, a lot of improvements to efficiency. 0.2.4 Anonymizer module, integration anonymization with any endpoint (using dynamic payload analysis and full payload anonymisation), dedicated /api/anonymize_text endpoint as memory only anonymization. Whole router may be run in FORCE_ANONYMISATION mode. 0.3.0 Anonymization available with three strategies: fast_masker , genai , prov_masker . 0.3.1 Refactoring lb.strategies to be more flexible modular. Introduced MaskerPipeline and GuardrailPipeline both configured via env. Removed genai-based masking endpoint. 0.4.0 The main repository is divided into dedicated ones: plugins, services, web — separate repositories. Clean up the whole repository. Examples of integration with llamaindex, langchain, openai, litellm and haystack. 0.4.1 Audit log is stored using GPG. Add bash script ( scripts/gen_and_export_gpg.sh to prepare GPG keys and simple scripts/decrypt_auditor_logs.sh to decrypt encrypted audit logs. Moved core functionality from base to module core module. Quickstart. 0.4.2 Fix first_available_optim Strategy. Add KeepAliveMonitor to periodically pings model endpoints to keep them warm. 0.4.3 Add custom Prometheus metrices for logging masker/guardrail inidents. Fix OpenAI compatible v1 /models endpoint. Introduce monitors: services and keep alive models. Fixed guardrail retunr in case when streaming. 0.4.4 Validate unique provider identifiers. Store all hosts with keep‑alive configured in a Redis. UtilsPlugin pipeline with LangChain based simple RAG plugin (extending context to GenAI with locally built databse). Add handling of v1/response endpoint 0.4.5 Fixed sreaming to LMStudio native. Refactor streaming module. 0.4.6 Added support for embeddings endpoints across all providers. Extended ApiModel and ApiTypesI with is_embedding flag. Added test_embeddings.py utility for verifying embedding models through the API. 0.4.7 Integration with native Anthropic API. Add translate , generative_answer and ping methods to LLMRouterClient (with tests). Refactor LLMRouterCkientServices to use self.model_cls . Add payload converter for vLLM. 0.5.0 Integration with P…"}]}