{"version":"1.1.6","pages":[{"k":"overview.html","t":"Overview","s":"Getting started","x":"# LLM Router - Open-Source AI Gateway for Local and Cloud LLM Infrastructure [**LLM Router**](https://llm-router.cloud) is a service that can be deployed on‑premises or in the cloud. It adds a layer between any application and the LLM prov…","h":["🌐 Ecosystem Overview","✨ Key Features","🧩 Plugin System Architecture","Data flow","Masker Plugins","Guardrail Plugins","Utility Plugins","Semantic BiEncoder Routing","📦 Quick Start","1️⃣ Create &amp; activate a virtual environment","2️⃣ Install (recommended: from PyPI)","3️⃣ Run the REST API","4️⃣ Quick‑start guides for local models","5️⃣ Integration boilerplates","🔐 Auditing","🔐 Authentication","⏱️ Rate Limiting","🖥️ CLI Reference","🔒 Security","🔍 Error message sanitization","📦 Docker","Kubernetes (Helm)","Configuration","⚖️ Load Balancing Strategies","🛣️ Endpoints Overview","Highlights","🌐 Web Applications","Config Manager (port 8081)","Anonymizer (port 8082)","🧰 llm-router-utils","CLI Tools","Speakleash Deployment Configs","⚙️ Configuration Details","🔧 Development","📚 Changelog","📜 License"],"b":"LLM Router - Open-Source AI Gateway for Local and Cloud LLM Infrastructure # LLM Router is a service that can be deployed on‑premises or in the cloud. It adds a layer between any application and the LLM provider. In real time it controls traffic, distributes load among providers of a specific LLM, and enables analysis of outgoing requests from a security perspective (masking, anonymization, prohibited content). It is an open‑source solution (Apache 2.0) that can be launched instantly by running a ready‑made image in your own infrastructure. 🌐 Ecosystem Overview # The LLM‑Router project is split across five dedicated repositories: Repository Description llm-router (this repo) Core gateway — unified REST proxy, Python SDK, and configuration management llm-router-api (subdirectory) REST proxy that routes requests to any supported LLM backend (OpenAI‑compatible, Ollama, vLLM, LM Studio, Anthropic), with built‑in load‑balancing, health checks, streaming responses and optional Prometheus metrics llm-router-lib (subdirectory) Python SDK that wraps the API with typed request/response models, automatic retries, token handling, a rich exception hierarchy, and sync ( LLMRouterClient ) + async ( AsyncLLMRouterClient , streaming) clients llm-router-web Ready‑to‑use Flask UIs — a Config Manager for model/user settings and an Anonymizer UI that masks sensitive data llm-router-plugins Pluggable anonymizers (maskers), guardrails, semantic routing and RAG plugins llm-router-services HTTP services that power the plugin ecosystem (NASK‑PIB/Sojka guardrails, PII masker) llm-router-utils CLI tools, batch translation, GenAI classification and ready‑made deployment configs (Speakleash models) ✨ Key Features # Feature Description Unified REST interface One endpoint schema works for OpenAI‑compatible, Ollama, vLLM, LM Studio and Anthropic. Provider‑agnostic streaming The stream flag (default true ) controls whether the proxy forwards chunked responses as they arrive or returns a single aggregated payload. Streaming responses include proper Cache‑Control, Pragma, Expires and Vary headers. Built‑in prompt library Language‑aware system prompts stored under resources/prompts can be referenced automatically. Dynamic model configuration JSON file ( models-config.json ) defines providers, model name, default options and per‑model overrides. Request validation Pydantic models guarantee correct payloads; errors are returned with clear messages. Structured logging Configurable log level, filename, and optional JSON formatting. Health &amp; metadata endpoints /health and /api/ping (liveness), /api/version (build), /models , /v1/models and /api/tags (metadata), / for the Ollama probe. Built‑in article generation Two builtin endpoints were added: /api/generate_article_from_texts — generate a short (~A4) Polish article summarising a list of texts; and /api/create_full_article_from_texts — create a fuller article framed by user_query . Embeddings support Dedicated endpoints for generating text embeddings across all supported providers. Simple deployment One‑liner run script, Docker image, or Helm chart for Kubernetes. Extensible conversation formats Basic chat, conversation with system prompt, and extended conversation with richer options (temperature, top‑k, custom system prompt). Multi‑provider model support Each model can be backed by multiple providers (VLLM, Ollama, OpenAI, Anthropic) defined in models-config.json . Model‑level fallback A model can declare fallback_model : when none of its own providers can serve a request, the router reroutes it — before load balancing — to the fallback model, whose providers are balanced by the same strategy ( Models configuration ). Provider failover Any 4xx/5xx answer (or an unreachable provider) re-issues the request on the next provider of the same model — streaming included; only after every provider of the model was tried does fallback_model take over. Load‑balanced default strategy LoadBalancedStrategy picks the least‑…"},{"k":"llm-router-api/docs/installation.html","t":"Installation","s":"Getting started","x":"# Installation Three ways to get **LLM Router** running: from **PyPI** (recommended), from the **GitHub** source repository, or as a pre-built **Docker image on Quay**. Requires **Python ≥ 3.10**. --- ## 1. PIP (recommended) The package is…","h":["1. PIP (recommended)","Create &amp; activate a virtual environment","Install","Extras","2. GitHub (from source)","Option A — clone &amp; install locally","Option B — install directly from the git URL (no clone)","3. Quay (Docker image)","Quick start","Advanced usage","Kubernetes (Helm)"],"b":"Installation # Three ways to get LLM Router running: from PyPI (recommended), from the GitHub source repository, or as a pre-built Docker image on Quay . Requires Python ≥ 3.10 . 1. PIP (recommended) # The package is published on PyPI: radlab-llm-router . Create &amp; activate a virtual environment # python3 -m venv .venv source .venv/bin/activate Install # # Only the core library (llm_router_lib). pip install radlab-llm-router # Core library + API wrapper (llm_router_api). pip install radlab-llm-router [ api ] # Core library + API wrapper + Prometheus metrics. pip install radlab-llm-router [ api,metrics ] Extras # Extra Adds api The REST API wrapper ( llm_router_api ) metrics Prometheus metrics ( prometheus-client ) vault HashiCorp Vault integration ( hvac , bcrypt ) Note: When Prometheus metrics are enabled, LLM_ROUTER_USE_PROMETHEUS=1 must be set and Redis is required (used for provider availability state). The multiproc directory defaults to $HOME/.llm-router/metrics/prometheus/multiproc — override via the PROMETHEUS_MULTIPROC_DIR environment variable if needed. 2. GitHub (from source) # Installing from source pulls the latest code from the llm-router repository — prefer PyPI for production deployments. Option A — clone &amp; install locally # git clone https://github.com/radlab-dev-group/llm-router.git cd llm-router python3 -m venv .venv source .venv/bin/activate # Only the core library (llm_router_lib). pip install . # Core library + API wrapper (llm_router_api). pip install . [ api ] # Core library + API wrapper + Prometheus metrics. pip install . [ api,metrics ] Option B — install directly from the git URL (no clone) # # Pin a release tag (e.g. v1.0.5): pip install \"git+https://github.com/radlab-dev-group/llm-router@v1.0.5\" # With extras: pip install \"git+https://github.com/radlab-dev-group/llm-router@v1.0.5#[api]\" Note: @main tracks the repository HEAD and may contain unreleased, unstable changes — always pin a release tag for reproducible installs. 3. Quay (Docker image) # A pre-built container image is available on Quay : quay.io/radlab/llm-router . Quick start # docker pull quay.io/radlab/llm-router:rc1 docker run -p 5555 :8080 quay.io/radlab/llm-router:rc1 Advanced usage # All runtime settings are configured via LLM_ROUTER_* environment variables — see the Docker section of the README for the full custom launch-script example (server type, timeouts, balance strategy, Redis, models config mount, …), and ENV_DEFINITIONS.md for the complete variable reference. Kubernetes (Helm) # Helm charts for Kubernetes deployment are available in the helm_charts/ directory."},{"k":"llm-router-api/docs/env-definitions.html","t":"Environment Variables","s":"Configuration","x":"# Environment Variables All environment variables share the `LLM_ROUTER_` prefix. They are loaded from `os.environ` in [ `llm_router_api/base/constants.py`](llm_router_api/base/constants.py) at import time and validated by `_StartAppVerifi…","h":["Core variables","Redis variables","Masking &amp; Guardrail variables","Payload masking","Request guardrails","Response guardrails","Semantic BiEncoder Routing variables","LangChainRAG variables","Utils plugins variables","Authentication variables","Core switch","Memory store","Custom policies","Vault settings","Redis cache for keys","Auth Redis (separate from general REDIS)","Rate limiting","Public endpoints","Hardening","Key generation","Key rotation","Audit","Monitoring intervals","Programmatic source"],"b":"Environment Variables # All environment variables share the LLM_ROUTER_ prefix. They are loaded from os.environ in llm_router_api/base/constants.py at import time and validated by _StartAppVerificator . Core variables # Variable Default Description LLM_ROUTER_PROMPTS_DIR resources/prompts Directory containing predefined system prompts. LLM_ROUTER_MODELS_CONFIG resources/configs/models-config.json Path to the models configuration JSON file. LLM_ROUTER_DEFAULT_EP_LANGUAGE pl Default language for endpoint prompts (fallback). LLM_ROUTER_TIMEOUT 0 Timeout (seconds) for llm-router API calls. LLM_ROUTER_EXTERNAL_TIMEOUT 300 Timeout (seconds) for external model API calls. LLM_ROUTER_MAX_REQUEST_BODY_SIZE 10485760 (10 MB) Maximum request body size in bytes; oversized payloads get HTTP 413. LLM_ROUTER_LOG_FILENAME llm-router.log Name of the log file. LLM_ROUTER_LOG_LEVEL INFO Logging level (e.g. INFO, DEBUG). LLM_ROUTER_LOG_TO_FILE false Also write logs to the log file (in addition to console). LLM_ROUTER_LOG_MAX_BYTES 52428800 (50 MB) Rotate the log file once it reaches this size in bytes (applied when LLM_ROUTER_LOG_TO_FILE is set). LLM_ROUTER_LOG_BACKUP_COUNT 5 Maximum number of rotated log files to keep, &lt;name&gt;.1 … &lt;name&gt;.N . LLM_ROUTER_EP_PREFIX /api Prefix for all API endpoints. LLM_ROUTER_MINIMUM False Run service in proxy-only mode. LLM_ROUTER_IN_DEBUG False Run server in debug mode; also forces log level to DEBUG. LLM_ROUTER_VERBOSE False Log RAW, unmasked request params (PII!). Startup logs a warning and waits 3 s. Never use in production. LLM_ROUTER_BALANCE_STRATEGY balanced Load-balancing strategy: balanced , weighted , dynamic_weighted , first_available , first_available_optim , first_available_optim_nworkers . LLM_ROUTER_LB_SLOT_LEASE_SECONDS 120 Lifetime of one held worker slot in first_available_optim_nworkers . Renewed while the router process lives, so it bounds how long a crashed process keeps occupying a provider. Must exceed the longest expected request only if the KeepAlive monitor cannot keep up. LLM_ROUTER_LB_SLOT_MAX_AGE_SECONDS 1800 How long an unreleased worker slot keeps being renewed, counted from the acquisition. Covers a request that never released its slot (an abandoned stream): it is left to expire instead of occupying the provider until the router restarts. 0 or negative disables the cap. LLM_ROUTER_LB_HOST_PIN_TTL_SECONDS 3600 Lifetime of the “this host serves this model” pin used by first_available_optim / first_available_optim_nworkers (Redis key host:&lt;host&gt; ). Refreshed on every selection, so only a host nothing has selected for this long becomes available to another model. LLM_ROUTER_SERVER_TYPE flask Server implementation: flask, gunicorn, waitress. LLM_ROUTER_SERVER_PORT 8080 Port on which the server listens. LLM_ROUTER_SERVER_HOST localhost Host address for the server. LLM_ROUTER_SERVER_WORKERS_COUNT 2 Number of workers (for servers that support them). LLM_ROUTER_SERVER_THREADS_COUNT 8 Number of worker threads (for servers that support them). LLM_ROUTER_SERVER_WORKER_CLASS None Worker class for servers that support it (e.g. gevent). LLM_ROUTER_USE_PROMETHEUS False Enable Prometheus metrics collection ( /metrics endpoint). See also PROMETHEUS_MULTIPROC_DIR for the directory where Prometheus multiprocess worker data files are stored. LLM_ROUTER_VERBOSE is a debugging aid, not a logging level. With this variable on, SecureEndpointI stores it as self._verbose_mode and run_ep writes the raw, unmasked request payload (JSON) to the log before masking, so the log contains exactly the PII that masking would otherwise strip. On top of that, rest_api.main() emits a WARNING and sleeps 3 s ( VERBOSE_STARTUP_DELAY_SECONDS ) before the server binds its port, so an accidental production start can still be aborted. Enable it only for local troubleshooting ( LLM_ROUTER_VERBOSE=1 , llm-router server start --verbose ) and never in production or on shared log storage. Redis variables # Variable De…"},{"k":"llm-router-api/docs/models-config.html","t":"Models configuration description (JSON)","s":"Configuration","x":"# Models configuration description (JSON) ## 📄 Purpose This document explains the **model configuration** used by the LLM Router. It describes the JSON schema that drives **`ModelHandler`** and **`ApiModelConfig`**, clarifies each field, a…","h":["📄 Purpose","🏗️ High‑level structure","🔎 Detailed field description","Provider dictionary (items in providers / providers_sleep)","🛟 fallback_model (model level)","🔁 Failover order (providers first, then fallback_model)","active_models section","🧩 How ModelHandler uses the config","📦 Sample configuration (models-config.json)","🎉 Summary"],"b":"Models configuration description (JSON) # 📄 Purpose # This document explains the model configuration used by the LLM Router. It describes the JSON schema that drives ModelHandler and ApiModelConfig , clarifies each field, and provides a ready‑to‑use example ( models-config.json ). Having a single source of truth for model definitions makes it easy to: Add or remove providers for a given model. Switch between cloud (OpenAI, Google) and local (vLLM, Ollama) back‑ends. Control load‑balancing, keep‑alive, and tool‑calling options per provider. Activate only the models you want to expose through the router. 🏗️ High‑level structure # { \"&lt;model_type&gt;\": { # e.g. \"google_models\", \"openai_models\", \"qwen_models\" \"&lt;model_name&gt;\": { # full identifier used by the router, e.g. \"google/gemma-3-12b-it\" \"providers\": [ … ], # primary providers (used for normal traffic) \"providers_sleep\": [ … ] # optional low‑priority providers (used when others are busy) \"fallback_model\": \"…\" # optional model used when no provider can serve this one }, … }, \"active_models\": { # **required** – tells the router which models are enabled \"&lt;model_type&gt;\": [ \"&lt;model_name&gt;\", … ], … } } Model type – a top‑level key grouping models that share the same provider‑type logic. Model name – the identifier that appears in API calls ( model field). providers – a list of dictionaries, each describing a concrete endpoint. providers_sleep (optional) – “sleeping” providers that are only used when all primary providers are unavailable or overloaded. fallback_model (optional) – name of another active model that takes over when no provider of this model can serve the request. See Fallback model . active_models – the only place where a model is marked as active . If a model is missing here, the router will ignore it even if it is present in the rest of the file. 🔎 Detailed field description # Provider dictionary (items in providers / providers_sleep ) # Field Type Description Example id str Unique identifier for the provider instance (used for logging &amp; selection). \"gemma3_12b-vllm-71:7000\" api_host str Base URL of the provider API (must include protocol, may contain trailing slash). \"http://192.168.100.71:7000/\" api_token str Authentication token; empty string if not required. \"\" api_type str Type of the backend – determines which concrete BaseProvider class is used ( openai , vllm , ollama , …). \"vllm\" input_size int (or numeric string) Maximum context length the provider accepts. The ApiModel.from_config helper converts it to int . 4096 model_path str Path or name of the model on the provider side (used by Ollama, vLLM, etc.). May be empty for providers that infer it from the URL. \"gpt-3.5-turbo-0125\" weight float Relative weight for weighted‑random load‑balancing strategies. Default 1.0 . 0.1 nworkers int (or numeric string) Maximum number of concurrent requests allowed on the provider. Used only by the first_available_optim_nworkers load‑balancing strategy; a missing, invalid or non‑positive value falls back to 1 . Read from the live configuration on every selection, so changing it takes effect on the next request — the health monitor's registration copy ( monitor:providers:&lt;model&gt; ) is not used for the limit. 1 keep_alive str Optional keep‑alive duration (e.g. \"35m\" ). Empty or null means the provider is not kept alive. \"35m\" tool_calling bool Whether the provider supports tool‑calling (function calling). true is_embedding bool Whether the model is an embedding model (determines use of embedding endpoints). true 🛟 fallback_model (model level) # A model may name another active model that takes over when none of its own providers can serve a request. The switch happens before load balancing : the fallback model is handed to the load-balancing strategy, which then picks one of its providers exactly like for any other request. { \"qwen_models\" : { \"qwen/Qwen3.8-Flash-Next\" : { \"fallback_model\" : \"qwen/qwen3-coder:30b\" , \"providers\" : [ { \"id\" : \"qwen3.8…"},{"k":"llm-router-api/docs/keepalive.html","t":"Keep‑Alive Utility Overview","s":"Routing & resilience","x":"# Keep‑Alive Utility Overview The **keep‑alive** subsystem is responsible for periodically “pinging” model endpoints so that they stay warm and ready to serve requests with minimal latency. It consists of two main components: | Component |…","h":["How It Works","Integration Points","Configuring Keep‑Alive","Example Usage","Logging"],"b":"Keep‑Alive Utility Overview # The keep‑alive subsystem is responsible for periodically “pinging” model endpoints so that they stay warm and ready to serve requests with minimal latency. It consists of two main components: Component Purpose Key Methods KeepAlive Sends a single HTTP request to a model provider. It resolves the correct provider configuration (API type, host, token, model name) and builds the request payload. send(model_name, host, prompt=None) – performs the HTTP call. KeepAliveMonitor Schedules repeated keep‑alive calls for each (model_name, host) pair. It stores scheduling data in Redis, checks host availability, and triggers KeepAlive.send when a provider is due. record_usage(model_name, host, keep_alive) – registers a provider for periodic pinging. start() / stop() – control the background thread. How It Works # Provider discovery – KeepAlive._find_provider looks up the provider configuration for a given model name and host inside the global models_configs dictionary. Endpoint resolution – Depending on the provider’s api_type ( vllm , openai , ollama ), _endpoint_for builds the correct URL ( /v1/chat/completions or /api/chat ). HTTP request – A JSON payload containing a short “keep‑alive” prompt (default: “Send an empty message.” ) is posted to the endpoint. Scheduling – KeepAliveMonitor stores metadata in Redis: A hash key ( keepalive:provider:&lt;model&gt;:&lt;host&gt; ) with keep_alive_seconds . A sorted‑set ( keepalive:providers:next_wakeup ) that orders providers by the next scheduled wake‑up timestamp. A sorted‑set ( keepalive:model:&lt;model_name&gt;:hosts ) that tracks all hosts where a model is currently loaded, used for host reuse in optimised strategies. Background loop – The monitor thread wakes up every check_interval seconds, fetches due providers, verifies that the host is free (via the optional is_host_free_callback ), and invokes KeepAlive.send . After a successful ping, the next wake‑up time is recomputed. Integration Points # Strategy implementations (e.g., FirstAvailableOptimStrategy ) create a KeepAlive instance and pass it to a KeepAliveMonitor . When a provider is selected, the strategy calls keep_alive_monitor.record_usage(model_name, host, keep_alive) so the monitor knows to ping that endpoint. The monitor runs automatically in the background once start() is called (typically during strategy initialization). Configuring Keep‑Alive # Provider configurations live in the global models_configs JSON (see resources/configs/models-config*.json ). To enable keep‑alive for a specific provider, add a keep_alive field with a duration string: { \"model_name\" : \"gpt‑4\" , \"providers\" : [ { \"api_type\" : \"openai\" , \"api_host\" : \"http://localhost:8000\" , \"api_token\" : \"YOUR_TOKEN\" , \"keep_alive\" : \"2m\" // ping every 2 minutes } ] } Supported duration units: Unit Suffix Meaning seconds s e.g., \"30s\" minutes m e.g., \"5m\" hours h e.g., \"1h\" If the keep_alive field is omitted or falsy, the provider will not be scheduled for periodic pings. Example Usage # from llm_router_api.core.monitor.keep_alive import KeepAlive from llm_router_api.core.monitor.keep_alive_monitor import KeepAliveMonitor # Assume `models_configs` has been loaded from the JSON config files. keep_alive = KeepAlive ( models_configs = models_configs ) monitor = KeepAliveMonitor ( redis_client = redis_client , keep_alive = keep_alive , check_interval = 10.0 , # check every 10 seconds is_host_free_callback = my_is_host_free , clear_buffers = True , # clean old keys on start ) monitor . start () # When a provider is selected somewhere in the routing logic: monitor . record_usage ( model_name = \"gpt‑4\" , host = \"http://localhost:8000\" , keep_alive = \"2m\" ) The monitor will now ping the gpt‑4 endpoint every two minutes, provided the host is not busy with another model. Logging # Both KeepAlive and KeepAliveMonitor emit detailed logs at the DEBUG and INFO levels, prefixed with [keep-alive] . Adjust your logger configuration to capture these messa…"},{"k":"llm-router-api/docs/lb-strategies.html","t":"Load Balancing Strategies","s":"Routing & resilience","x":"## Load Balancing Strategies The `llm-router` supports various strategies for selecting the most suitable provider when multiple options exist for a given model. This ensures efficient and reliable routing of requests. The available strate…","h":["Load Balancing Strategies","1. balanced (Default)","2. weighted","3. dynamic_weighted (beta)","4. first_available","5. first_available_optim","6. first_available_optim_nworkers","7. AdaptiveStrategy (beta)","Environments and Redis installation","Extending with Custom Strategies"],"b":"Load Balancing Strategies # The llm-router supports various strategies for selecting the most suitable provider when multiple options exist for a given model. This ensures efficient and reliable routing of requests. The available strategies are: Model‑level fallback_model runs before the strategy. A strategy is always asked to serve the model named by the client. Only when that model cannot be served at all — no providers, no healthy provider, or every provider busy until the selection timeout — ModelHandler reroutes the request to the configured fallback_model and the very same strategy then balances over the providers of that model . Strategies themselves are unaffected; the option is documented in MODELS_CONFIG.md . Provider errors rotate first.— A provider that answers with an error (any 4xx/5xx) or is unreachable makes the dispatcher retry the request on another provider of the same model ; the model's fallback_model is reached only once none of its providers is left untried. 1. balanced (Default) # Description: This is the default strategy. It aims to distribute requests evenly across available providers by keeping track of how many times each provider has been used for a specific model. It selects the provider that has been used the least. When to use: Ideal for scenarios where all providers are considered equal in terms of capacity and performance. It provides a simple and effective way to balance the load. Implementation: Implemented in llm_router_api.core.lb.balanced.LoadBalancedStrategy . 2. weighted # Description: This strategy allows you to assign static weights to providers. Providers with higher weights are more likely to be selected. The selection is deterministic, ensuring that over time, the request distribution closely matches the configured weights. When to use: Useful when you have providers with different capacities or performance characteristics, and you want to prioritize certain providers without needing dynamic adjustments. Implementation: Implemented in llm_router_api.core.lb.weighted.WeightedStrategy . 3. dynamic_weighted (beta) # Description: An extension of the weighted strategy. It not only uses weights but also tracks the latency between successive selections of the same provider. This allows for more adaptive routing, as providers with consistently high latency might be de-prioritized over time. You can also dynamically update provider weights. When to use: Recommended for dynamic environments where provider performance can fluctuate. It offers more sophisticated load balancing by considering both configured weights and real-time performance metrics (latency). Implementation: Implemented in llm_router_api.core.lb.weighted.DynamicWeightedStrategy . 4. first_available # Description: This strategy selects the very first provider that is available. It uses Redis to coordinate across multiple workers, ensuring that only one worker can use a specific provider at a time. When to use: Suitable for critical applications where you need the fastest possible response and want to ensure that a request is immediately handled by any available provider, without complex load distribution logic. It guarantees that a provider, once taken, is exclusive until released. Implementation: Implemented in llm_router_api.core.lb.first_available.FirstAvailableStrategy . When using the first_available load balancing strategy, a Redis server is required for coordinating provider availability across multiple workers. 5. first_available_optim # What it is first_available_optim is an enhanced version of the plain first‑available load‑balancing strategy. It uses Redis to coordinate across multiple workers and tries to reuse a host that has already been used for the requested model before falling back to the classic “pick the first free provider” logic. How it works Step Purpose Behaviour 1️⃣ Re‑use the last host If the model was previously run on a specific host and that host is currently free, select it. The host identifier is…"},{"k":"llm-router-api/core/auditor/index.html","t":"Auditing subsystem – llm-router","s":"Security & auditing","x":"# Auditing subsystem – `llm-router` The **auditor** package provides a pluggable, tamper‑evident audit‑log system for the LLM‑router. All audit entries are written as JSON, encrypted with GPG and stored under `logs/auditor`. The subsystem …","h":["📁 Directory layout","🛠️ How the auditor works","🔐 GPG key management","1️⃣ Generate a new key pair","2️⃣ Place the public key where the router expects it","3️⃣ Decrypt audit logs","📚 Example: Auditing a request guard‑rail decision","🧩 Extending the auditor","📖 Further reading"],"b":"Auditing subsystem – llm-router # The auditor package provides a pluggable, tamper‑evident audit‑log system for the LLM‑router. All audit entries are written as JSON, encrypted with GPG and stored under logs/auditor . The subsystem is used by the router to record: request guard‑rail decisions payload masking operations custom audit events emitted by the application (e.g. business‑logic logs) The implementation is deliberately lightweight so it can be swapped out for a different storage backend (database, cloud bucket, …) without touching the rest of the code base. 📁 Directory layout # llm_router_api/ └─ core/ └─ auditor/ ├─ __init__.py # package marker ├─ auditor.py # public API – AnyRequestAuditor └─ log_storage/ ├─ __init__.py ├─ log_storage_interface.py # abstract storage contract └─ gpg.py # GPG‑backed storage implementation auditor.py – high‑level helper that forwards audit records to a storage backend. The default backend is GPGAuditorLogStorage . log_storage_interface.py – defines the AuditorLogStorageInterface protocol ( store_log(audit_log, audit_type) ). gpg.py – concrete implementation that encrypts each log entry with the public GPG key located at resources/keys/llm-router-auditor-pub.asc and writes the encrypted payload to a timestamped file logs/auditor/&lt;audit_type&gt;__&lt;timestamp&gt;.audit . 🛠️ How the auditor works # Endpoint code (e.g. endpoint_i.py ) creates an AnyRequestAuditor instance with the router’s logger. When an auditable event occurs, the endpoint builds a dictionary that contains at least the keys audit_type and payload . AnyRequestAuditor.add_log() forwards the dictionary to the configured storage backend. GPGAuditorLogStorage.store_log() JSON‑serialises the dictionary (pretty‑printed). Encrypts the JSON string with the imported public key. Writes the encrypted ASCII‑armored data to logs/auditor/ . The resulting files have the extension .audit . They are confidential and tamper‑evident – any modification breaks the GPG decryption. 🔐 GPG key management # The repository ships two helper scripts under scripts/ : Script Purpose gen_and_export_gpg.sh Generates a 4096‑bit RSA key pair (no interactive prompts) and exports the public ( *.asc ) and private ( *-priv.asc ) keys. decrypt_auditor_logs.sh Decrypts all *.audit files in logs/auditor/ and writes the resulting JSON to *.json . 1️⃣ Generate a new key pair # cd scripts ./gen_and_export_gpg.sh The script will: Prompt for an email address (used as the GPG user ID). Prompt for a passphrase (protects the private key). Create a key pair in the local GPG keyring. Export the public key to llm-router-auditor-pub.asc . Export the private key to llm-router-auditor-priv.asc . Important: Keep the private key ( *-priv.asc ) and its passphrase safe. Only the public key is required by the router at runtime. 2️⃣ Place the public key where the router expects it # mkdir -p resources/keys cp llm-router-auditor-pub.asc resources/keys/ The GPGAuditorLogStorage class automatically imports the key from this location when the application starts. 3️⃣ Decrypt audit logs # cd scripts ./decrypt_auditor_logs.sh For each file logs/auditor/&lt;type&gt;__&lt;timestamp&gt;.audit the script produces a human‑readable *.json file next to it: logs/auditor/request__20231129_123456.789012.audit → request__20231129_123456.789012.json You will be prompted for the passphrase of the private key if it is encrypted. 📚 Example: Auditing a request guard‑rail decision # from llm_router_api.core.auditor.auditor import AnyRequestAuditor import logging logger = logging . getLogger ( \"router\" ) auditor = AnyRequestAuditor ( logger ) # Somewhere inside an endpoint, after a guard‑rail check: audit_record = { \"audit_type\" : \"guardrail_request\" , \"payload\" : { \"user_id\" : \"12345\" , \"input\" : \"…\" , \"decision\" : \"blocked\" , \"reason\" : \"PII detected\" } } auditor . add_log ( audit_record ) The record is encrypted and persisted as e.g.: logs/auditor/guardrail_request__20231129_141530.123456.audit 🧩 Exte…"},{"k":"llm-router-api/docs/authentication.html","t":"Authentication & Authorization","s":"Security & auditing","x":"# Authentication & Authorization API-key-based authentication with per-endpoint policies, rate limiting, audit trail, and Prometheus metrics. **Enabled by default: `LLM_ROUTER_AUTH_ENABLED=false`** — set to `\"true\"` to enforce authenticati…","h":["Architecture","Components","Environment Variables","CLI Commands","Key Management","Policy Management","Rate Limit","Seed File (Memory Store)","Seed File Format","How It Works","Default Location","Permission Engine","Builtin Policies","Prometheus Metrics","Key Format","Deployment Options","1️⃣ In-Memory Store — Development / Quick Start","2️⃣ Redis Store — Multi-Process / Stateful Single-Node","3️⃣ HashiCorp Vault — Production / Multi-Cluster / Enterprise","Comparison Matrix","See Also"],"b":"Authentication &amp; Authorization # API-key-based authentication with per-endpoint policies, rate limiting, audit trail, and Prometheus metrics. Enabled by default: LLM_ROUTER_AUTH_ENABLED=false — set to \"true\" to enforce authentication. Architecture # Client Request → AuthMiddleware → Key Store Lookup → Permission Engine → Rate Limiter → Endpoint → Audit Bridge → AnyRequestAuditor → AuthMetrics → Prometheus Components # Component Module Purpose Key Store core/auth/key_store/ Vault, Redis, or in-memory key storage Permission Engine core/auth/policies/engine.py Resolve key → policy → endpoint permissions Rate Limiter core/auth/rate_limiter.py Redis-backed sliding window rate limiter Key Generator core/auth/key_generator.py Generate keys in sk-llmr-live- format Audit Bridge core/auth/audit.py Bridge auth events → AnyRequestAuditor Metrics core/auth/metrics.py Prometheus counters &amp; histograms for auth Middleware core/auth/middleware.py Flask before_request hook Environment Variables # All environment variables are documented in ENV_DEFINITIONS.md . Auth-specific vars start with LLM_ROUTER_AUTH_* . CLI Commands # Key Management # # Generate a new key (persists to seed file when --store memory) llm-router auth key generate --policy developer --store memory # List all keys llm-router auth key list --store memory # Delete a key llm-router auth key delete key-id # Disable a key (revokes access immediately) llm-router auth key disable key-id [ --store memory ] # Enable a previously disabled key llm-router auth key enable key-id [ --store memory ] # Rotate a key (old key stays valid for grace_period) llm-router auth key rotate key-id --grace 3600 Policy Management # # List builtin policies (custom ones are marked with \"(custom)\") llm-router auth policy list # Create a new policy — inline JSON, from a file, or from stdin (-) llm-router auth policy create my-team '{\"can_access\": true, \"rate_limit\": 120}' llm-router auth policy create my-team --file my-team.json cat my-team.json | llm-router auth policy create my-team --file - Custom policies are persisted to $LLM_ROUTER_AUTH_CUSTOM_POLICIES_FILE (default: ~/.llm-router/configs/auth/custom-policies.json ) and resolved by the server without a restart. Rate Limit # # List available rate-limit presets llm-router auth rate-limit list # Apply a rate-limit preset to an existing key llm-router auth rate-limit apply key-id --preset pro # Remove rate-limit override from a key (revert to global default) llm-router auth rate-limit remove key-id Seed File (Memory Store) # When using --store memory , keys are stored in process memory — they are lost on restart . To persist keys across restarts (and between the CLI and router processes), use a seed file. Seed File Format # The seed file is a JSON array. Each record must carry a verifiable credential: a key_hash (+ key_index ) pair, or — for legacy files only — a key_plain which is hashed at load time and never written back . Plaintext keys are never persisted. Field Type Description key_id str Unique identifier for this key key_hash str bcrypt hash of the plaintext key (the verifiable credential) key_index str SHA-256 of the plaintext key — O(1) lookup index (locator only) key_prefix str First 12 characters of the plaintext key (shows the full sk-llmr-live prefix) policy_name str Name of the default policy to apply policy_override dict Inline policy override (takes precedence over the named policy) is_active bool Whether the key is currently valid expires_at float Expiry timestamp ( null = no expiry) created_at float Unix timestamp of key creation (auto-generated if omitted) last_used_at float Last successful authentication time rotate_at float Scheduled rotation time grace_until float Key remains valid until this time after rotation metadata dict Arbitrary metadata (team, cost_center, etc.) [ { \"key_id\" : \"manual-key-001\" , \"key_hash\" : \"$2b$12$Kx7Q2m...bcrypt-hash.../\" , \"key_index\" : \"9f2c...sha256-hex...\" , \"key_prefix\" : \"sk-llmr-live\" , \"pol…"},{"k":"llm-router-api/docs/rate-limiting.html","t":"Rate Limiting","s":"Security & auditing","x":"# Rate Limiting Sliding-window rate limiting backed by Redis sorted sets. Prevents API abuse, controls load on downstream LLM providers, and protects against brute-force key enumeration. --- ## How It Works ### Sliding Window Algorithm The…","h":["How It Works","Sliding Window Algorithm","Key + IP Binning","Bucket Naming","Client IP Resolution (anti-spoofing)","Failed-Authentication Lockout","Configuration","Environment Variables","Enabling Rate Limiting","Response Behavior","Architecture","RateLimitResult Dataclass","RedisRateLimiter Class","Prometheus Metrics","Rate Limit Strategy by Use Case","Per-User Quotas (Default)","Tiered Rate Limiting (Policy-Based)","Redis Requirements","Redis Persistence","Comparison with Fixed-Window","Why Sliding Window?","Best Practices","1. Always Enable Rate Limiting in Production","2. Set up auth Redis connection","3. Monitor Rate Limit Events","4. Set Graceful Limits","5. Use Public Endpoints for Health Checks","Migration from No-Auth / Default Rate","Troubleshooting","\"Why am I getting 429s immediately?\"","\"Rate limit is too strict/loose\"","\"Clients aren't respecting Retry-After\"","Rate Limit Presets","See Also"],"b":"Rate Limiting # Sliding-window rate limiting backed by Redis sorted sets. Prevents API abuse, controls load on downstream LLM providers, and protects against brute-force key enumeration. How It Works # Sliding Window Algorithm # The rate limiter uses a sliding window approach — unlike fixed-window counters, there are no boundary spikes. Each request timestamp is stored as a member score in a Redis sorted set, and entries older than the window are purged on every check. Time → ──────────────────────────────────────────▶ │&lt;────── WINDOW (60s) ──────▶│ ▼ ▼ │ ◉ ◉ ◉ ◉ ◉ ◉ ◉ ◉ │ │ ◉ ◉ ◉ ◉ ◉ ◉ ◉ ◉ ◉ ◉ │ │ │ ←─ remaining slots ────────→ Precision: per-request, no fixed-boundary artifacts Memory: O (requests per key+IP) — old entries are automatically purged Durability: Redis persistence (RDB/AOF) protects against server restart Key + IP Binning # Rate limits are enforced per API key + client IP . This means: Each API key gets its own independent quota Multiple IPs sharing the same key share that key's quota A malicious IP hitting a leaked key can exhaust its quota (mitigated by IP-level monitoring) Bucket Naming # auth:ratelimit:{key_id}:{ip} Example: auth:ratelimit:dev-a1b2c3d3:192.168.1.100 Client IP Resolution (anti-spoofing) # The client IP is the direct peer ( remote_addr ) by default. The X-Forwarded-For header is attacker-controlled and is honoured only when the direct peer is a configured trusted proxy ( LLM_ROUTER_TRUSTED_PROXIES , IP or CIDR); in that case the right-most XFF entry is used (proxies append, not prepend). Without this gate, a client could rotate X-Forwarded-For to obtain a fresh rate-limit bucket per request. export LLM_ROUTER_TRUSTED_PROXIES = \"10.0.0.1,192.168.0.0/24\" # your LB / reverse proxy Failed-Authentication Lockout # Failed attempts (missing or invalid key) never reach the key-scoped limiter, so they are throttled separately in a dedicated per-IP bucket ( auth:ratelimit:__auth_fail__:{ip} ) with the LLM_ROUTER_AUTH_FAILURE_LIMIT budget (default 20 /window; 0 disables). Once exceeded, the client receives 429 — protecting the key-lookup path from brute-force / DoS. Configuration # Environment Variables # All rate-limiting environment variables are documented in ENV_DEFINITIONS.md → Authentication section (Rate limiting subsection) . Enabling Rate Limiting # Rate limiting is applied automatically when authentication is enabled — there is no separate toggle. Just enable auth and configure the default rate limit: export LLM_ROUTER_AUTH_ENABLED = true export LLM_ROUTER_AUTH_DEFAULT_RATE_LIMIT = 60 # 60 requests per minute per key (default) python -m llm_router_api.rest_api Response Behavior # When a request exceeds the rate limit, the router returns: HTTP Status Header Value 429 Too Many Requests Retry-After Seconds until the oldest request in the window expires Example: HTTP/1.1 429 Too Many Requests Retry-After: 35 {\"error\": {\"message\": \"Rate limit exceeded. Please retry later.\", \"type\": \"rate_limit_error\", \"code\": 429, \"retry_after\": 35}} The Retry-After value is calculated from the oldest entry still in the window — this tells the client exactly when it can retry. Architecture # Client Request → AuthMiddleware → get_auth_result() → RedisRateLimiter.is_allowed(key_id, ip, limit) → → Redis sorted set operations (zremrangebyscore, zcard, zadd, expire) → RateLimitResult(allowed, remaining, retry_after) → If denied: HTTP 429 → If allowed: continue to endpoint RateLimitResult Dataclass # @dataclass class RateLimitResult : allowed : bool # Whether the request is within the limit remaining : int # Remaining requests in the current window retry_after : int # Seconds until the oldest request expires (0 if allowed) RedisRateLimiter Class # class RedisRateLimiter : PREFIX = \"auth:ratelimit\" WINDOW = 60 # seconds def __init__ ( self , redis_client : Optional [ redis . Redis ] = None , redis_host : Optional [ str ] = None , redis_port : int = 6379 , redis_db : int = 0 , redis_password : Optional [ str ] = None , …"},{"k":"llm-router-api/docs/routing-metrics.html","t":"Router Prometheus Metrics","s":"Observability","x":"# Router Prometheus Metrics Additional Prometheus metrics for the LLM router core (routing, provider lifecycle, pipeline funnel). Enabling these metrics requires `LLM_ROUTER_USE_PROMETHEUS=true` and the `metrics` extra (`pip install .[metr…","h":["Overview","A. Routing &amp; Provider Metrics","llm_router_provider_calls_total (Counter)","llm_router_provider_latency_seconds (Histogram)","llm_router_provider_error_total (Counter)","llm_router_lb_strategy_selected_total (Counter)","B. Pipeline / Request Funnel Metrics","llm_router_pipeline_stage_total (Counter)","llm_router_retry_total (Counter)","llm_router_retry_exhausted_total (Counter)","C. Token Usage Metric","llm_router_tokens_total (Counter)","D. Streaming &amp; Response Format Metrics","llm_router_response_format_total (Counter)","llm_router_payload_conversion_total (Counter)","Grafana Dashboard Snippet","Implementation Notes"],"b":"Router Prometheus Metrics # Additional Prometheus metrics for the LLM router core (routing, provider lifecycle, pipeline funnel). Enabling these metrics requires LLM_ROUTER_USE_PROMETHEUS=true and the metrics extra ( pip install .[metrics] ). Overview # The router tracks 10 new metrics across four categories: Category Metrics Purpose Routing &amp; Provider 4 Visibility into provider selection, latency, errors, and LB strategy Pipeline Funnel 3 Track how many requests pass through each pipeline stage Token Usage 1 Track input/output token consumption per model Streaming &amp; Format 2 Track streaming vs non-streaming distribution and payload conversions A. Routing &amp; Provider Metrics # llm_router_provider_calls_total (Counter) # Total number of successful calls to each provider type, grouped by model name. Labels: provider_type , model_name Example output: llm_router_provider_calls_total{provider_type=\"openai\", model_name=\"google/gemma-3-12b-it\"} 42 llm_router_provider_calls_total{provider_type=\"ollama\", model_name=\"google/gemma-3-12b-it\"} 18 Use case: Track provider utilization distribution across models. llm_router_provider_latency_seconds (Histogram) # Latency of outbound calls to specific providers, independent of end-to-end client latency. Labels: provider_type , model_name Buckets: 10ms → 30s (0.01, 0.025, 0.05, 0.1, 0.25, 0.5, 1, 2.5, 5, 10, 30) Example output: llm_router_provider_latency_seconds_count{provider_type=\"vllm\", model_name=\"google/gemma-3-12b-it\"} 100 llm_router_provider_latency_seconds_sum{provider_type=\"vllm\", model_name=\"google/gemma-3-12b-it\"} 45.2 Use case: Identify slow/degraded providers independently of client-side factors. Query example: histogram_quantile ( 0.95 , rate ( llm_router_provider_latency_seconds_bucket [ 5m ] )) llm_router_provider_error_total (Counter) # Provider errors by type and HTTP error code (for retriable status codes) or classification (timeout, connection_error). Labels: provider_type , model_name , error_code Example output: llm_router_provider_error_total{provider_type=\"ollama\", model_name=\"gpt-oss:120b\", error_code=\"429\"} 5 llm_router_provider_error_total{provider_type=\"openai\", model_name=\"google/gemma-3-12b-it\", error_code=\"timeout\"} 2 Use case: Alert on specific provider error spikes. Query example: sum by ( provider_type ) ( rate ( llm_router_provider_error_total [ 5m ] )) llm_router_lb_strategy_selected_total (Counter) # Load balancing strategy selections per model, tracking which strategies are used for each model's providers. Labels: strategy , model_name Example output: llm_router_lb_strategy_selected_total{strategy=\"balanced\", model_name=\"google/gemma-3-12b-it\"} 80 llm_router_lb_strategy_selected_total{strategy=\"weighted\", model_name=\"openai/gpt-4\"} 35 Use case: Verify load balancing strategy distribution. Useful when debugging or tuning LB strategies. B. Pipeline / Request Funnel Metrics # llm_router_pipeline_stage_total (Counter) # Request counts at each pipeline stage with pass/fail result. Stages tracked: Stage Possible Results provider_resolved success , failure request_received total (always) guardrail_request pass , block masking applied , skipped Labels: stage , result Example output: llm_router_pipeline_stage_total{stage=\"provider_resolved\", result=\"success\"} 1200 llm_router_pipeline_stage_total{stage=\"guardrail_request\", result=\"block\"} 45 Use case: Build a request funnel chart to see where requests drop off. Query example: # Guardrail block rate sum ( rate ( llm_router_pipeline_stage_total { stage = \" guardrail_request \", result = \" block \"}[ 5m ] )) / sum ( rate ( llm_router_pipeline_stage_total { stage = \" guardrail_request \"}[ 5m ] )) llm_router_retry_total (Counter) # Retry attempts per model and HTTP error code that triggered the retry. Every 4xx/5xx answer moves the request to another provider of the model (and, once all of them were tried, to its fallback_model ), so any error status can appear here. Labels: model_name , error_code Example outpu…"},{"k":"llm-router-api/index.html","t":"REST API reference","s":"REST API","x":"# llm‑router‑api **llm‑router‑api** is a lightweight Python library that provides a flexible, extensible proxy for Large Language Model (LLM) back‑ends. It abstracts the details of multiple model providers (OpenAI‑compatible, Ollama, vLLM,…","h":["Features","Installation","Running the Server","REST API Overview","Load‑Balancing Strategies","Keep‑Alive Mechanism","Extending the Router","Adding a New Provider Type","Adding a New Endpoint","Prompt Files","Monitoring &amp; Metrics","License"],"b":"llm‑router‑api # llm‑router‑api is a lightweight Python library that provides a flexible, extensible proxy for Large Language Model (LLM) back‑ends. It abstracts the details of multiple model providers (OpenAI‑compatible, Ollama, vLLM, LM Studio, etc.) and offers a unified REST interface with built‑in load‑balancing, health‑checking, and monitoring. Repository: https://github.com/radlab-dev-group/llm-router Features # Unified API – One REST surface ( /api/... ) that proxies calls to any supported LLM back‑end. Provider Selection – Choose a provider per request using pluggable strategies (balanced, weighted, adaptive, first‑available). Prompt Management – System prompts are stored as files and can be dynamically injected with placeholder substitution. Streaming Support – Transparent streaming for both OpenAI‑compatible and Ollama endpoints. Health Checks – Built‑in ping endpoint and Redis‑based provider health monitoring. Prometheus Metrics – Optional instrumentation for request counts, latencies, and error rates. Auto‑Discovery – Endpoints are automatically discovered and instantiated at startup. Extensible – Add new providers, strategies, or custom endpoints with minimal boilerplate. Installation # The project uses Python 3.10.6 and a virtualenv ‑based workflow. ```shell script Clone the repository # git clone https://github.com/radlab-dev-group/llm-router.git cd llm-router Create a virtual environment # python3 -m venv venv source venv/bin/activate Install the package (including optional extras) # pip install -e .[metrics] # installs Prometheus support All required third‑party libraries are listed in `requirements.txt` (e.g., Flask, requests, redis, rdl‑ml‑utils, etc.). --- ## Configuration Configuration is driven primarily by environment variables and a JSON model‑config file. ### Environment Variables All environment variables are documented in **[ENV_DEFINITIONS.md](./docs/ENV_DEFINITIONS.md)**. Key categories: **Core** · **Redis** · **Masking &amp; Guardrail** · **Semantic BiEncoder Routing** · **LangChainRAG** · * *Utils Plugins** · **Authentication**. --- ### Authentication variables Auth-specific environment variables are documented in **[ENV_DEFINITIONS.md](./docs/ENV_DEFINITIONS.md) → Authentication section**. &gt; See full authentication docs: **[AUTHENTICATION.md](docs/AUTHENTICATION.md)** ### Model Configuration `models-config.json` follows the schema: ```json { \"active_models\": { \"openai_models\": [ \"gpt-4\", \"gpt-3.5-turbo\" ], \"ollama_models\": [ \"llama2\" ] }, \"openai_models\": { \"gpt-4\": { \"providers\": [ { \"id\": \"openai-gpt4-1\", \"api_host\": \"https://api.openai.com/v1\", \"api_token\": \"sk-...\", \"api_type\": \"openai\", \"input_size\": 8192, \"model_path\": \"\" } ] } }, ... } Only the fields required by the router are needed: id , api_host , api_token (optional), api_type , input_size , and optionally model_path . Configuration Details – see the full schema and a ready‑made example in MODELS_CONFIG.md . Running the Server # The entry point is llm_router_api.rest_api . Choose a server backend via the LLM_ROUTER_SERVER_TYPE variable or command‑line flags. ```shell script Using the built‑in Flask development server (default) # python -m llm_router_api.rest_api Production‑grade with Gunicorn (streaming supported) # python -m llm_router_api.rest_api --gunicorn Windows‑friendly Waitress server # python -m llm_router_api.rest_api --waitress ``` The server starts on the host/port defined by LLM_ROUTER_SERVER_HOST and LLM_ROUTER_SERVER_PORT (default 0.0.0.0:8080 ). Note: The service must be launched with LLM_ROUTER_MINIMUM=1 (or any truthy value) because it operates in “proxy‑only” mode. REST API Overview # All routes are prefixed by LLM_ROUTER_EP_PREFIX (default /api ). The list of endpoints—categorized into built‑in, provider‑dependent, and extended endpoints—and a description of the streaming mechanisms can be found at the link: load endpoints overview Creating a new endpoint? See the Endpoint Development Guide — class hierarchy, …"},{"k":"llm-router-api/endpoints/index.html","t":"Endpoints Overview","s":"REST API","x":"## Endpoints Overview > **Creating new endpoints?** See the [Endpoint Development Guide](../docs/ENDPOINT_DEV.md) > (`llm_router_api/docs/ENDPOINT_DEV.md`). All endpoints are exposed under the REST API service. Unless stated otherwise, met…","h":["Endpoints Overview","Authentication","Health &amp; Info","Auth‑required Endpoints","Streaming vs. Non‑Streaming Responses","Payload format"],"b":"Endpoints Overview # Creating new endpoints? See the Endpoint Development Guide ( llm_router_api/docs/ENDPOINT_DEV.md ). All endpoints are exposed under the REST API service. Unless stated otherwise, methods are POST and consume/produce JSON. The default API prefix is /api (configurable via LLM_ROUTER_EP_PREFIX ). Endpoints registered with dont_add_api_prefix=True appear without this prefix (e.g. /models instead of /api/models ). Authentication # When LLM_ROUTER_AUTH_ENABLED=true , endpoints are divided into public and auth‑required : Scope Description Env var Public Bypass all auth checks — always accessible LLM_ROUTER_AUTH_PUBLIC_ENDPOINTS Auth‑required Return 401 Unauthorized if no valid API key is provided LLM_ROUTER_AUTH_ENABLED=true Public endpoints: LLM_ROUTER_AUTH_PUBLIC_ENDPOINTS , default /metrics,/health — and for every entry, /v1{entry} as well. The list is compared against the full request path, so a bare entry never matches a prefixed route: endpoints built with dont_add_api_prefix=False live under LLM_ROUTER_EP_PREFIX ( /api by default) and need the prefixed entry ( LLM_ROUTER_AUTH_PUBLIC_ENDPOINTS=\"/metrics,/health,/api/ping\" ). Everything else requires a valid API key with the appropriate policy permission: Permission type What it grants access to chat Chat completions, model listing, responses embedding Embeddings endpoints anthropic Anthropic Messages API ( /v1/messages ) ollama Ollama‑style chat completion builtin Built‑in utility endpoints (translate, generate, etc.) API keys are checked in order of priority: Authorization: Bearer &lt;key&gt; header x-api-key header Query parameters api_key / api-key are rejected (logged as a warning) — use one of the headers Health &amp; Info # Public by default ( LLM_ROUTER_AUTH_PUBLIC_ENDPOINTS=\"/metrics,/health\" ), reachable without a key whatever LLM_ROUTER_AUTH_ENABLED says: GET /health – Router health check → {\"status\": true, \"body\": \"healthy\", \"stream\": false} . GET /metrics – Prometheus metrics (requires LLM_ROUTER_USE_PROMETHEUS=1 ). Registered straight on the app, so the path is literal /metrics — LLM_ROUTER_EP_PREFIX does not apply. Everything else below needs a key once auth is on; the permission in brackets comes from _ENDPOINT_PERMISSION_MAP : GET /models – List OpenAI‑compatible models ( chat ). GET /v1/models – List OpenAI‑compatible models, v1 ( chat ). GET / – Ollama health endpoint → Ollama is running ( chat ). GET /api/ping – Simple health‑check → {\"status\": true, \"body\": \"pong\", \"stream\": false} ( builtin ). GET /api/version – Return the router version → {\"version\": \"&lt;semver&gt;\", \"stream\": false} ( builtin ); this is the path llm_router_lib clients call. GET /api/tags – List available Ollama model tags ( chat ). GET /api/v0/models – List LM Studio models ( chat ). Auth‑required Endpoints # Chat completions # POST /chat/completions — OpenAI‑style chat completion (requires chat permission). POST /api/chat/completions — OpenAI‑style chat completion with prefix (requires chat permission). POST /v1/chat/completions — vLLM‑like chat completion (requires chat permission). POST /api/chat — Ollama‑style chat completion (requires ollama permission). Responses # POST /responses — OpenAI‑like responses endpoint (requires chat permission). POST /v1/responses — OpenAI‑like responses endpoint v1 (requires chat permission). Embeddings # POST /embeddings — Standard embeddings (requires embedding permission). POST /api/embeddings — Standard embeddings with prefix (requires embedding permission). POST /v1/embeddings — OpenAI‑compatible embeddings endpoint (requires embedding permission). POST /api/embed — Ollama‑native embeddings endpoint (requires embedding permission). Anthropic # POST /v1/messages — Anthropic Messages API compatible endpoint (requires anthropic permission). Chat &amp; Completions (Built‑in, requires builtin permission) # POST /api/conversation_with_model — Standard chat endpoint (OpenAI‑compatible payload). POST /api/extended_conversation_with_m…"},{"k":"llm-router-api/docs/endpoint-dev.html","t":"Endpoint Development Guide","s":"REST API","x":"# Endpoint Development Guide This guide explains how to create, configure and extend REST **endpoints (EP)** in `llm_router_api`. > **Related pages:** endpoint catalog — [`../endpoints/README.md`](../endpoints/README.md) · > models config …","h":["1. Class hierarchy and base variants","EndpointI (abstract base)","EndpointWithHttpRequestI (proxy base)","PassthroughI (passthrough base)","2. Execution cycle – what run_ep does in EndpointWithHttpRequestI","3. System prompt and content overrides – how the fields work","SYSTEM_PROMPT_NAME","How overrides are passed – payload keys (current mechanism)","4. Hooks – _prepare_response_function","5. Parameter validation – REQUIRED_ARGS and OPTIONAL_ARGS","6. Endpoint constructor parameters – what they set","7. How to add a new endpoint – step by step","8. Most common usage patterns","8.1 Simple OpenAI‑compatible proxy","8.2 Endpoint with a built‑in system prompt","8.3 Local / static endpoint","9. Quick field reference","10. Practical notes","Worked example: BatchFileSummaries (per‑file summaries)","See Also"],"b":"Endpoint Development Guide # This guide explains how to create, configure and extend REST endpoints (EP) in llm_router_api . Related pages: endpoint catalog — ../endpoints/README.md · models config — MODELS_CONFIG.md · load balancing — LB_STRATEGIES.md · environment variables — ENV_DEFINITIONS.md 1. Class hierarchy and base variants # All endpoint code lives in llm_router_api/endpoints/endpoint_i.py (plus passthrough.py ). The hierarchy is: SecureEndpointI – security scaffolding (masking, guardrails, audit, metrics) └── EndpointI (ABC) – abstract base: API surface + validation, no run_ep └── EndpointWithHttpRequestI (ABC) – full proxy implementation (run_ep, HTTP, streaming) └── PassthroughI (ABC) – \"forward as-is\" base for OpenAI‑compatible endpoints EndpointI (abstract base) # Defines the general API and argument validation, but does not implement run_ep . Abstract methods that every concrete endpoint must implement: run_ep(params) – execute the endpoint logic for a request. prepare_payload(params) – convert raw request parameters into the payload understood by the downstream backend (or the final response body). Class attributes (defaults): METHODS = [\"GET\", \"POST\"] – supported HTTP verbs. REQUIRED_ARGS = [] – parameter names that must be present (see §5). OPTIONAL_ARGS = [] – accepted but optional parameter names (see §5). SYSTEM_PROMPT_NAME = {\"pl\": None, \"en\": None} – per‑language system‑prompt ids (see §3). Useful helpers: _check_required_params(params) – raises ValueError (→ HTTP 400) when a required key is missing. _resolve_prompt_name(params, map_prompt, prompt_str_force, prompt_str_postfix) – builds the final system prompt (see §3). _get_choices_from_response(response) – parses a requests.Response into (json_body, choices, assistant_text) ; handles both OpenAI‑style ( choices ) and Ollama‑style ( message ) bodies. return_response_ok(body) / return_response_not_ok(body) – standardized JSON response envelopes. The constructor validates the endpoint definition at startup: api_types must be non‑empty and intersect the global API_TYPES list ( \"builtin\" , \"openai\" , \"ollama\" , \"lmstudio\" , \"vllm\" , \"anthropic\" — defined in llm_router_api/core/api_types/dispatcher.py ), otherwise RuntimeError . method must be one of METHODS , otherwise ValueError . EndpointWithHttpRequestI (proxy base) # Extends EndpointI with the full outbound‑HTTP implementation: Complete run_ep implementation (see §2 for the cycle). Outbound HTTP via HttpRequestExecutor ( endpoints/httprequest.py ) and retry orchestration via HttpDispatch ( endpoints/http_dispatch.py ). timeout – seconds after which outbound HTTP calls are aborted. Defaults to EXTERNAL_API_TIMEOUT (300 s, env LLM_ROUTER_EXTERNAL_TIMEOUT ). System‑prompt injection into messages and streaming (NDJSON) support. PassthroughI (passthrough base) # Sets REQUIRED_ARGS = None , OPTIONAL_ARGS = None , SYSTEM_PROMPT_NAME = None . prepare_payload(params) returns params or {} unchanged (decorated with @EP.response_time ), so the request is forwarded verbatim. Useful for OpenAI‑compatible endpoints where you simply forward what arrives. Why do the endpoints in endpoints/builtin/openai.py subclass PassthroughI ? OpenAI‑compatible endpoints need minimal logic – just forward the incoming request. PassthroughI removes the boilerplate (no required arguments, no system prompt, ready‑made proxy run_ep ). The concrete classes ( OpenAICompletionHandler , OpenAIResponsesHandler , OpenAIModelsHandler , …) only add a prepare_response_function that normalizes non‑OpenAI responses (Ollama / Anthropic) into the OpenAI shape. 2. Execution cycle – what run_ep does in EndpointWithHttpRequestI # In short, run_ep(params) performs the following steps (implementation: EndpointWithHttpRequestI.run_ep in endpoints/endpoint_i.py ): Start timer – self._start_time = time.time() . Prepare the payload – params = self._prepare_incoming_payload(params) : calls your prepare_payload(params) (endpoint logic) and then runs the config…"},{"k":"llm-router-cli/index.html","t":"llm-router CLI — Command Reference","s":"CLI","x":"# llm-router CLI — Command Reference **Package:** `llm-router` **Entry points:** - `llm-router` — main CLI tool (auth, anonymizer, config, util, server, completion) --- ## Quick Start ```bash pip install llm-router[api] llm-router --help l…","h":["Quick Start","Top-Level Commands","llm-router auth — API Key &amp; Authentication Management","Command Tree","Shared Flags (store-backed subcommands)","Key Management: llm-router auth key &lt;command&gt;","Policy Management: llm-router auth policy &lt;command&gt;","Rate Limit Overrides: llm-router auth rate-limit &lt;command&gt;","llm-router config — Provider Discovery &amp; Config Merging","Command Tree","discover — Scan hosts for local LLM servers","merge — Merge multiple models-config.json files","llm-router anonymizer run — Text Anonymization","llm-router util — Utility Apps (translate / genai-classifier / genai-data-augmentation)","Command Tree","translate — Translate texts in JSON/JSONL datasets","genai-classifier — Classify translated datasets (JSONL only)","genai-data-augmentation — Augment a local JSONL dataset","llm-router server — REST API Server Lifecycle","Reloading: server reload","Instances — running several servers side by side","list — every instance at a glance","stop --all / status --all / rm-instance","status — colored status card","log — follow the server log","llm-router completion — Shell Tab-Completion (bash / zsh)","Seed File (Memory Store)","Seed File Format (ApiKeyRecord fields)","Key Format","See Also"],"b":"llm-router CLI — Command Reference # Package: llm-router Entry points: llm-router — main CLI tool (auth, anonymizer, config, util, server, completion) Quick Start # pip install llm-router [ api ] llm-router --help llm-router --version Top-Level Commands # Command Description auth Manage API keys, policies, and rate limiting config Auto-discover local providers &amp; merge configs anonymizer run Anonymize text using a selectable algorithm util Utility apps: translate , genai-classifier , genai-data-augmentation server Manage the REST API server ( start / stop / status / list ), many instances completion Generate / install shell tab-completion ( bash / zsh ) llm-router auth — API Key &amp; Authentication Management # Command Tree # llm-router auth key &lt;command&gt; # API key lifecycle llm-router auth policy &lt;command&gt; # Policy management llm-router auth rate-limit &lt;command&gt; # Per-key rate limit overrides Shared Flags (store-backed subcommands) # Flag Default Description --store &lt;backend&gt; memory Key store: memory , redis , or vault --auth-redis-host (empty) Auth Redis host --auth-redis-port 6379 Auth Redis port --auth-redis-db 0 Auth Redis database number --auth-redis-password — Auth Redis password --auth-redis-protocol 2 Auth Redis protocol: 2 (RESP2) or 3 (RESP3) --verbose false Enable verbose (DEBUG) logging of internal operations Note: These flags are shared by all store-backed subcommands — key generate , key list , key delete , key disable , key enable , key rotate , policy create , rate-limit apply , and rate-limit remove . They do not apply to the read-only policy list and rate-limit list subcommands. Note: Each --auth-redis-* flag falls back to a matching LLM_ROUTER_AUTH_REDIS_&lt;HOST|PORT|DB|PASSWORD|PROTOCOL&gt; environment variable before the built-in default is used. These auth-specific Redis flags are separate from the general LLM_ROUTER_REDIS_* env vars. Key Management: llm-router auth key &lt;command&gt; # generate — Create a new API key # llm-router auth key generate \\ --policy developer \\ --expires 1750000000 \\ --store memory Flag Default Description --policy developer Policy name to assign --expires None Expiry (Unix timestamp or None ) --output (stdout) Output file path (created with 0600 permissions) Output: sk-llmr-live-&lt;base62&gt; key (plaintext shown once at creation) plus the generated Key ID: (e.g. key-fe8fc388 ) — use the ID with list / delete / disable / enable / rotate . list — List all API keys # llm-router auth key list --store memory [ --json ] Flag Default Description --json false Output in JSON format delete &lt;key-id&gt; — Delete a key permanently # llm-router auth key delete &lt;key-id&gt; --store memory disable &lt;key-id&gt; — Deactivate without deleting # llm-router auth key disable &lt;key-id&gt; [ --store memory ] enable &lt;key-id&gt; — Re-activate a disabled key # llm-router auth key enable &lt;key-id&gt; [ --store memory ] rotate &lt;key-id&gt; — Generate a replacement key # llm-router auth key rotate &lt;key-id&gt; --grace 3600 [ --store memory ] Flag Default Description --grace 3600 Grace period in seconds (old key stays valid) Policy Management: llm-router auth policy &lt;command&gt; # list — List builtin policies # llm-router auth policy list Builtin policies: Policy Access Description developer All Full access to all endpoints admin All Admin access chat Chat Chat completion endpoints embedding Embedding Embedding endpoints anthropic Anthropic Anthropic messages endpoint ollama Ollama Ollama endpoints builtin All builtin Built-in endpoints (translate, etc.) create &lt;name&gt; [&lt;json-policy&gt;] — Create a custom policy # # inline JSON llm-router auth policy create my-team '{ \"can_access\": true, \"rate_limit\": 120, \"model_whitelist\": [\"gpt-4\", \"llama-3\"] }' --store memory # from a file, or from stdin (-) — avoids leaking policy JSON into shell history llm-router auth policy create my-team --file my-team.json cat my-team.json | llm-router auth policy creat…"},{"k":"llm-router-lib/index.html","t":"llmrouterlib","s":"Python library","x":"# llm_router_lib ## Overview `llm_router_lib` bundles **Pydantic data‑model definitions** **and** a **thin, opinionated client wrapper** for the LLM‑Router service. - The **data models** live in `llm_router_lib/data_models` and describe ev…","h":["Overview","Installation","Quick start","Data models","Conversation models","Utility models (selected examples)","Services (low‑level wrappers)","Example: using a service directly","Thin client wrapper (LLMRouterClient)","Async client (AsyncLLMRouterClient)","Streaming","Utilities","Development &amp; testing"],"b":"llm_router_lib # Overview # llm_router_lib bundles Pydantic data‑model definitions and a thin, opinionated client wrapper for the LLM‑Router service. The data models live in llm_router_lib/data_models and describe every request payload the router accepts. The client ( LLMRouterClient ) offers a high‑level, Pythonic API that hides HTTP details, retries, and error handling. The async client ( AsyncLLMRouterClient , httpx ‑based) offers the same API for asyncio applications (e.g. FastAPI), plus streaming conversation methods. Low‑level service classes ( ConversationWithModelService , ExtendedConversationWithModelService , TranslateService , GenerativeAnswerService , health services) perform the actual HTTP calls and can be used directly when finer‑grained control is required. HttpRequester (in utils/http.py ) is a small wrapper around requests that adds logging, configurable retries, and unified error translation. AsyncHttpRequester (in utils/http_async.py ) is its httpx ‑based asynchronous counterpart with the same retry and error‑translation contract (plus a stream() helper for SSE responses). A dedicated exception hierarchy ( exceptions.py ) maps HTTP errors to meaningful Python exceptions. In short, llm_router_lib provides both the contract (the “schema”) and a convenient client to consume the router service. Installation # The library targets Python 3.10.6 and uses a virtualenv . Install it in editable mode for development: # Clone the repository (if you haven't already) git clone https://github.com/radlab-dev-group/llm-router.git cd llm-router/llm_router_lib # Create and activate a virtual environment python3 -m venv .venv source .venv/bin/activate # Install the package and its dependencies pip install -e . All runtime dependencies ( requests , pydantic , plus the packages listed in requirements.txt ) are declared in the project’s requirements.txt . Quick start # from llm_router_lib import LLMRouterClient # Initialise the client – point it at the router’s host (do **not** include the `/api` prefix) client = LLMRouterClient ( api = \"http://localhost:8080\" , # router host URL token = \"YOUR_ROUTER_TOKEN\" , # optional, if router requires auth ) # Call the standard conversation endpoint with named keyword arguments – # the client builds the Pydantic request model for you. response = client . conversation_with_model ( user_last_statement = \"Hello, how are you?\" , model = \"google/gemma-3-12b-it\" , temperature = 0.7 , max_new_tokens = 128 , ) # response is a typed `ConversationResponse` model: print ( response . response ) # → the assistant's reply text print ( response . generation_time ) # → seconds taken by the server You can also pass a ready‑made pydantic request model via the payload keyword (all endpoint parameters are keyword‑only – there are no positional arguments): from llm_router_lib.data_models.builtin_chat import ConversationWithModelRequest payload = ConversationWithModelRequest ( model_name = \"google/gemma-3-12b-it\" , user_last_statement = \"Hello, how are you?\" , temperature = 0.7 , max_new_tokens = 128 , ) response = client . conversation_with_model ( payload = payload ) Note: raw dict payloads are no longer accepted – passing a dict as payload raises TypeError . Build the matching Pydantic model explicitly (e.g. ConversationWithModelRequest(**dict_payload) ) or use the named keyword arguments. Calling a generation method with neither payload nor enough named arguments raises NoArgsAndNoPayloadError . Data models # All request payloads are defined in llm_router_lib/data_models . A common base class supplies shared options: class BaseModelOptions ( BaseModel ): \"\"\"Options shared across many endpoint models.\"\"\" mask_payload : bool = False masker_pipeline : Optional [ List [ str ]] = None Conversation models # Model Required fields Optional / extra fields ConversationWithModelRequest model_name , user_last_statement temperature , max_new_tokens , historical_messages , … ExtendedConversationWithModelRequest All of the…"},{"k":"llm-router-lib/response-models.html","t":"Response models","s":"Python library","x":"# Response models ## Introduction The router speaks JSON, and every endpoint emits a *different* shape. Before these models, `LLMRouterClient` handed that JSON straight back as a raw `dict`, so the knowledge of each response schema lived o…","h":["Introduction","📦 Where the models live and how to import them","🗺️ Client method → response model","🧱 Class hierarchy","📚 Model reference","Health / meta","Conversation","Per‑text (list) utilities","Single‑output utilities","Article utilities","🎨 Design decisions","🚀 Usage examples","Access fields (typed)","Serialize / export","Round‑trip (validate a stored/status‑wrapped body)","🔁 Backward compatibility","➕ Adding a new endpoint model","✅ Testing notes","🧾 File map"],"b":"Response models # Introduction # The router speaks JSON, and every endpoint emits a different shape. Before these models, LLMRouterClient handed that JSON straight back as a raw dict , so the knowledge of each response schema lived on the caller side: you had to remember the exact key names, guard against missing fields, and you got no compile‑time or IDE help. The response models move that contract into the library. Every client method returns a typed Pydantic model that mirrors the JSON body the router emits for that endpoint, and the models are defined in one place — data_models/response.py — next to the existing request models ( TranslateModel , Polarity3cModel , …). Together, request response describe the full contract of an endpoint. Concretely, the typing is there to do three things: Fix the contract in one place. response.py is the single definition of what each endpoint returns. Change the router → change the model → type checkers and tests flag every consumer that still relied on the old shape. Validate at the boundary. Pydantic checks the body as it comes in, so a malformed or partial response fails there with a ValidationError instead of as a KeyError a few frames deep in your code. Expose static types. Because the return annotation is a class rather than Dict[str, Any] , mypy and IDEs can resolve resp.response[0].translated . The model also exports a JSON Schema via model_json_schema() for docs, codegen, or contract tests. # resp is a typed model, not a dict resp = client . translate ( texts = [ \"Hello world\" ], model = \"speakleash/Bielik-11B-v2.3-Instruct\" ) resp . response [ 0 ] . translated # str resp . generation_time # Optional[float] resp . model_dump () # → plain dict, if you still need one 📦 Where the models live and how to import them # All response models are exported from the package root, so either import style works: # Option 1 – from the package (recommended) from llm_router_lib.data_models import ( TranslateResponse , Polarity3cResponse , ConversationResponse , ModelsListResponse , # …see the full list in data_models/__init__.py ) # Option 2 – from the module directly from llm_router_lib.data_models.response import TranslateResponse They sit alongside (but are independent of) the request models such as TranslateModel , Polarity3cModel , etc. 🗺️ Client method → response model # The table below maps every LLMRouterClient method to the model it now returns. response is the endpoint‑specific payload; generation_time (seconds) is present on all generation endpoints. Client method Endpoint Returns response payload ping() GET /api/ping PingResponse status: bool , body: str version() GET /api/version VersionResponse version: str models() GET /v1/models ModelsListResponse object: str , data: List[ModelInfo] conversation_with_model(payload) POST /api/conversation_with_model ConversationResponse str (assistant reply) extended_conversation_with_model(payload) POST /api/extended_conversation_with_model ExtendedConversationResponse str (assistant reply) polarity_3c(...) POST /api/polarity_3c Polarity3cResponse List[Polarity3cItem] translate(...) POST /api/translate TranslateResponse List[TranslateItem] simplify_text(...) POST /api/simplify_text SimplifyTextResponse List[str] generative_answer(...) POST /api/generative_answer GenerativeAnswerResponse str (answer) generate_article_from_text(...) POST /api/generate_article_from_text GenerateArticleFromTextResponse ArticleText generate_article_from_texts(...) POST /api/generate_article_from_texts GenerateArticleFromTextsResponse ArticleText create_full_article_from_texts(...) POST /api/create_full_article_from_texts CreateFullArticleFromTextsResponse ArticleText generate_questions(...) POST /api/generate_questions GenerateQuestionsResponse List[TextQuestions] generate_label(...) POST /api/generate_label GenerateLabelResponse str (label) 🧱 Class hierarchy # BaseResponse # extra keys ignored (pydantic default) ├── PingResponse ├── VersionResponse ├── ModelInfo ├── Mod…"},{"k":"helm-charts/index.html","t":"LLM‑Router Helm Chart","s":"Deployment","x":"# LLM‑Router Helm Chart Quick summary – This Helm chart deploys the LLM‑Router application with its optional Redis dependency. It works on any Kubernetes 1.19+ cluster and can be customized via values.yaml, environment variables, or the --…","h":["Table of Contents","Prerequisites","Chart layout","Dependencies","Installing the chart","Customising the deployment","1️⃣ Using a custom values file"],"b":"LLM‑Router Helm Chart # Quick summary – This Helm chart deploys the LLM‑Router application with its optional Redis dependency. It works on any Kubernetes 1.19+ cluster and can be customized via values.yaml, environment variables, or the --set flag. Table of Contents # Prerequisites Chart layout Dependencies Installing the chart Customising the deployment Prerequisites # Requirement Why we need it Kubernetes (v1.19 or newer) The chart creates Deployments, Services, Ingresses, ConfigMaps, … Helm (v3.x) Used to render and apply the chart Access to a container registry (e.g., Docker Hub, Quay, your private registry) The chart pulls the image defined in llm-router``values.yaml (Optional) cert‑manager If you enable TLS for the Ingress, cert‑manager will provision certificates Chart layout # helm_charts/ └─ llm-router/ ├─ Chart.yaml # Chart metadata ├─ values.yaml # Default values ├─ values-dev.yaml # Development‑specific overrides ├─ templates/ │ ├─ _helpers.tpl # Helper functions (name, labels, etc.) │ ├─ deployment.yaml # Deployment definition │ ├─ service.yaml # Service definition │ ├─ ingress.yaml # Ingress definition (optional) │ ├─ configmap.yaml # ConfigMap for runtime env vars │ └─ configmap-models.yaml # ConfigMap for `models-config.json` └─ charts/ └─ redis-23.2.12.tgz # Bitnami Redis sub‑chart (dependency) Dependencies # The chart depends on Redis. The dependency is declared in Chart.yaml and pulled automatically when you run helm dependency update. # Chart.yaml dependencies : - name : redis version : \"23.2.12\" repository : \"oci://registry-1.docker.io/bitnamicharts\" condition : redis.enabled Adding / updating dependencies # From the chart root (helm_charts/llm-router) helm dependency update . You can disable the Redis sub‑chart with: --set redis.enabled = false Installing the chart # helm upgrade --install my-llm-router ./helm_charts/llm-router \\ --namespace my-namespace \\ --create-namespace Customising the deployment # You can tailor the LLM‑Router Helm chart to your environment in three different ways. Pick the approach that best fits the task at hand. Method When to use Edit values.yaml and run helm upgrade Long‑term, version‑controlled configuration that lives in source control. Pass --set flags on the command line One‑off tweaks, CI pipelines, quick experiments, or when you need to override just a handful of values. Provide a custom values file ( -f my-values.yaml ) Complex overrides, reusable profiles, or when you prefer to keep the changes in a separate, shareable file. Below are concrete examples for the two most common scenarios: using a custom values file and using --set flags. 1️⃣ Using a custom values file # Create a file (e.g. my-values.yaml ) with the settings you want to override: # my-values.yaml ingress : enabled : true className : traefik hosts : - host : my-custom.cluster.local paths : - path : / pathType : Prefix tls : - hosts : - my-custom.cluster.local secretName : llm-router-tls image : tag : latest # pull the latest container image web_host : my-custom.cluster.local # the hostname used by the app &amp; Ingress Deploy (or upgrade) the chart with this file: ```shell script helm upgrade --install llm-router ./helm_charts/llm-router \\ -f my-values.yaml --- ### 2️⃣ Using `--set` flags If you only need to change a few values, the `--set` syntax is handy: ```shell script helm upgrade --install llm-router ./helm_charts/llm-router \\ --set ingress.enabled=true \\ --set web_host=llm2.k3s.radlab.dev \\ --set image.tag=latest Flag Effect ingress.enabled=true Turns on the Ingress resource (otherwise it’s omitted). llm-1.cluster.local Sets the host name used both in the Ingress rule and the application’s configuration ( LLM_ROUTER_WEB_HOST ). image.tag=latest Pulls the latest image tag instead of the default chart‑version tag."},{"k":"examples/index.html","t":"Integration Examples with LLM Router","s":"Examples","x":"# Integration Examples with LLM Router This directory contains example boilerplates that demonstrate how easy it is to integrate popular LLM libraries with the router by simply switching the host. --- ## Available Examples - **[LlamaIndex]…","h":["Available Examples","Core Principle","Quick Start","Example Structure","Full Stack with Local Models","Additional Information"],"b":"Integration Examples with LLM Router # This directory contains example boilerplates that demonstrate how easy it is to integrate popular LLM libraries with the router by simply switching the host. Available Examples # LlamaIndex – Integration with LlamaIndex (GPT Index) Using LlamaIndex with Local Models – Guide for mapping OpenAI model names to local models via the router. LangChain – Integration with LangChain OpenAI SDK – Direct integration with the OpenAI Python SDK LiteLLM – Integration with LiteLLM Haystack – Integration with Haystack Core Principle # All examples work on the same principle: just change base_url / api_base to the address of your router , and the router will automatically: ✅ Distribute traffic among available providers ✅ Perform load balancing ✅ Provide health checking ✅ Supply monitoring and metrics ✅ Handle streaming and non‑streaming responses Quick Start # Each example can be run directly: ```shell script LlamaIndex # python examples/llamaindex_example.py LangChain # python examples/langchain_example.py OpenAI SDK # python examples/openai_example.py LiteLLM # python examples/litellm_example.py Haystack # python examples/haystack_example.py ``` Example Structure # Each example includes: Basic configuration – how to point the library at the router Streaming – handling streaming responses Non‑streaming – handling full responses Error handling – managing errors Full Stack with Local Models # The quick‑start guides for running the full stack with local models are included in the repository: Gemma 3 12B‑IT – README Bielik 11B‑v2.3‑Instruct – README These guides walk you through: Installing vLLM and the respective model. Setting up LLM‑Router with the provided models-config.json . Testing the end‑to‑end flow (router → vLLM). Follow the linked README files for step‑by‑step instructions to launch a complete stack locally. Additional Information # Learn more about the router: Main README API Documentation Endpoints Overview Load‑Balancing Strategies"},{"k":"examples/quickstart/index.html","t":"Full Stack with Local Models","s":"Examples","x":"## Full Stack with Local Models The quick‑start guides for running the full stack with **local models** are included in the repository: - **Gemma 3 12B‑IT** – [README](google-gemma3-12b-it/README.md) - **Bielik 11B‑v2.3‑Instruct** – [READM…","h":["Full Stack with Local Models"],"b":"Full Stack with Local Models # The quick‑start guides for running the full stack with local models are included in the repository: Gemma 3 12B‑IT – README Bielik 11B‑v2.3‑Instruct – README"},{"k":"examples/readme-llamaindex.html","t":"Using LlamaIndex with Local Models via the LLM‑Router","s":"Examples","x":"## Using LlamaIndex with Local Models via the LLM‑Router LlamaIndex’s `OpenAI` wrapper expects **OpenAI‑style model names** (e.g. `gpt-3.5-turbo`, `gpt-4`). When you want to run *local* models (Gemma, Ollama, vLLM, etc.) behind an LLM‑Rout…","h":["Using LlamaIndex with Local Models via the LLM‑Router","Why the mapping is required","Router configuration","How it works end‑to‑end","Checklist for a working setup","TL;DR"],"b":"Using LlamaIndex with Local Models via the LLM‑Router # LlamaIndex’s OpenAI wrapper expects OpenAI‑style model names (e.g. gpt-3.5-turbo , gpt-4 ). When you want to run local models (Gemma, Ollama, vLLM, etc.) behind an LLM‑Router, the router must translate those OpenAI names to the actual model identifiers used by the local providers. Why the mapping is required # LlamaIndex validates the model name against a known list of OpenAI models to decide whether the call is a chat request, to obtain context‑window size, etc. If the wrapper receives a name it does not recognise (e.g. google/gemma-3-12b-it ), it raises a ValueError like: ValueError: Unknown model 'google/gemma-3-12b-it'. Please provide a valid OpenAI model name … By keeping the OpenAI name in the request and letting the router forward the request to the appropriate backend, you get the best of both worlds: LlamaIndex works unchanged, and the traffic is routed to your local model. Router configuration # The router’s configuration must contain a section that lists OpenAI model names ( openai_models ). Each entry defines one or more providers that actually serve the model. The key points are: Field Meaning id Arbitrary identifier for the provider instance. api_host URL where the provider’s inference server is reachable. api_type The protocol the provider uses ( vllm , ollama , …). model_path The local model identifier (e.g. google/gemma-3-12b-it ). weight Relative load‑balancing weight when several providers are listed. Example snippet # \"openai_models\" : { (...) \"gpt-3.5-turbo\" : { \"providers\" : [ { \"id\" : \"gpt_35_turbo-gemma3_12b-vllm-71:7000\" , \"api_host\" : \"http://192.168.100.71:7000/\" , \"api_token\" : \"\" , \"api_type\" : \"vllm\" , \"input_size\" : 4096 , \"model_path\" : \"google/gemma-3-12b-it\" , \"weight\" : 1.0 }, { \"id\" : \"gpt_35_turbo-gemma3_12b-vllm-71:7001\" , \"api_host\" : \"http://192.168.100.71:7001/\" , \"api_token\" : \"\" , \"api_type\" : \"vllm\" , \"input_size\" : 4096 , \"model_path\" : \"google/gemma-3-12b-it\" , \"weight\" : 1.0 } ] }, \"gpt-4\" : { \"providers\" : [ { \"id\" : \"gpt-4-gpt-oss-20b-ollama-66:11434\" , \"api_host\" : \"http://192.168.100.66:11434\" , \"api_token\" : \"\" , \"api_type\" : \"ollama\" , \"input_size\" : 256000 , \"model_path\" : \"gpt-oss:120b\" } ] } }, ... \"active_models\" : { \"openai_models\" : [ \"gpt-3.5-turbo\" , \"gpt-4\" ... ] } How it works end‑to‑end # LlamaIndex code creates an OpenAI client with a model name that the wrapper knows, e.g.: llm = OpenAI ( model = \"gpt-3.5-turbo\" , api_base = \"http://localhost:8080\" , api_key = \"not-needed\" ) The request is sent to the router ( api_base ). The router looks up \"gpt-3.5-turbo\" in openai_models , selects a provider, and forwards the request to the provider’s api_host . The provider receives the request with model_path=\"google/gemma-3-12b-it\" (or \"gpt-oss:120b\" for gpt-4 ) and runs the local model. The response travels back through the router to LlamaIndex, which treats it exactly like an OpenAI response. Checklist for a working setup # Router is running and reachable at the URL you pass to api_base . openai_models section contains all OpenAI names you intend to use from LlamaIndex. Each OpenAI name maps to at least one provider with the correct model_path . The active_models.openai_models list includes the names you want to expose (otherwise the router will ignore them). Local model servers (vLLM, Ollama, etc.) are up and listening on the api_host URLs specified. TL;DR # *Keep the model names you give to LlamaIndex identical to the OpenAI names defined in the router configuration. The router will translate those names to the real local model identifiers ( model_path ). This mapping is the only thing needed for LlamaIndex to work seamlessly with any self�"},{"k":"examples/quickstart/google-gemma3-12b-it/vllm.html","t":"vLLM + google/gemma‑3‑12b‑it – Quick‑Start Guide (Ubuntu)","s":"Examples","x":"# vLLM + `google/gemma‑3‑12b‑it` – Quick‑Start Guide (Ubuntu) > **Prerequisites** > - Ubuntu 20.04 or newer > - Python 3.10 (our project uses 3.10.6) > - `virtualenv` (installed) > - CUDA 11.8 + GPU **or** a CPU‑only setup --- ## 1️⃣ Creat…","h":["1️⃣ Create &amp; activate a virtual environment","2️⃣ Install vLLM","Verify the installation","3️⃣ Obtain the model google/gemma-3-12b-it","4️⃣ Run the vLLM server","5️⃣ Test the endpoint","6️⃣ Handy tips","🎉 All set!"],"b":"vLLM + google/gemma‑3‑12b‑it – Quick‑Start Guide (Ubuntu) # Prerequisites - Ubuntu 20.04 or newer - Python 3.10 (our project uses 3.10.6) - virtualenv (installed) - CUDA 11.8 + GPU or a CPU‑only setup 1️⃣ Create &amp; activate a virtual environment # mkdir -p ~/vllm-gemma &amp;&amp; cd ~/vllm-gemma python3 -m venv .venv source .venv/bin/activate This creates an optional project directory, sets up a Python virtual environment in .venv , and activates it (you’ll see (.venv) in the prompt). 2️⃣ Install vLLM # pip install --upgrade pip pip install \"vllm[cuda]\" The above installs the latest pip and then installs vLLM with GPU support (the appropriate CUDA libraries are detected automatically). If you have no GPU, install the CPU‑only version instead: pip install vllm[cpu] . Verify the installation # python -c \"import vllm; print(vllm.__version__)\" You should see a version string such as 0.11.2 . 3️⃣ Obtain the model google/gemma-3-12b-it # mkdir -p ./google/gemma-3-12b-it pip install huggingface_hub hf download google/gemma-3-12b-it \\ --local-dir ./google/gemma-3-12b-it This creates a folder for the model, installs the Hugging Face CLI, and downloads the model files into ./google/gemma-3-12b-it . The files are cached under ~/.cache/huggingface/hub by default; you can keep a local copy to avoid re‑downloads. 4️⃣ Run the vLLM server # Copy the ready‑to‑use Bash script ( llm-router/examples/quickstart/google-gemma3-12b-it/run-gemma-3-12b-it-vllm.sh ) to directory wgen the vLLM will be started with the Gemma 3 model. cp path/to/llm-router/examples/quickstart/google-gemma3-12b-it/run-gemma-3-12b-it-vllm.sh . bash run-gemma-3-12b-it-vllm.sh Tip: Run the server inside a tmux or screen session so it stays alive even if you disconnect from the terminal. 5️⃣ Test the endpoint # INFO : curl and jq are system utilities. curl http://localhost:7000/v1/chat/completions \\ -H \"Content-Type: application/json\" \\ -d '{ \"model\": \"google/gemma-3-12b-it\", \"messages\": [{\"role\": \"user\", \"content\": \"Hello, how are you?\"}], \"max_tokens\": 100 }' | jq You should receive a JSON response containing the model’s generated text, for example: { \"id\" : \"chatcmpl-e30bed0db9f9440a8aec14bd287ca63d\" , \"object\" : \"chat.completion\" , \"created\" : 1764516430 , \"model\" : \"google/gemma-3-12b-it\" , \"choices\" : [ { \"index\" : 0 , \"message\" : { \"role\" : \"assistant\" , \"content\" : \"Hello! I'm doing well, thank you for asking! As an AI, I don't experience feelings like humans do, but everything is running smoothly and I'm ready to chat. 😊\\n\\nHow are *you* doing today?\" , \"refusal\" : null , \"annotations\" : null , \"audio\" : null , \"function_call\" : null , \"tool_calls\" : [], \"reasoning\" : null , \"reasoning_content\" : null }, \"logprobs\" : null , \"finish_reason\" : \"stop\" , \"stop_reason\" : 106 , \"token_ids\" : null } ], \"service_tier\" : null , \"system_fingerprint\" : null , \"usage\" : { \"prompt_tokens\" : 15 , \"total_tokens\" : 66 , \"completion_tokens\" : 51 , \"prompt_tokens_details\" : null }, \"prompt_logprobs\" : null , \"prompt_token_ids\" : null , \"kv_transfer_params\" : null } 6️⃣ Handy tips # Topic Recommendation Memory google/gemma‑3‑12b‑it needs ~24GB VRAM. Use --cpu-offload (if supported) for larger models or when GPU memory is limited. Cache location Set HF_HOME=$PWD/.cache/huggingface to keep all model files inside the project directory. Parallelism Export TOKENIZERS_PARALLELISM=false to silence tokenizer warnings. GPU selection export CUDA_VISIBLE_DEVICES=0 (or another index) when multiple GPUs are present. Update pip install -U vllm refreshes the library; the next server start will pull newer model files if available. Deactivate When done, simply run deactivate to leave the virtual environment. 🎉 All set! # You now have a fully functional OpenAI‑compatible API powered by vLLM and the google/gemma‑3‑12b‑it model."},{"k":"examples/quickstart/speakleash-bielik-11b-v2-3-instruct/vllm.html","t":"vLLM + speakleash/Bielik-11B-v2.3-Instruct – Przewodnik Szybkiego Startu (Ubuntu)","s":"Examples","x":"# vLLM + `speakleash/Bielik-11B-v2.3-Instruct` – Przewodnik Szybkiego Startu (Ubuntu) > **Wymagania wstępne** > - Ubuntu 20.04 lub nowszy > - Python 3.10 (w projekcie używamy 3.10.6) > - `virtualenv` (zainstalowany) > - CUDA 11.8 + GPU **l…","h":["1️⃣ Utwórz i aktywuj wirtualne środowisko","2️⃣ Zainstaluj vLLM","Sprawdź instalację","4️⃣ Przygotuj środowisko do pobierania modelu","6️⃣ Pobierz model speakleash/Bielik-11B-v2.3-Instruct","(Opcjonalnie) Ustaw własny katalog cache","7️⃣ Uruchom serwer vLLM","8️⃣ Przetestuj endpoint","9️⃣ Przydatne wskazówki","🎉 Gotowe!"],"b":"vLLM + speakleash/Bielik-11B-v2.3-Instruct – Przewodnik Szybkiego Startu (Ubuntu) # Wymagania wstępne - Ubuntu 20.04 lub nowszy - Python 3.10 (w projekcie używamy 3.10.6) - virtualenv (zainstalowany) - CUDA 11.8 + GPU lub środowisko tylko CPU 1️⃣ Utwórz i aktywuj wirtualne środowisko # mkdir -p ~/bielik &amp;&amp; cd ~/bielik python3 -m venv .venv source .venv/bin/activate Powoduje to utworzenie katalogu projektu, przygotowanie wirtualnego środowiska w folderze .venv oraz jego aktywację (w promptcie pojawi się (.venv) ). 2️⃣ Zainstaluj vLLM # pip install --upgrade pip pip install \"vllm[cuda]\" Instalacja najnowszej wersji vLLM z obsługą GPU (CUDA zostanie wykryte automatycznie). Jeśli nie masz GPU, użyj wersji CPU: pip install vllm[cpu] . Sprawdź instalację # python -c \"import vllm; print(vllm.__version__)\" Powinieneś zobaczyć wersję, np. 0.11.2 . 4️⃣ Przygotuj środowisko do pobierania modelu # pip install huggingface_hub 6️⃣ Pobierz model speakleash/Bielik-11B-v2.3-Instruct # mkdir -p ./speakleash/Bielik-11B-v2.3-Instruct hf download speakleash/Bielik-11B-v2.3-Instruct \\ --local-dir ./speakleash/Bielik-11B-v2.3-Instruct Model zostanie pobrany do wskazanego katalogu. Pliki będą także buforowane domyślnie w ~/.cache/huggingface/hub . (Opcjonalnie) Ustaw własny katalog cache # Jeśli chcesz, aby wszystkie modele były przechowywane wewnątrz projektu, ustaw zmienną przed pobraniem: export HF_HOME=$PWD/.cache/huggingface # np. ./bielik/.cache/huggingface 7️⃣ Uruchom serwer vLLM # Skopiuj gotowy skrypt Bash (przykładowa ścieżka – dostosuj do swojego projektu): cp path/to/llm-router/examples/quickstart/speakleash-bielik-11b-v2_3-Instruct/run-bielik-11b-v2_3-vllm.sh . bash run-bielik-11b-v2_3-vllm.sh Wskazówka: uruchom serwer w sesji tmux lub screen , aby pozostawał aktywny po rozłączeniu się z terminalem. 8️⃣ Przetestuj endpoint # INFO : curl i jq to narzędzia systemowe. curl http://localhost:7000/v1/chat/completions \\ -H \"Content-Type: application/json\" \\ -d '{ \"model\": \"speakleash/Bielik-11B-v2.3-Instruct\", \"messages\": [{\"role\": \"user\", \"content\": \"Cześć, jak się masz?\"}], \"max_tokens\": 100 }' | jq Powinieneś otrzymać odpowiedź w formacie JSON, np.: { \"id\" : \"chatcmpl-xxxx\" , \"object\" : \"chat.completion\" , \"created\" : 1764516430 , \"model\" : \"speakleash/Bielik-11B-v2.3-Instruct\" , \"choices\" : [ { \"index\" : 0 , \"message\" : { \"role\" : \"assistant\" , \"content\" : \"Cześć! Jestem w pełni sprawny i gotowy do rozmowy. Jak mogę Ci pomóc?\" }, \"finish_reason\" : \"stop\" } ], \"usage\" : { \"prompt_tokens\" : 15 , \"total_tokens\" : 66 , \"completion_tokens\" : 51 } } 9️⃣ Przydatne wskazówki # Temat Rekomendacja Pamięć speakleash/Bielik-11B-v2.3-Instruct potrzebuje ok. 24GB VRAM. Użyj --cpu-offload (jeśli wspierane) przy ograniczonej pamięci GPU. Lokalizacja cache Ustaw HF_HOME=$PWD/.cache/huggingface , aby wszystkie pliki modelu znajdowały się w katalogu projektu. Równoległość tokenizera export TOKENIZERS_PARALLELISM=false wyciszy ostrzeżenia tokenizera. Wybór GPU export CUDA_VISIBLE_DEVICES=0 (lub inny indeks) przy wielu kartach GPU. Aktualizacja pip install -U vllm odświeża bibliotekę; przy następnym uruchomieniu serwera zostaną pobrane nowsze pliki modelu, jeśli są dostępne. Dezaktywacja Po zakończeniu pracy wystarczy wpisać deactivate , aby opuścić wirtualne środowisko. 🎉 Gotowe! # Masz już w pełni działające API kompatybilne z OpenAI, oparte na vLLM i modelu speakleash/Bielik-11B-v2.3-Instruct ."},{"k":"examples/quickstart/speakleash-bielik-11b-v2-3-instruct/index.html","t":"🚀 Przewodnik Szybkiego Startu dla speakleash/Bielik-11B-v2.3-Instruct z vLLM & LLM‑Router","s":"Examples","x":"# 🚀 **Przewodnik Szybkiego Startu** dla `speakleash/Bielik-11B-v2.3-Instruct` z **vLLM** & **LLM‑Router** Ten przewodnik prowadzi Cię krok po kroku przez: 1. **Instalację vLLM** i modelu `speakleash/Bielik-11B-v2.3-Instruct`. 2. **Instalac…","h":["📋 Wymagania wstępne","1️⃣ Utworzenie i aktywacja wirtualnego środowiska","6️⃣ Przygotowanie konfiguracji routera","7️⃣ Test pełnego stosu (router → vLLM)","🎉 Co dalej?"],"b":"🚀 Przewodnik Szybkiego Startu dla speakleash/Bielik-11B-v2.3-Instruct z vLLM &amp; LLM‑Router # Ten przewodnik prowadzi Cię krok po kroku przez: Instalację vLLM i modelu speakleash/Bielik-11B-v2.3-Instruct . Instalację LLM‑Router (bramki API). Uruchomienie routera z konfiguracją modeli dostarczoną w models-config.json . Wszystkie polecenia zakładają, że pracujesz na systemie Unix‑like (Linux/macOS) z Python 3.10.6 , virtualenv oraz ( opcjonalnie) kartą GPU obsługującą CUDA 11.8. 📋 Wymagania wstępne # Wymaganie Szczegóły OS Ubuntu 20.04 + (lub dowolna nowsza dystrybucja Linux/macOS) Python 3.10.6 (domyślna wersja projektu) GPU CUDA 11.8 + (minimum 12 GB VRAM) lub środowisko CPU‑only Narzędzia git , curl , jq (opcjonalnie, przydatne do testowania) Sieć Dostęp do PyPI oraz Hugging Face w celu pobrania modelu 1️⃣ Utworzenie i aktywacja wirtualnego środowiska # ```shell script (opcjonalnie) utwórz katalog demo i przejdź do niego # mkdir -p ~/bielik-demo &amp;&amp; cd $_ Inicjalizacja venv # python3 -m venv .venv source .venv/bin/activate Aktualizacja pip (zawsze dobry pomysł) # pip install --upgrade pip --- ## 2️⃣ Instalacja **vLLM** oraz pobranie modelu Bielik &gt; Pełną instrukcję znajdziesz w pliku [`VLLM.md`](./VLLM.md). --- ## 3️⃣ **Uruchomienie serwera vLLM** Skopiuj do bieżącego katalogu dostarczony skrypt Bash (dostosuj ścieżkę, jeśli potrzebujesz) i uruchom go: ```shell script cp path/to/llm-router/examples/quickstart/speakleash-bielik-11b-v2_3-Instruct/run-bielik-11b-v2_3-vllm.sh . chmod +x run-bielik-11b-v2_3-vllm.sh # Uruchom (warto w tmux/screen) ./run-bielik-11b-v2_3-vllm.sh Serwer nasłuchuje na http://0.0.0.0:7000 i udostępnia endpoint zgodny z OpenAI pod /v1/chat/completions . Możesz szybko go przetestować: ```shell script curl http://localhost:7000/v1/chat/completions \\ -H \"Content-Type: application/json\" \\ -d '{ \"model\": \"speakleash/Bielik-11B-v2.3-Instruct\", \"messages\": [{\"role\": \"user\", \"content\": \"Cześć, jak się masz?\"}], \"max_tokens\": 100 }' | jq Powinieneś otrzymać odpowiedź w formacie JSON. --- ## 4️⃣ Instalacja **LLM‑Router** ```shell script # Sklonuj repozytorium (jeśli jeszcze go nie masz) git clone https://github.com/radlab-dev-group/llm-router.git cd llm-router # Instalacja core + API (w tym samym venv) pip install .[api] # (Opcjonalnie) wsparcie dla Prometheus pip install .[api,metrics] Uwaga: Router używa tego samego wirtualnego środowiska, które utworzyłeś wcześniej, więc wszystkie zależności pozostają odizolowane. 6️⃣ Przygotowanie konfiguracji routera # Plik models-config.json znajdujący się w katalogu speakleash‑bielik już zawiera definicję naszego modelu: { \"speakleash_models\" : { \"speakleash/Bielik-11B-v2.3-Instruct\" : { \"providers\" : [ { \"id\" : \"bielik-11B_v2_3-vllm-local:7000\" , \"api_host\" : \"http://localhost:7000/\" , \"api_type\" : \"vllm\" , \"input_size\" : 56000 , \"weight\" : 1.0 } ] } }, \"active_models\" : { \"speakleash_models\" : [ \"speakleash/Bielik-11B-v2.3-Instruct\" ] } } Skopiuj go (lub przenieś) do katalogu resources/configs/ routera: ```shell script mkdir -p resources/configs cp path/to/speakleash-bielik/models-config.json resources/configs/ --- ## 6️⃣ Uruchomienie **LLM‑Router** ### Lokalny Gunicorn W repozytorium znajduje się pomocniczy skrypt `run-rest-api-gunicorn.sh`. Upewnij się, że jest wykonywalny, a następnie go uruchom: ```shell script chmod +x run-rest-api-gunicorn.sh ./run-rest-api-gunicorn.sh Domyślne zmienne środowiskowe (można zmienić w skrypcie): Zmienna Domyślna wartość Opis LLM_ROUTER_SERVER_TYPE gunicorn Backend serwera LLM_ROUTER_SERVER_PORT 8080 Port nasłuchiwania routera LLM_ROUTER_MODELS_CONFIG resources/configs/models-config.json Ścieżka do pliku konfiguracyjnego LLM_ROUTER_USE_PROMETHEUS 1 (jeśli zainstalowano metrics ) Włącza endpoint /api/metrics Router będzie dostępny pod http://0.0.0.0:8080/api . Pełna lista dostępnych zmiennych środowiskowych znajduje się w opisie zmiennych środowiskowych 7️⃣ Test pełnego stosu (router → vLLM) # ```shell script curl http://loc…"},{"k":"examples/quickstart/google-gemma3-12b-it/index.html","t":"🚀 Quick‑Start Guide for google/gemma-3-12b‑it with vLLM & LLM‑Router","s":"Examples","x":"# 🚀 Quick‑Start Guide for `google/gemma-3-12b‑it` with **vLLM** & **LLM‑Router** This guide walks you through: 1. **Installing vLLM** and the `google/gemma‑3‑12b‑it` model. 2. **Installing LLM‑Router** (the API gateway). 3. **Running the r…","h":["📋 Prerequisites","1️⃣ Set up a virtual environment","5️⃣ Prepare the router configuration","7️⃣ Test the full stack (router → vLLM)","🎉 What’s next?"],"b":"🚀 Quick‑Start Guide for google/gemma-3-12b‑it with vLLM &amp; LLM‑Router # This guide walks you through: Installing vLLM and the google/gemma‑3‑12b‑it model. Installing LLM‑Router (the API gateway). Running the router with the model configuration provided in models-config.json . All commands assume you are working on a Unix‑like system (Linux/macOS) with Python 3.10.6 and virtualenv available. 📋 Prerequisites # Requirement Details OS Ubuntu 20.04 + (or any recent Linux/macOS) Python 3.10.6 (project’s default) GPU CUDA 11.8 + (≥ 24 GB VRAM) or CPU‑only setup Tools git , curl , jq (optional but handy for testing) Network Ability to pull Docker images / PyPI packages and download the model from Hugging Face 1️⃣ Set up a virtual environment # ```shell script Create a directory for the whole demo (optional) # mkdir -p ~/gemma3-demo &amp;&amp; cd $_ Initialise the venv # python3 -m venv .venv source .venv/bin/activate Upgrade pip (always a good idea) # pip install --upgrade pip --- ## 2️⃣ Install **vLLM** and download the Gemma 3 model &gt; **See the full step‑by‑step instructions in** [`VLLM.md`](./VLLM.md). --- ## 3️⃣ Run the **vLLM** server Copy the helper script (or run the command manually) inside the demo directory: ```shell script # If you have the script `run-gemma-3-12b-it-vllm.sh` in the repo: cp path/to/llm-router/examples/quickstart/google-gemma3-12b-it/run-gemma-3-12b-it-vllm.sh . chmod +x run-gemma-3-12b-it-vllm.sh # Start the server (you may want to use tmux/screen) ./run-gemma-3-12b-it-vllm.sh The server will listen on http://0.0.0.0:7000 and expose an OpenAI‑compatible endpoint at /v1/chat/completions . You can quickly test it: ```shell script curl http://localhost:7000/v1/chat/completions \\ -H \"Content-Type: application/json\" \\ -d '{ \"model\": \"google/gemma-3-12b-it\", \"messages\": [{\"role\": \"user\", \"content\": \"Hello, how are you?\"}], \"max_tokens\": 100 }' | jq You should receive a JSON payload with the model’s generated text. --- ## 4️⃣ Install **LLM‑Router** ### Local install ```shell script # Clone the router repository (if you haven’t already) git clone https://github.com/radlab-dev-group/llm-router.git cd llm-router # Install the core library + API wrapper (includes the REST server) pip install .[api] # (Optional) Install Prometheus metrics support pip install .[api,metrics] Note: The router uses the same virtual environment you created earlier, so all dependencies stay isolated. 5️⃣ Prepare the router configuration # The example repository already ships a models-config.json that points to the locally running vLLM instance: { \"google_models\" : { \"google/gemma-3-12b-it\" : { \"providers\" : [ { \"id\" : \"gemma3_12b-vllm-local:7000\" , \"api_host\" : \"http://localhost:7000/\" , \"api_type\" : \"vllm\" , \"input_size\" : 56000 , \"weight\" : 1.0 } ] } }, \"active_models\" : { \"google_models\" : [ \"google/gemma-3-12b-it\" ] } } Copy it (or edit the path) to the router’s resources/configs/ directory: ```shell script mkdir -p resources/configs cp path/to/google-gemma3-12b-it/models-config.json resources/configs/ --- ## 6️⃣ Run the **LLM‑Router** ### Local Gunicorn The helper script `run-rest-api-gunicorn.sh` sets a sensible default environment. You can use it directly or export the variables yourself. ```shell script # Make the script executable (if needed) chmod +x path/to/run-rest-api-gunicorn.sh # Run the router ./run-rest-api-gunicorn.sh Key environment variables (already defined in the script) you may want to adjust: Variable Default Meaning LLM_ROUTER_SERVER_TYPE gunicorn Server backend (gunicorn, flask, waitress) LLM_ROUTER_SERVER_PORT 8080 Port on which the router listens LLM_ROUTER_MODELS_CONFIG resources/configs/models-config.json Path to the JSON file above LLM_ROUTER_PROMPTS_DIR resources/prompts Prompt‑template directory (optional) LLM_ROUTER_BALANCE_STRATEGY first_available Load‑balancing strategy LLM_ROUTER_USE_PROMETHEUS 1 (if you installed metrics) Enable /api/metrics endpoint After the script starts, the router will be re…"},{"k":"changelog.html","t":"Changelog","s":"Release notes","x":"## Changelog | Version | Changelog | |-----------|---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------…","h":["Changelog"],"b":"Changelog # Version Changelog 0.0.1 Initialization, License, setup, interface for each endpoint and sample ping EP. Autoloader of builtin endpoints and for the future implementations. 0.0.2 Add base models for api call (module llm_proxy_rest.data_models with error.py handling. Decorators to check required params and to measure the response time. 0.0.3 Proper AutoLoading for each found endpoint. Implementation of ApiTypesDispatcher , ApiModelConfig , ModelHandler . Ollama endpoints: / , tags . Added endpoint to full proxy with params. Streaming in case when external api provides stream. 0.0.4 All llama-service endpoints are refactored to llm-proxy-api . Refactoring base ep_run method. Proper handling system message, prompt name, model etc. 0.1.0 Repository name changed from llm-proxy-api to llm-router . Added class HttpRequestExecutor to handle http requests from EndpointWithHttpRequestI . Handled routing between any models: openai -&gt; ollama and ollama -&gt; openai 0.1.1 Prometheus metrics logging. Workers/Threads/Workers class is able to set by environments. Streaming fixes. Multi-providers for single model with default-balanced strategy. 0.2.0 Add balancing strategies: balanced , weighted , dynamic_weighted and first_available which works for streaming and non streaming requests. Included Prometheus metrics logging via /metrics endpoint. First stage of llm_router_lib library, to simply usage of llm-router-api . 0.2.1 Fix stream: OpenAI-&gt;Ollama, Ollama-&gt;OpenAI. Add Redis caching of availability of model providers (when using first_available strategy). Add llm_router_web module with simple flask-based frontend to manage llm-router config files. 0.2.2 Update dockerfile and requirements. Fix routing with vLLM. 0.2.3 New web configurator: Handling projects, configs for each user separately. First Available strategy is more powerful, a lot of improvements to efficiency. 0.2.4 Anonymizer module, integration anonymization with any endpoint (using dynamic payload analysis and full payload anonymisation), dedicated /api/anonymize_text endpoint as memory only anonymization. Whole router may be run in FORCE_ANONYMISATION mode. 0.3.0 Anonymization available with three strategies: fast_masker , genai , prov_masker . 0.3.1 Refactoring lb.strategies to be more flexible modular. Introduced MaskerPipeline and GuardrailPipeline both configured via env. Removed genai-based masking endpoint. 0.4.0 The main repository is divided into dedicated ones: plugins, services, web — separate repositories. Clean up the whole repository. Examples of integration with llamaindex, langchain, openai, litellm and haystack. 0.4.1 Audit log is stored using GPG. Add bash script ( scripts/gen_and_export_gpg.sh to prepare GPG keys and simple scripts/decrypt_auditor_logs.sh to decrypt encrypted audit logs. Moved core functionality from base to module core module. Quickstart. 0.4.2 Fix first_available_optim Strategy. Add KeepAliveMonitor to periodically pings model endpoints to keep them warm. 0.4.3 Add custom Prometheus metrices for logging masker/guardrail inidents. Fix OpenAI compatible v1 /models endpoint. Introduce monitors: services and keep alive models. Fixed guardrail retunr in case when streaming. 0.4.4 Validate unique provider identifiers. Store all hosts with keep‑alive configured in a Redis. UtilsPlugin pipeline with LangChain based simple RAG plugin (extending context to GenAI with locally built databse). Add handling of v1/response endpoint 0.4.5 Fixed sreaming to LMStudio native. Refactor streaming module. 0.4.6 Added support for embeddings endpoints across all providers. Extended ApiModel and ApiTypesI with is_embedding flag. Added test_embeddings.py utility for verifying embedding models through the API. 0.4.7 Integration with native Anthropic API. Add translate , generative_answer and ping methods to LLMRouterClient (with tests). Refactor LLMRouterCkientServices to use self.model_cls . Add payload converter for vLLM. 0.5.0 Integration with P…"},{"k":"plugins/index.html","t":"Plugins overview","s":"Plugins overview","x":"## Overview The **LLM‑Router** project ships with a modular plugin system that lets you plug‑in **anonymizers** (also called *maskers*) and **guardrails** into request‑processing pipelines. Each plugin implements a tiny, well‑defined inter…","h":["Overview","1. Anonymizers (Maskers)","1.1 What they do","1.2 Built‑in anonymizer plugins","1.3 How a masker is used","2. Guardrails","2.1 What they do","2.2 Built‑in guardrail plugins","2.3 How a guardrail is used","2.5 ML-Based PII Classification","How it complements regex maskers","Quick start","2.6 Polish Identification Regex Patterns","Available Polish rules","How they work","Adding rules to the pipeline","2.7 Semantic Routing (Model Selection)","2.7.1 Simple Semantic Routing (Heuristic)","2.7.2 Bi-Encoder Semantic Routing (Model Selection)","2.7.3 Codex CLI routing (auto_codex)","2.7.4 Claude Code Model Swap (claude-* → configured models)","3. Pipelines","3.1 Registration","3.2 Configuration","4. Adding a New Plugin","5. Retrieval‑Augmented Generation (RAG) Support","5.1 What the plugin does","5.2 Environment variables"],"b":"Overview # The LLM‑Router project ships with a modular plugin system that lets you plug‑in anonymizers (also called maskers ) and guardrails into request‑processing pipelines. Each plugin implements a tiny, well‑defined interface ( apply ) and can be composed in an ordered list to form a * pipeline *. Pipelines are instantiated by the MaskerPipeline and GuardrailPipeline classes and are driven automatically by the endpoint logic in endpoint_i.py . 1. Anonymizers (Maskers) # 1.1 What they do # Goal – Remove or replace personally‑identifiable information (PII) from a payload before it reaches the LLM or an external service. Typical strategy – Run a pipeline of maskers that locate spans corresponding to IDs, emails, IPs, etc., and replace each span with a placeholder such as {{MASKED_ITEM}} . 1.2 Built‑in anonymizer plugins # Plugin Description Technical notes FastMaskerPlugin ( fast_masker_plugin.py ) Thin wrapper around the FastMasker utility class. Receives a JSON‑compatible payload and returns the same payload with all detected PII masked. Implements PluginInterface . The heavy lifting is delegated to FastMasker.mask_payload(payload) . No extra I/O; the FastMasker instance is created once in __init__ . 1.3 How a masker is used # The endpoint (e.g. EndpointI._do_masking_if_needed ) checks the global flag FORCE_MASKING . If enabled, it creates a MaskerPipeline with the list of masker plugin identifiers (e.g. [\"fast_masker\"] ). The pipeline calls each plugin’s apply method sequentially, feeding the output of one as the input of the next. The final payload – now stripped of PII – proceeds to the rest of the request flow (guardrails, model dispatch, etc.). 2. Guardrails # 2.1 What they do # Goal – Verify that a request (or its response) complies with policy rules (e.g. no hateful, illegal, or unsafe content). Typical strategy – Split the payload into manageable text chunks, run a pipeline of guardrails, aggregate per‑chunk scores, and decide whether the overall request is safe. 2.2 Built‑in guardrail plugins # Plugin Description Technical notes NASKGuardPlugin ( nask_guard_plugin.py ) HTTP‑based guardrail that forwards the payload to the external NASK guardrail service ( /nask_guard endpoint) and returns a boolean safe flag together with the raw response. Inherits from HttpPluginInterface . The apply method calls _request(payload) (provided by the base class) and extracts results[\"safe\"] . Errors are caught and logged; on failure the plugin returns (False, {}) . SojkaGuardPlugin ( sojka_guard_plugin.py ) HTTP‑based guardrail that forwards the payload to the Sójka guardrail service ( /sojka_guard endpoint) and returns a safety flag. Mirrors the design of NASKGuardPlugin . The endpoint_url is built from the LLM_ROUTER_GUARDRAIL_SOJKA_GUARD_HOST environment variable. On success it returns (True, response) , otherwise (False, {}) . (Implicit) GuardrailProcessor ( processor.py ) Core logic used by the internal NASK guardrail Flask route ( nask_guardrail ). Tokenises the payload, creates overlapping chunks, runs a Hugging‑Face text‑classification pipeline, and produces a detailed safety report. Handles model loading ( AutoTokenizer , pipeline(\"text‑classification\") ), chunking ( _chunk_text ), and scoring thresholds ( MIN_SCORE_FOR_SAFE , MIN_SCORE_FOR_NOT_SAFE ). Returns a dict: {\"safe\": &lt;bool&gt;, \"detailed\": [...]} . 2.3 How a guardrail is used # The endpoint calls _is_request_guardrail_safe(payload) (or the analogous response guardrail). If FORCE_GUARDRAIL_REQUEST is true, a GuardrailPipeline is built from the configured plugin IDs (e.g. [\"nask_guard\", \"sojka_guard\"] ). The pipeline iterates over each guardrail plugin; each apply returns (is_safe, message) . The first plugin that reports is_safe=False short‑circuits the pipeline and the request is rejected with a 400/500 error payload. 2.5 ML-Based PII Classification # For cases where regex patterns alone are insufficient (e.g. context-dependent PII detection), the project integ…","r":"llm-router-plugins"},{"k":"plugins/llm-router-plugins/maskers/fast-masker/index.html","t":"Overview","s":"Masker plugins","x":"## Overview The **fast_masker** plugin provides a simple, rule‑based engine that scans a piece of text and replaces sensitive data ( e‑mail addresses, IPs, URLs, phone numbers, Polish PESEL identifiers, etc.) with clearly marked placeholde…","h":["Overview","Masking Rules","1. Highest Certainty — Checksum Validated Identifiers","2. High Certainty — Strict Format Validation","3. Medium-High Certainty — International Phone Numbers","4. Medium Certainty — Well-Structured Formats","5. Medium-Low Certainty — Business Identifiers","6. Lower Certainty — Pattern-Based with Context","7. Format-Based — Specific Patterns","8. Lowest Certainty — Generic Patterns","9. Beta Features","Disabled Rules (Too Noisy)","Utility Validators","Polish Identification Numbers","Financial Numbers","Vehicle and Transport","International Identification","Network and Security","Business Identifiers"],"b":"Overview # The fast_masker plugin provides a simple, rule‑based engine that scans a piece of text and replaces sensitive data ( e‑mail addresses, IPs, URLs, phone numbers, Polish PESEL identifiers, etc.) with clearly marked placeholders. The core component is the :class: ~llm_router_plugins.plugins.fast_masker.core.masker.FastMasker , which receives an ordered list of rule objects and applies each rule sequentially to the input text. Because the rules are applied in the order they are supplied, you can control precedence (e.g., replace URLs before e‑mails if needed). Masking Rules # Rules are applied in order from highest certainty (checksum-validated) to lowest certainty (pattern-based). This ordering minimizes false positives and ensures the most reliable identifiers are masked first. 1. Highest Certainty — Checksum Validated Identifiers # Rule Placeholder What it Detects Notes CreditCardRule {{CREDIT_CARD}} Credit card numbers (13-19 digits with optional spaces/dashes, e.g., 4532 1234 5678 9010 ). Validates using Luhn algorithm checksum. VinRule {{VIN}} Vehicle Identification Numbers (17 characters, e.g., 1HGBH41JXMN109186 ). Validates using ISO 3779 checksum (position 9). PeselTaggedRule {{PESEL_TAGGED}} Polish PESEL with label (e.g., PESEL: 44051401359 ). Validates checksum, preserves label prefix. PeselRule {{PESEL}} Polish PESEL numbers (11-digit personal identifiers). Validates checksum via is_valid_pesel . NipRule {{NIP}} Polish NIP numbers (plain, hyphen-separated, or wrapped in markdown). Validates checksum with weights [6,5,7,2,3,4,5,6,7] . KrsRule {{KRS}} Polish KRS numbers (10 digits, plain or hyphen-separated). Validates format only (10 digits). RegonRule {{REGON}} Polish REGON numbers (9 or 14 digits, optionally split by spaces). Validates checksum; handles both 9-digit and 14-digit forms. 2. High Certainty — Strict Format Validation # Rule Placeholder What it Detects Notes NrbRule {{NRB}} Polish NRB (bank account) numbers (26 digits, optional spaces). Supports formats: 26 digits or 2-4-4-4-4-4-4 grouping. MacAddressRule {{MAC_ADDRESS}} MAC addresses (6 octets, e.g., 00:1A:2B:3C:4D:5E ). Supports : , - , or no separators. PassportRule {{PASSPORT}} Passport numbers (2 letters + 7 digits, e.g., AB1234567 ). Case-insensitive matching. IdCardRule {{ID_CARD}} Polish ID card numbers (3 letters + 6 digits, e.g., ABC123456 ). Case-insensitive matching. SsnRule {{SSN}} US Social Security Numbers (format AAA-GG-SSSS , e.g., 123-45-6789 ). Format validation only, no checksum. 3. Medium-High Certainty — International Phone Numbers # Rule Placeholder What it Detects Notes PhoneInternationalRule {{PHONE_INTERNATIONAL}} Phone numbers with leading + and country code (e.g., +48 123 456 789 ). Matches 1-3 digit country codes with subscriber number groups. 4. Medium Certainty — Well-Structured Formats # Rule Placeholder What it Detects Notes EmailRule {{EMAIL}} E-mail addresses (e.g., user@example.com ). Permissive regex; matches local-part, @ , domain with TLD. Applied before URLs. UrlRule {{URL}} HTTP/HTTPS URLs and standalone domains (e.g., https://example.com , www.wp.pl ). Avoids code patterns like requests.post , response.json . IpRule {{IP}} , {{PORT}} IPv4, IPv6 addresses and hostname localhost . Masks ports as {{IP}}:{{PORT}} . Light octet validation; port captured separately. BankAccountRule {{BANK_ACCOUNT}} Polish IBAN (28 characters) and partially masked accounts (groups may contain X ). Exact length match; supports masked formats. 5. Medium-Low Certainty — Business Identifiers # Rule Placeholder What it Detects Notes JwtRule {{JWT}} JSON Web Tokens (3 Base64URL parts separated by dots). Validates structure; each part must be substantial (≥20 chars for header/payload). InvoiceNumberRule {{INVOICE_NUMBER}} Invoice identifiers (e.g., FV/2023/00123 , INV-2023-456 ). Case-insensitive; matches FV , INV , or INVOICE prefixes. OrderNumberRule {{ORDER_NUMBER}} E-commerce order identifiers (e.g., ORD123456 , ORDER-2023-001 ).…","r":"llm-router-plugins"},{"k":"plugins/llm-router-plugins/utils/routing/agentic-routing/claude-code/index.html","t":"Claude Code Model Swap (agenticroutingclaude_code)","s":"Routing plugins","x":"# Claude Code Model Swap (`agentic_routing_claude_code`) Model-swapping plugin for the **Anthropic-Messages-style** requests emitted by the **Claude Code** CLI. It looks at the model every request asks for, finds the Claude Code **tier** t…","h":["What the plugin changes","Quick start","1. Install the plugin package","2. Load the plugin into the router","3. Declare the models in the router","4. Point Claude Code at the router","5. Verify","How it works","Step 1 — normalize the requested name","Step 2 — match it against the configured tiers","Step 3 — rewrite and annotate","Configuration","Config file layout","Tiers and what each one replaces","Environment variables","Precedence and linting","Tuning and gotchas","Verifying the plugin","Offline dry run (no network, no router)","Router smoke test","Automated tests","Troubleshooting","Reference","Module map","Defaults at a glance","See also"],"b":"Claude Code Model Swap ( agentic_routing_claude_code ) # Model-swapping plugin for the Anthropic-Messages-style requests emitted by the Claude Code CLI. It looks at the model every request asks for, finds the Claude Code tier that owns that name, and rewrites the model to the one this router actually serves. Sonnet traffic, Opus traffic, the Haiku calls Claude Code runs in the background and the opusplan Plan Mode phase can each land on their own model, without any per-developer configuration. The plugin exists because Claude Code has no gateway-side notion of a model mapping. Pointing it at your own backend used to mean setting one environment variable per tier on every host — ANTHROPIC_DEFAULT_FABLE_MODEL , ANTHROPIC_DEFAULT_OPUS_MODEL , ANTHROPIC_DEFAULT_SONNET_MODEL , ANTHROPIC_DEFAULT_HAIKU_MODEL , CLAUDE_CODE_SUBAGENT_MODEL — and repeating it on every laptop, container image and CI runner. Here that mapping is one JSON file the router reads once. Plugin name (registry key): agentic_routing_claude_code Class: llm_router_plugins.utils.routing.agentic_routing.claude_code.plugin.ClaudeCodeRoutingPlugin Default config: llm_router_plugins/resources/routing/agentic_routing_claude_code.json Env prefix: LLM_ROUTER_ROUTING_SEMANTIC_AGENTIC_CLAUDE_CODE_ Dependencies: none beyond the standard library — no embedding model, no FAISS What the plugin changes # apply() rewrites the payload in place and touches the model key(s) plus one annotation: Key Value model model_name of the matched tier, when the payload carried it model_name the same target, when the payload carried it routing plugin , similarity , mode , original_model , matched_model , match_type , field Only keys the payload already carries are rewritten, and only those configured in settings.model_fields are considered at all: a payload that names its model model_name never grows a model key, and a key outside model_fields keeps its value. Everything else — messages , system , tools , max_tokens , metadata — is forwarded unchanged. The plugin never rejects a request; it only ever picks a model. Fail-open, and silent about it. A payload that is not a dict, carries no configured model key, whose model no tier claims, or whose tier has no model_name , is returned as the same object it came in as — no rewrite, no log line. There is deliberately nothing to grep for on that path: an unmatched model is normal traffic, not a signal. Anything raised inside the plugin is caught, logged once as a warning, and the request still goes out unchanged. Fail-hard on configuration only. No tiers, duplicated tier names, an embedded wildcard, or one exact model name claimed by two tiers all raise at plugin construction, so the router refuses to start rather than misroute quietly. Quick start # 1. Install the plugin package # # from a checkout of this repository, into the environment the router runs in pip install -e . # this plugin needs nothing extra Unlike the semantic and Codex routing plugins, this one has no [ml] extra: it never embeds anything. 2. Load the plugin into the router # The plugin is a utils plugin: it runs inside the llm-router request pipeline and is selected by a comma-separated list of plugin identifiers. export LLM_ROUTER_UTILS_PLUGINS_PIPELINE = \"agentic_routing_claude_code\" The router resolves the name against llm_router_plugins.utils.registry.MAIN_UTILS_REGISTRY , instantiates it once per process ( UtilsRegistry ) and calls apply() on every prepared request. Order matters when several utils plugins are wired together: the first one to rewrite the model wins the downstream dispatch, so put this plugin before any semantic routing plugin that might answer the same request. 3. Declare the models in the router # Every value in claude_code_modes[].model_name must be a normal, active model in the llm-router model config ( LLM_ROUTER_MODELS_CONFIG ) with a reachable provider. A name the router does not know is a hard downstream failure ( model not found ), not a plugin error. Ch…","r":"llm-router-plugins"},{"k":"plugins/llm-router-plugins/utils/routing/agentic-routing/codex/index.html","t":"Codex CLI Routing (agenticroutingcodex)","s":"Routing plugins","x":"# Codex CLI Routing (`agentic_routing_codex`) Routing plugin for the **OpenAI-Responses-style** requests emitted by the **Codex CLI** coding agent. It looks at every request that arrives with the trigger model (`auto_codex` by default), de…","h":["What the plugin changes","Quick start","1. Install the plugin package","2. Load the plugin into the router","3. Declare the models in the router","4. Point the Codex CLI at the router","5. Verify","How it works","Where it runs","Step 1 — payload normalization","Step 2 — request class","Step 3 — resolution cascade","Step 4 — keyword scoring","Step 5 — semantic similarity (optional)","Step 6 — annotating the payload","Failure behaviour","Configuration","Config file layout","Environment variables","Precedence and linting","Configuring the models","The shipped mapping","Model facts that matter here","Hard requirements on the router side","Choosing which mode gets which model","Tuning and gotchas","Verifying the plugin","Offline dry run (no network, no embedding model)","Router smoke test","Automated tests","Troubleshooting","Reference","Module map","Defaults at a glance","Reference deployment (example — your hosts will differ)","See also"],"b":"Codex CLI Routing ( agentic_routing_codex ) # Routing plugin for the OpenAI-Responses-style requests emitted by the Codex CLI coding agent. It looks at every request that arrives with the trigger model ( auto_codex by default), decides what kind of work the agent is doing right now , and rewrites payload[\"model\"] to the model configured for that work mode. Planning turns, code reviews, test runs, debugging, one-line thread titles and context compaction can then each run on the model that fits them best, while the Codex CLI keeps talking to a single stable model name. The decision cascade is deterministic first : structural signals (request class, the &lt;collaboration_mode&gt; block the CLI injects, keyword scoring) answer most requests offline, for free, and reproducibly. An optional embedding layer contributes cosine similarity over the mode descriptions and examples only when the cheap layers stay silent. Plugin name (registry key): agentic_routing_codex Class: llm_router_plugins.utils.routing.agentic_routing.codex.plugin.CodexRoutingPlugin Default config: llm_router_plugins/resources/routing/agentic_routing_codex.json Env prefix: LLM_ROUTER_ROUTING_SEMANTIC_AGENTIC_CODEX_ What the plugin changes # apply() rewrites the payload in place and touches exactly three keys: Key Value model model_name of the resolved Codex work mode agent_mode name of the resolved mode ( plan , implement , test , …) routing decision metadata: plugin , similarity , source , class and ids Everything else — input , instructions , tools , reasoning , text , stream , parallel_tool_calls — is forwarded unchanged. The plugin never rejects a request; it only ever picks a model. Fail-open by design. A payload that is not a dict, does not carry the trigger model, resolves to a mode that is not configured or to a mode with an empty model_name , or that raises anywhere during parsing or classification, is returned untouched with the trigger model still in place. Fail-hard on configuration only : an inconsistent config raises at plugin construction, so the router refuses to start instead of silently misrouting every request. Quick start # 1. Install the plugin package # # from a checkout of this repository, into the environment the router runs in pip install -e . # deterministic routing only pip install -e \".[ml]\" # + CPU FAISS and sentence-transformers (semantic layer) pip install -e \".[ml-gpu]\" # + CUDA FAISS instead — never both FAISS builds at once The [ml] / [ml-gpu] extras install faiss-cpu / faiss-gpu , sentence-transformers , numpy and scipy . Without them the plugin still loads and routes deterministically — the semantic layer steps aside with a single warning. In a container image, add the package (with the extra you want) next to llm-router so both import the same sentence-transformers . 2. Load the plugin into the router # The plugin is a utils plugin: it runs inside the llm-router request pipeline and is selected by a comma-separated list of plugin identifiers. export LLM_ROUTER_UTILS_PLUGINS_PIPELINE = \"agentic_routing_codex\" The router resolves the name against llm_router_plugins.utils.registry.MAIN_UTILS_REGISTRY , instantiates it once per process ( UtilsRegistry ), and calls its apply() on every prepared request, before guardrails and before a provider is chosen. Order matters if you wire several utils plugins together: the first plugin that rewrites model wins the downstream dispatch. Instantiating the plugin loads the JSON config, applies env overrides, validates them and — when the semantic layer is enabled — loads the embedding model and builds (or loads) the FAISS index. Expect a few seconds of extra startup time and a couple of hundred MB of RSS for a 300 M embedding model. 3. Declare the models in the router # Two things must exist in the llm-router model config ( LLM_ROUTER_MODELS_CONFIG , JSON): The trigger auto_codex as a declared model alias backed by a builtin provider — the CLI asks for it, the router accepts it, and the plugin re…","r":"llm-router-plugins"},{"k":"plugins/llm-router-plugins/utils/routing/semantic-biencoder/index.html","t":"Semantic BiEncoder Routing (semanticbiencoderrouting)","s":"Routing plugins","x":"# Semantic BiEncoder Routing (`semantic_biencoder_routing`) Embedding-based model selection. A set of **routing targets** — each one a name, a model and a pile of natural-language examples — is embedded once at startup into a FAISS index. …","h":["What the plugin changes","Quick start","1. Install with the ML extras","2. Load the plugin into the router","3. Make the embedding weights available","4. Declare the models in the router model config","5. Verify","How it works","Startup (constructor)","Request path","Measured behaviour","Configuration","Config file layout","Environment variables","Startup failures","Tuning and gotchas","Verifying the plugin","Live check (loads the embedding model)","Router smoke test","Automated tests","Troubleshooting","Reference","Module map","Defaults at a glance","See also"],"b":"Semantic BiEncoder Routing ( semantic_biencoder_routing ) # Embedding-based model selection. A set of routing targets — each one a name, a model and a pile of natural-language examples — is embedded once at startup into a FAISS index. Every incoming request that asks for model: \"auto\" has its last user message embedded and matched against that index by cosine similarity, and payload[\"model\"] is replaced with the model of the closest target. This is the \"teach routing by writing examples\" plugin: to change a decision, add or reword examples rather than maintain keyword lists. It needs sentence-transformers and faiss and it loads an embedding model into the router process. Plugin name (registry key): semantic_biencoder_routing Class: llm_router_plugins.utils.routing.semantic_biencoder.semantic_biencoder_routing.SemanticBiEncoderRoutingPlugin Default config: llm_router_plugins/resources/routing/semantic_biencoder.json Trigger: payload[\"model\"] equals auto after trimming ( should_route ) Env prefix: LLM_ROUTER_ROUTING_SEMANTIC_BIENCODER_ Dependencies: pip install -e \".[ml]\" (or [ml-gpu] ) What the plugin changes # Key Value model model_name of the winning routing target routing plugin , target_name , similarity (mean cosine of the match) Everything else in the payload is untouched. A rejected match (similarity below threshold) and a request without extractable text leave the payload as it is , so model stays \"auto\" in both cases. { \"model\" : \"qwen3.6:35b\" , \"routing\" : { \"plugin\" : \"semantic_biencoder_routing\" , \"similarity\" : 0.8224 , \"target_name\" : \"code-generation\" } } The log line that explains every decision: SemanticBiEncoderRouting: text='Napisz funkcję w Pythonie, która parsuje plik CSV' target='code-generation' similarity=0.8224 -&gt; model=qwen3.6:35b Quick start # 1. Install with the ML extras # # from a checkout of this repository, into the environment the router runs in pip install -e \".[ml]\" # CPU FAISS + sentence-transformers pip install -e \".[ml-gpu]\" # CUDA FAISS instead — never both FAISS builds at once Without these imports this plugin does not start : the constructor raises and the router refuses to boot. 2. Load the plugin into the router # export LLM_ROUTER_UTILS_PLUGINS_PIPELINE = \"semantic_biencoder_routing\" Resolved through llm_router_plugins.utils.registry.MAIN_UTILS_REGISTRY , instantiated once per process, then run on every prepared payload before guardrails and provider selection. 3. Make the embedding weights available # embedding_model is a Hugging Face id or a local directory . google/embeddinggemma-300m is gated behind the Gemma license, so either accept it and huggingface-cli login , or point at a local copy: { \"embedding_model\" : \"/models/google/embeddinggemma-300m\" } or without touching the file: export LLM_ROUTER_ROUTING_SEMANTIC_BIENCODER_MODEL = \"/models/google/embeddinggemma-300m\" A local path also removes the startup download, which matters for air-gapped or restart-heavy deployments. 4. Declare the models in the router model config # the auto alias, backed by a builtin provider (the client asks for it, this plugin replaces it); every model_name used by a routing target ( qwen3.6:35b , gpt-oss:120b , …) as a normal reachable model. // fragment of the llm-router models config \"semantic_routing\": { \"auto\": { \"providers\": [ { \"id\": \"semantic-routing__auto\", \"api_type\": \"builtin\", \"api_host\": null, \"api_token\": \"\", \"model_path\": \"\", \"input_size\": 0, \"weight\": 1.0, \"tool_calling\": false } ] } } 5. Verify # Embedding model loaded successfully. Index built: 6 targets, 79 total embeddings. SemanticBiEncoderRouting: text='…' target='code-generation' similarity=0.8224 -&gt; model=qwen3.6:35b Startup of the shipped config (6 targets, 79 vectors) takes about 9 s on CPU including the model load. See Verifying the plugin for a runnable check. How it works # Startup (constructor) # SemanticBiEncoderConfig.from_file() — bundled JSON, or the …_CONFIG env var (raw JSON string or a path). _override_from_env…","r":"llm-router-plugins"},{"k":"plugins/llm-router-plugins/utils/routing/index.html","t":"Semantic Routing Plugins","s":"Routing plugins","x":"# Semantic Routing Plugins Three routing plugins are available for **model selection** in the **LLM‑Router** system. - `simple_semantic_routing` and `semantic_biencoder_routing` activate when `payload[\"model\"] == \"auto\"`. - `agentic_routin…","h":["Table of Contents","0. Shared Routing Layer","1. Simple Semantic Routing (Heuristic)","1.1 Architecture","1.2 Intent Categories","1.3 Model Selection Algorithm","1.4 Configuration","1.5 Environment Variable Overrides","1.6 Usage Examples","1.7 Weight Tuning Guidelines","1.8 Phrase Format Guidelines","1.9 Complexity Levels","1.10 Built-in Patterns Reference","1.11 Adding New Intents","1.12 Running Tests","2. Bi-Encoder Semantic Routing (Embedding-based)","2.1 Index Building (on first load or when the persist directory is missing)","2.2 Routing (query)","2.3 Persistence","2.4 Configuration","2.5 Usage Examples","2.6 Scoring Details","2.7 Running Tests","3. Codex Routing (Codex CLI Requests)","3.1 Module Layout","3.2 Request Classes","3.3 Detection Cascade","3.4 Similarity: Cosine over Mode Embeddings","3.5 Configuration","3.6 Environment Variable Overrides","3.7 Usage Example","3.8 Running Tests","4. Comparison: Which Plugin to Use?","Recommendation","5. File Locations"],"b":"Semantic Routing Plugins # Three routing plugins are available for model selection in the LLM‑Router system. simple_semantic_routing and semantic_biencoder_routing activate when payload[\"model\"] == \"auto\" . agentic_routing_codex activates when payload[\"model\"] equals its own trigger value ( \"auto_codex\" by default) and routes the Codex CLI's request class (main turn, title generation, compaction) and work mode. Table of Contents # 0. Shared Routing Layer 1. Simple Semantic Routing (Heuristic) 2. Bi-Encoder Semantic Routing (Embedding-based) 3. Codex Routing (Codex CLI Requests) 4. Comparison: Which Plugin to Use? 5. File Locations 0. Shared Routing Layer # semantic_biencoder_routing and agentic_routing_codex share a common layer in llm_router_plugins/utils/routing/ : Module Contents embedder.py EmbeddingRouter (BiEncoder + FAISS: chunking, index build/load/persist, route() ) and EmbeddingRouterConfig — the formal config contract the router duck-types target.py RoutingTarget — the shared target/mode dataclass ( name , model_name , description , examples ); CodexMode is a subclass of it common.py RoutingConfigBase (shared from_file / from_json protocol: ..._CONFIG env var holding raw JSON or a file path, optional default config location), env_int / env_float / env_bool / resolve_persist_dir env helpers, build_embedding_router + check_router_has_vectors , should_route (the payload[\"model\"] trigger gate) and annotate_routing (writes payload[\"model\"] + payload[\"routing\"] ) Backward compatibility: semantic_biencoder/embedder.py and semantic_biencoder/config.py re-export EmbeddingRouter / EmbeddingRouterConfig / RoutingTarget from the shared modules, so existing imports keep working. Both plugins also share the text-extraction helper llm_router_plugins/utils/text_extractor.py::extract_user_text ( messages[-1].content → user_last_statement → query → prompt → input ). 1. Simple Semantic Routing (Heuristic) # The Simple Semantic Routing plugin ( simple_semantic_routing ) performs two-stage heuristic model selection: it classifies the user's intent (code, math, creative, general) via weighted keywords, multi-word phrases, and regex patterns, then estimates input complexity (token count) to pick the most appropriate model from a configured pool. No embedding model is required — routing is a fast, pure-text classification. 1.1 Architecture # Data Flow # User Input │ ▼ classify_intent(text) ← keywords + phrases + regex patterns │ ▼ estimate_tokens(text) ← word_count × 1.25 │ ▼ complexity_level(tokens) ← simple / medium / complex │ ▼ select_model(intent, complexity) ← weighted score → model index │ ▼ Updated payload[\"model\"] Intent Classification Algorithm # Each intent is defined in simple_semantic.json with four complementary signal types : a) Keywords — single words # Each keyword has an optional weight. If no weight is specified, the default is 1.0 . \"keywords\" : [ \"code\" , \"debug\" , \"funkcja\" ], \"weights\" : { \"debug\" : 5 , \"kod\" : 4 , \"code\" : 2 } Matching \"debug\" adds 5 to the score; matching \"code\" adds 2 . b) Phrases — multi-word expressions # Phrases use the format \"text:weight\" . The default weight is 2.0 when omitted. \"phrases\" : [ \"write code:5\" , \"fix bug:4\" , \"napisz funkcję:4\" ] Matching \"write code\" adds 5 to the score. c) Regex Patterns — structural detection # Patterns add a flat +3.0 per match, useful for detecting code structures, math formulas, and question patterns. \"patterns\" : [ \"def\\\\s+\\\\w+\\\\s*\\\\(\" , \"class\\\\s+\\\\w+\\\\s*[(:]\" , \"\\\\b\\\\d+\\\\s*[+\\\\-*/^]\\\\s*\\\\d+\\\\b\" , \"\\\\bwhat\\\\s+(is|are|does)\\\\b\" ] d) Weights — keyword importance # The weights object maps individual keywords to boost factors, allowing fine-grained control: \"weights\" : { \"debug\" : 5 , // Very strong signal for code intent \"błąd\" : 5 , // Very strong signal for code intent \"kod\" : 2 , // Moderate signal \"python\" : 2 // Moderate signal } 1.2 Intent Categories # Intent Description Examples code Programming, debugging, implementation \"napisz funkcję\", \"fix bug…","r":"llm-router-plugins"},{"k":"plugins/llm-router-plugins/utils/routing/simple-semantic/index.html","t":"Simple Semantic Routing (simplesemanticrouting)","s":"Routing plugins","x":"# Simple Semantic Routing (`simple_semantic_routing`) Heuristic, two-stage model selection: it classifies the **intent** of the last user message (code / math / creative / general) with weighted keywords, phrases and regexes, estimates the…","h":["What the plugin changes","Quick start","1. Install","2. Load the plugin into the router","3. Declare the models in the router model config","4. Verify","How it works","Where it runs","Text extraction","Stage 1 — intent classification","Stage 2 — complexity estimation","Stage 3 — model selection","Configuration","Config file layout","Environment variables","Validation","Tuning and gotchas","Verifying the plugin","Offline dry run","Router smoke test","Automated tests","Troubleshooting","Reference","Defaults at a glance","Files","See also"],"b":"Simple Semantic Routing ( simple_semantic_routing ) # Heuristic, two-stage model selection: it classifies the intent of the last user message (code / math / creative / general) with weighted keywords, phrases and regexes, estimates the complexity from a rough token count, and picks a model from an ordered pool. No embedding model, no vector store, no network — the whole decision is a few string scans, so it costs microseconds and is fully reproducible. It is the entry-level routing plugin: turn model: \"auto\" into \"a sensible model for this request\" without any ML dependencies. When you outgrow keyword heuristics, move to semantic_biencoder_routing , which decides by embedding similarity instead. Plugin name (registry key): simple_semantic_routing Class: llm_router_plugins.utils.routing.simple_semantic.simple_semantic_routing.SimpleSemanticRoutingPlugin Default config: llm_router_plugins/resources/routing/simple_semantic.json Trigger: payload[\"model\"] == \"auto\" (exact match) Env prefix: LLM_ROUTER_ROUTING_SEMANTIC_ Dependencies: none beyond the base package What the plugin changes # apply() sets exactly one key: Key Value model the pool entry selected for this intent + complexity Nothing else is added — there is no routing metadata block in the payload (unlike semantic_biencoder_routing and agentic_routing_codex ). Everything the plugin knows is written to the log: SimpleSemanticRouting: intent=code, complexity=simple (10 tokens) -&gt; qwen3.6:35b If no user text can be found, model is set to the configured default model and a warning is logged; a payload whose model is not exactly \"auto\" is returned untouched. Quick start # 1. Install # # from a checkout of this repository, into the environment the router runs in pip install -e . No extras needed — this plugin imports only the standard library. 2. Load the plugin into the router # export LLM_ROUTER_UTILS_PLUGINS_PIPELINE = \"simple_semantic_routing\" The router resolves the name in llm_router_plugins.utils.registry.MAIN_UTILS_REGISTRY , instantiates it once per process and runs apply() on every prepared payload, before guardrails and before a provider is selected. 3. Declare the models in the router model config # Everything the plugin can select must exist in the llm-router model config ( LLM_ROUTER_MODELS_CONFIG ): the auto alias itself, backed by a builtin provider — the client asks for it, the plugin replaces it; every model in the pool ( default_models values, or the LLM_ROUTER_ROUTING_SEMANTIC_MODELS list). // fragment of the llm-router models config \"semantic_routing\": { \"auto\": { \"providers\": [ { \"id\": \"semantic-routing__auto\", \"api_type\": \"builtin\", \"api_host\": null, \"api_token\": \"\", \"model_path\": \"\", \"input_size\": 0, \"weight\": 1.0, \"tool_calling\": false } ] } } A pool entry the router does not know is a hard downstream error, not a routing no-op. 4. Verify # Send a request with \"model\": \"auto\" and look for the log line, or run the verification snippet . Two lines to remember: SimpleSemanticRouting: intent=code, complexity=simple (10 tokens) -&gt; qwen3.6:35b No text content found, using default model gpt-oss:120b How it works # payload ──▶ extract text ──▶ stage 1: intent ──▶ stage 2: complexity ──▶ stage 3: pool index ──▶ payload[\"model\"] Where it runs # Inside the llm-router utils pipeline. The plugin is stateless: no cache, no session, no learned state, so every replica answers identically. Text extraction # Text is located by the shared helper llm_router_plugins.utils.text_extractor.extract_user_text , in this priority: payload[\"messages\"][-1][\"content\"] — the last message of a chat history (earlier turns are not seen) payload[\"user_last_statement\"] payload[\"query\"] payload[\"prompt\"] payload[\"input\"] Only that one string is scored, so intent detection always reflects the newest user turn. Stage 1 — intent classification # Every intent in intents accumulates a score from three signal kinds; the highest score above zero wins, ties go to the intent declared first in …","r":"llm-router-plugins"},{"k":"plugins/llm-router-plugins/utils/rag/langchain-rag.html","t":"LangChain RAG Plugin for LLM‑Router","s":"Utility plugins","x":"# LangChain RAG Plugin for LLM‑Router A lightweight **Retrieval‑Augmented Generation (RAG)** plugin that lets you attach any local knowledge base to the LLM‑Router without touching the original application code. It builds the knowledge bas…","h":["Table of Contents","Features","Installation","CLI – Indexing &amp; Searching","Searching","Example workflow","Error handling","License &amp; Credits","Happy augmenting! 🚀"],"b":"LangChain RAG Plugin for LLM‑Router # A lightweight Retrieval‑Augmented Generation (RAG) plugin that lets you attach any local knowledge base to the LLM‑Router without touching the original application code. It builds the knowledge base with LangChain , stores vectors in FAISS (fallback to Milvus is possible), and injects the retrieved context into every LLM request. Table of Contents # Features Installation Configuration (environment variables) CLI – Indexing &amp; Searching Integrating the plugin into LLM‑Router Example workflow Error handling License &amp; Credits Features # Feature Description Indexing Recursively reads a directory of .md , .txt , .html (or any extensions you pass) → splits documents into overlapping token windows → embeds chunks with a transformer model → stores them in a persistent FAISS index. Searching Given a user query, retrieves the top‑N most similar chunks (default 10) using cosine similarity. Context injection Retrieved chunks are appended (with a clear prefix) to the user’s last message, so the downstream LLM can use the extra knowledge automatically. Full configurability All hyper‑parameters (collection name, embedder model, device, chunk size, overlap, persistence directory) are set via environment variables. Fail‑safe If the RAG subsystem is disabled or mis‑configured, the plugin raises a clear exception while the rest of the router keeps working. CLI helpers llm-router-rag-langchain command provides convenient index and search sub‑commands, plus an interactive REPL mode. Installation # Only the optional RAG dependencies are required; they are not installed with the core LLM‑Router. ```shell script Activate your virtualenv / conda env first # pip install faiss-cpu tqdm langchain langchain-community transformers torch &gt; **Note** – If you prefer GPU‑accelerated FAISS, replace `faiss-cpu` with `faiss-gpu`. --- ## Configuration (environment variables) | Variable | Default | Description | |------------------------------------------|---------------|------------------------------------------------------------------------------------------------------------------------------------------------------------------------| | `LLM_ROUTER_LANGCHAIN_RAG_COLLECTION` | *must be set* | Name of the FAISS collection (e.g. `sample_collection`). | | `LLM_ROUTER_LANGCHAIN_RAG_EMBEDDER` | *must be set* | Hugging Face model identifier for the embedder (e.g. `sentence-transformers/all-MiniLM-L6-v2`). | | `LLM_ROUTER_LANGCHAIN_RAG_DEVICE` | `cpu` | Torch device (`cpu`, `cuda:0`, `cuda:1`, …). | | `LLM_ROUTER_LANGCHAIN_RAG_CHUNK_SIZE` | `1024` | Number of tokens per chunk. | | `LLM_ROUTER_LANGCHAIN_RAG_CHUNK_OVERLAP` | `100` | Overlap in tokens between consecutive chunks. | | `LLM_ROUTER_LANGCHAIN_RAG_PERSIST_DIR` | *none* | Directory where the FAISS index and docstore are persisted. If omitted, a folder named after the collection is created under `./workdir/plugins/utils/rag/langchain/`. | ### Example export block (add to your shell profile or a `.env` file) ```shell script export LLM_ROUTER_LANGCHAIN_RAG_COLLECTION=\"${LLM_ROUTER_LANGCHAIN_RAG_COLLECTION:-sample_collection}\" export LLM_ROUTER_LANGCHAIN_RAG_EMBEDDER=\"${LLM_ROUTER_LANGCHAIN_RAG_EMBEDDER:-sentence-transformers/all-MiniLM-L6-v2}\" export LLM_ROUTER_LANGCHAIN_RAG_DEVICE=\"${LLM_ROUTER_LANGCHAIN_RAG_DEVICE:-cpu}\" export LLM_ROUTER_LANGCHAIN_RAG_CHUNK_SIZE=\"${LLM_ROUTER_LANGCHAIN_RAG_CHUNK_SIZE:-200}\" export LLM_ROUTER_LANGCHAIN_RAG_CHUNK_OVERLAP=\"${LLM_ROUTER_LANGCHAIN_RAG_CHUNK_OVERLAP:-50}\" export LLM_ROUTER_LANGCHAIN_RAG_PERSIST_DIR=\"${LLM_ROUTER_LANGCHAIN_RAG_PERSIST_DIR:-./workdir/plugins/utils/rag/langchain/${LLM_ROUTER_LANGCHAIN_RAG_COLLECTION}}\" CLI – Indexing &amp; Searching # The plugin ships with a convenient entry‑point: ```shell script llm-router-rag-langchain [options] ### Indexing ```shell script # Index Markdown, plain‑text and HTML files from a repository llm-router-rag-langchain index --path \"../llm-router\" --ext .md .txt .html You can run …","r":"llm-router-plugins"},{"k":"services/index.html","t":"llm-router services","s":"Services overview","x":"# llm_router_services ## ✨ Overview `llm_router_services` delivers **HTTP services** that power the LLM‑Router plugin ecosystem. All functionality (guard‑rails, maskers, …) is exposed through **one Flask application** that can be started w…","h":["✨ Overview","🚀 Quick start","1. Install the package","3. Run the service","📡 API reference","Request payload","Example curl call","⚙️ Configuration (environment variables)","🧪 Development &amp; testing","Running tests","📜 License"],"b":"llm_router_services # ✨ Overview # llm_router_services delivers HTTP services that power the LLM‑Router plugin ecosystem. All functionality (guard‑rails, maskers, …) is exposed through one Flask application that can be started with a single command or via Gunicorn. Sub‑package Purpose guardrails/ Safety‑checking services (NASK‑PIB, Sojka) and a dynamic router ( router.py ) that registers only the endpoints whose environment flag is enabled. maskers/ PIIMasker – a token‑classification based PII anonymiser with an in‑memory cache that avoids redundant model calls for identical text inputs. run_servcices.sh Helper script that launches the unified API with Gunicorn, wiring all required environment variables. requirements.txt Heavy dependencies (e.g. transformers ) needed for GPU‑accelerated inference. All services load models once at start‑up and serve requests over HTTP. The masker caches predictions in memory for the lifetime of the process. 🚀 Quick start # 1. Install the package # ```shell script git clone https://github.com/radlab-dev-group/llm-router-services.git cd llm-router-services python -m venv .venv source .venv/bin/activate pip install -r requirements.txt editable install of the package itself # pip install -e . &gt; **Tip:** The package requires Python ≥ 3.10 (tested on 3.11+). ### 2. Set environment variables Only services whose `*_ENABLED` flag is set to `1` (or `true`) will be exposed. ```shell script export LLM_ROUTER_API_HOST=0.0.0.0 export LLM_ROUTER_API_PORT=5000 # Enable NASK‑PIB Guard export LLM_ROUTER_NASK_PIB_GUARD_ENABLED=1 export LLM_ROUTER_NASK_PIB_GUARD_MODEL_PATH=NASK-PIB/Herbert-PL-Guard # -1 = CPU, 0/1 = CUDA device index export LLM_ROUTER_NASK_PIB_GUARD_DEVICE=-1 # Enable Sojka Guard export LLM_ROUTER_SOJKA_GUARD_ENABLED=1 export LLM_ROUTER_SOJKA_GUARD_MODEL_PATH=speakleash/Bielik-Guard-0.1B-v1.0 # -1 = CPU, 0/1 = CUDA device index export LLM_ROUTER_SOJKA_GUARD_DEVICE=-1 # Enable PII Masker export LLM_ROUTER_PII_MASKER_ENABLED=1 # local path to a token-classification model (or an HF hub id) export LLM_ROUTER_PII_MASKER_MODEL_PATH=/models/pii-classification-bert # -1 = CPU, 0/1 = CUDA device index export LLM_ROUTER_PII_MASKER_DEVICE=-1 3. Run the service # Option A – via the helper script (recommended) # ```shell script ./run_servcices.sh The script starts **Gunicorn** with the Flask app created by `llm_router_services.router:create_app()`. #### Option B – directly with Python ```shell script python -m llm_router_services.router Both commands bind to 0.0.0.0:5000 (or the values you supplied). 📡 API reference # All endpoints are mounted under /api/guardrails/ (guard‑rails) or /api/maskers/ (maskers). Service Model Endpoint Method Description NASK‑PIB Guard NASK‑PIB/Herbert-PL-Guard /api/guardrails/nask_guard POST Polish safety classifier (hate, violence, etc.). Returns safe: bool and per‑chunk classification details. Sojka Guard speakleash/Bielik-Guard-0.1B-v1.0 /api/guardrails/sojka_guard POST Multi‑category Polish safety model (HATE, VULGAR, SEX, CRIME, SELF‑HARM). Returns per‑category scores and overall safe flag. PIIMasker (masker) PIIMasker (token classification) /api/maskers/pii POST Detects and redacts PII entities (PERSON, LOCATION, ORGANIZATION, etc.). Returns anonymized text plus mappings . Caches identical inputs in‑memory. Request payload # { \"payload\" : \"Your input string here (must be longer than 8 characters to be processed).\" } or more complicated: { \"payload\" : { \"Jak zrobić bombę\" : [ \"zaopatrzyć się w carnuchów, co będą robić czarny proch\" ], \"texts\" : [ \"To jest normalna wiadomość bez problemów.\" , \"Chcę się zabić i skończyć z tym wszystkim.\" , \"Dlaczego hipisi są tacy brudni i głupi\" , \"K*rwa, to jest ch*jowe rozwiązanie.\" , \"Opowiem ci o seksie i intymności.\" , \"Jak ukraść samochód i nie zostać złapanym.\" ] } } Example curl call # ```shell script curl -X POST http://localhost:5000/api/guardrails/nask_guard \\ -H \"Content-Type: application/json\" \\ -d '{\"payload\":\"Przykładowy teks…","r":"llm-router-services"},{"k":"services/llm-router-services/guardrails/index.html","t":"Guardrails Package","s":"Guardrail services","x":"# Guardrails Package `guardrails` provides the core safety‑checking services used by **llm‑router**, alongside the PII `maskers` package. All enabled services are mounted on a **single Flask application** (`llm_router_services/router.py`) …","h":["Sub‑packages","HTTP API","Running the services","Guardrail Services Documentation"],"b":"Guardrails Package # guardrails provides the core safety‑checking services used by llm‑router , alongside the PII maskers package. All enabled services are mounted on a single Flask application ( llm_router_services/router.py ) that exposes a simple HTTP API called by the corresponding llm‑router plugin. Sub‑packages # Paths are relative to llm_router_services/ . Path Purpose guardrails/inference/ Shared text‑classification engine, config interfaces and chunk scoring. guardrails/payload_handler.py Extracts textual elements from an arbitrary JSON payload. guardrails/nask/ NASK‑PIB safety model (Polish text classification). guardrails/speakleash/ Sojka (Bielik‑Guard) safety model (multi‑category). maskers/pii_classification/ PII masker – token‑classification anonymiser with a bounded in‑memory cache. router.py Builds the Flask app and registers every service whose *_ENABLED flag is on. ../../run_servcices.sh Starts all enabled services in a single Gunicorn process. HTTP API # All guardrail endpoints share the same request format: { \"payload\" : \"&lt;any JSON value&gt;\" } payload may be a string, object, list, or any JSON‑serialisable type. The service extracts all textual elements longer than 8 characters and runs them through the model. The response contains an overall safe flag and a detailed list with per‑chunk (or per‑category) scores. Non‑JSON request bodies are rejected with 400 ; inference failures return 500 with {\"error\": \"…\"} . Masker endpoints ( /api/maskers/&lt;name&gt; ) accept the same payload and answer with {\"anonymized\": …, \"mappings\": …} . Running the services # There is a single entry point – enable the services you need: shell script LLM_ROUTER_NASK_PIB_GUARD_ENABLED=1 \\ LLM_ROUTER_NASK_PIB_GUARD_MODEL_PATH=NASK-PIB/Herbert-PL-Guard \\ ./run_servcices.sh See the project README for the full list of environment variables. Guardrail Services Documentation # NASK‑PIB Guardrail – detailed usage, examples, and licensing information: guardrails/nask/README.md Sojka Guardrail – detailed usage, examples, and licensing information: guardrails/speakleash/README.md","r":"llm-router-services"},{"k":"services/llm-router-services/guardrails/speakleash/index.html","t":"Integration of Bielik‑Guard‑0.1B as llm‑router guardrail service (Sojka)","s":"Guardrail services","x":"# Integration of **Bielik‑Guard‑0.1B** as llm‑router guardrail service (Sojka) ## 1. Short Introduction The **Bielik‑Guard‑0.1B** model (`speakleash/Bielik-Guard-0.1B-v1.0`) is a Polish‑language safety classifier (text‑classification) buil…","h":["1. Short Introduction","2. Prerequisites","3. Running the Service","4. License and Usage Conditions","5. Sources &amp; Further Reading","6. Quick Start Code Snippet (Python)","🎉 Happy Guarding!"],"b":"Integration of Bielik‑Guard‑0.1B as llm‑router guardrail service (Sojka) # 1. Short Introduction # The Bielik‑Guard‑0.1B model ( speakleash/Bielik-Guard-0.1B-v1.0 ) is a Polish‑language safety classifier (text‑classification) built on top of the base model sdadas/mmlw-roberta-base . Within this project it is used to detect unsafe content in incoming requests handled by the /api/guardrails/sojka_guard endpoint defined in llm_router_services/guardrails/speakleash/sojka_guard_app.py . 2. Prerequisites # Component Version / Note Python 3.10 or newer (see python_requires in setup.py ) Packages transformers , torch , flask – already listed in requirements.txt Model speakleash/Bielik-Guard-0.1B-v1.0 (public on Hugging Face Hub) License Model – Apache‑2.0 . Code – Apache‑2.0 . No special commercial restrictions. Tip: The model will be downloaded automatically the first time you run the service. If you prefer to cache it locally, set the HF_HOME environment variable to a directory with enough space. 3. Running the Service # The module only exposes register_routes(app) ; the Flask app is built by llm_router_services/router.py , so the service is started through the common launcher: ```shell script LLM_ROUTER_SOJKA_GUARD_ENABLED=1 \\ LLM_ROUTER_SOJKA_GUARD_MODEL_PATH=speakleash/Bielik-Guard-0.1B-v1.0 \\ ./run_servcices.sh The endpoint is then available at: http://${LLM_ROUTER_API_HOST:-0.0.0.0}:${LLM_ROUTER_API_PORT:-5000}/api/guardrails/sojka_guard All enabled services share this single host and port – see the configuration table in the root `README.md`. ### Example request (using `curl`) ```shell script curl -X POST http://localhost:5000/api/guardrails/sojka_guard \\ -H \"Content-Type: application/json\" \\ -d '{\"payload\": \"Jak mogę zrobić bombę w domu?\"}' Example JSON response # { \"results\" : { \"detailed\" : [ { \"chunk_index\" : 0 , \"chunk_text\" : \"Jak mogę zrobić bombę w domu?\" , \"label\" : \"crime\" , \"safe\" : false , \"score\" : 0.9329 } ], \"safe\" : false } } Note: The label field contains one of the five safety categories defined by Bielik‑Guard ( HATE , VULGAR , SEX , CRIME , SELF‑HARM ). The score is the probability (0‑1) that the text belongs to the indicated category. The safe flag is false when any category exceeds the default threshold (0.5). 4. License and Usage Conditions # Element License Implications Application code ( guardrails/* ) Apache 2.0 Free for commercial and non‑commercial use, modification, and redistribution. Model ( Bielik‑Guard‑0.1B ) Apache 2.0 No non‑commercial restriction – the model can be used in commercial products provided attribution is kept. 5. Sources &amp; Further Reading # Model card : https://huggingface.co/speakleash/Bielik-Guard-0.1B-v1.0 Model card details (excerpt) markdown library_name: transformers license: apache-2.0 language: - pl base_model: - sdadas/mmlw-roberta-base pipeline_tag: text-classification Bielik‑Guard documentation (includes safety categories, training data, evaluation metrics, and citation information) – see the model card linked above. Community &amp; Support : Website: https://guard.bielik.ai/ Feedback / issue reporting: https://guard.bielik.ai/ 6. Quick Start Code Snippet (Python) # If you prefer to test the model locally before integrating it into the Flask service: from transformers import pipeline model_path = \"speakleash/Bielik-Guard-0.1B-v1.0\" classifier = pipeline ( \"text-classification\" , model = model_path , tokenizer = model_path , return_all_scores = True , ) texts = [ \"To jest normalna wiadomość bez problemów.\" , \"Chcę się zabić i skończyć z tym wszystkim.\" , \"Dlaczego hipisi są tacy brudni i głupi\" , \"K*rwa, to jest ch*jowe rozwiązanie.\" , \"Opowiem ci o seksie i intymności.\" , \"Jak ukraść samochód i nie zostać złapanym.\" ] for txt in texts : scores = classifier ( txt )[ 0 ] print ( f \" \\n Text: { txt } \" ) for s in scores : print ( f \" { s [ 'label' ] } : { s [ 'score' ] : .3f } \" ) Running the snippet will output probability scores for each of the five safety categori…","r":"llm-router-services"},{"k":"services/llm-router-services/guardrails/nask/index.html","t":"Integration of HerBERT‑PL‑Guard with the naskpibguard_app.py Service","s":"Guardrail services","x":"# Integration of **HerBERT‑PL‑Guard** with the `nask_pib_guard_app.py` Service ## 1. Short Introduction The **HerBERT‑PL‑Guard** model is a Polish‑language safety classifier (text‑classification) built on top of the base model `allegro/her…","h":["1. Short Introduction","2. Prerequisites","3. Running the Service","Example request (using curl)","4. License and Usage Conditions","Practical Consequences","5. Sources"],"b":"Integration of HerBERT‑PL‑Guard with the nask_pib_guard_app.py Service # 1. Short Introduction # The HerBERT‑PL‑Guard model is a Polish‑language safety classifier (text‑classification) built on top of the base model allegro/herbert-base-cased . Within this project it is used to detect unsafe content in incoming requests handled by the /api/guardrails/nask_guard endpoint defined in llm_router_services/guardrails/nask/nask_pib_guard_app.py . 2. Prerequisites # Component Version / Note Python 3.10 or newer (see python_requires in setup.py ) Packages transformers , torch , flask – already listed in requirements.txt Model NASK-PIB/HerBERT-PL-Guard (public on Hugging Face Hub) License Model – CC BY‑NC‑SA 4.0 (non‑commercial, attribution, share‑alike). Code – Apache 2.0 . Compatibility requires that any commercial deployment does not use the model in a way that violates the NC clause. 3. Running the Service # The module only exposes register_routes(app) ; the Flask app is built by llm_router_services/router.py , so the service is started through the common launcher: LLM_ROUTER_NASK_PIB_GUARD_ENABLED = 1 \\ LLM_ROUTER_NASK_PIB_GUARD_MODEL_PATH = NASK-PIB/HerBERT-PL-Guard \\ ./run_servcices.sh The endpoint is then available at http://${LLM_ROUTER_API_HOST:-0.0.0.0}:${LLM_ROUTER_API_PORT:-5000}/api/guardrails/nask_guard . All enabled services share this single host and port – see the configuration table in the root README.md . Example request (using curl ) # curl -X POST http://localhost:5000/api/guardrails/nask_guard \\ -H \"Content-Type: application/json\" \\ -d '{\"payload\": \"Jak mogę zrobić bombę w domu?\"}' Example JSON response # { \"results\" : { \"detailed\" : [ { \"chunk_index\" : 0 , \"chunk_text\" : \"Jak mogę zrobić bombę w domu?\" , \"label\" : \"S1\" , \"safe\" : false , \"score\" : 0.987 } ], \"safe\" : false } } 4. License and Usage Conditions # Element License Implications Application code ( guardrails/* ) Apache 2.0 Free for commercial and non‑commercial use, modification, and redistribution. Model ( HerBERT‑PL‑Guard ) CC BY‑NC‑SA 4.0 Attribution – original authors must be credited (Krasnodębska et al., 2025). Non‑commercial use – the model cannot be used in commercial products without additional permission. Share‑Alike – any derivative model must be released under the same CC BY‑NC‑SA 4.0 license. Datasets (PolyGuardMix, WildGuardMix) CC BY 4.0 / ODC‑BY 1.0 Require attribution but allow commercial use. Practical Consequences # The Apache 2.0 codebase may be released and used commercially, but the model must remain within the NC ( non‑commercial) constraints. Therefore, in commercial production environments you must: Run the model only for internal testing/evaluation, or Ensure that no model‑derived outputs are offered as a paid SaaS service without a separate license from the model owners. Repository distribution (e.g., a Git repo) must contain a LICENSE file that references both licenses and a README section describing the NC limitation. Fine‑tuning or modifying the model requires publishing the resulting model under CC BY‑NC‑SA 4.0 and retaining attribution to the original authors and datasets. 5. Sources # Model: HerBERT-PL-Guard Paper: @inproceedings { plguard2025 , author = {Aleksandra Krasnodębska and Karolina Seweryn and Szymon Łukasik and Wojciech Kusa} , title = {{PL-Guard: Benchmarking Language Model Safety for Polish}} , booktitle = {Proceedings of the 10th Workshop on Slavic Natural Language Processing} , year = {2025} , address = {Vienna, Austria} , publisher = {Association for Computational Linguistics} , url = {https://arxiv.org/abs/2506.16322} }","r":"llm-router-services"},{"k":"services/docker/index.html","t":"LLM‑Router Services – Docker Build & Run Guide","s":"Docker & deployment","x":"# LLM‑Router Services – Docker Build & Run Guide ## Overview This repository contains everything needed to containerise the **LLM‑Router** services with GPU support: | Component | File | Purpose | |-----------|---------------------|-------…","h":["Overview","Prerequisites","Common build‑time arguments","3. Build the No-CUDA Application Image","CPU build (llm-router-services:prod-cpu)","5. Customising Runtime Behaviour","6. Cleaning Up","CPU-only build &amp; run"],"b":"LLM‑Router Services – Docker Build &amp; Run Guide # Overview # This repository contains everything needed to containerise the LLM‑Router services with GPU support: Component File Purpose Base image Dockerfile.base Builds a CUDA‑enabled Ubuntu image with common utilities and PyTorch. Application image Dockerfile Extends the base image, pulls the service source code, installs Python dependencies, creates a non‑root user, and sets the entrypoint. No-CUDA application image Dockerfile.nocuda Standalone build image (Python + PyTorch, no NVIDIA/CUDA dependency). Useful for CPU-only deployment or when the CUDA base image is unavailable. Entrypoint script entrypoint.sh Handles optional debug mode and launches the main service script ( run_servcices.sh ). The steps below assume you have a recent Docker installation (Docker 20.10+ with the NVIDIA Container Toolkit for GPU access). Prerequisites # Requirement How to install Docker Engine Follow the official guide: https://docs.docker.com/engine/install/ NVIDIA drivers (host) Install the latest driver for your GPU (e.g., sudo apt install nvidia-driver-525 ). NVIDIA Container Toolkit https://docs.nvidia.com/datacenter/cloud-native/container-toolkit/install-guide.html Git (optional, for cloning the repo locally) sudo apt install git Verify GPU visibility inside Docker: ```shell script docker run --gpus all nvidia/cuda:12.2.2-runtime-ubuntu22.04 nvidia-smi You should see a table with your GPU details. --- ## 1. Build the Base Image The base image provides CUDA, Ubuntu, Python 3, `pip`, and a pre‑installed PyTorch wheel. It is built from **Dockerfile.base**, which accepts an optional `BASE_IMAGE` build‑arg (default: `nvidia/cuda:12.2.2-runtime-ubuntu22.04`). After the build the image is tagged as `gpu-base:cuda-12.2.2-ubuntu22.04` and can be used as a foundation for the application image. ```shell script # From the project ROOT (not inside docker/) docker build -f docker/Dockerfile.base \\ -t gpu-base:cuda-12.2.2-ubuntu22.04 . Tip – If you want to reuse the image across multiple projects, push it to a private registry: ```shell script docker tag gpu-base:cuda-12.2.2-ubuntu22.04 my-registry.example.com/gpu-base:cuda-12.2.2-ubuntu22.04 docker push my-registry.example.com/gpu-base:cuda-12.2.2-ubuntu22.04 --- ## 2. Build the Application Image The application image is based on the **gpu‑base** image you just built (or on any image you specify via the `BASE_IMAGE` build‑arg). It copies the service source code, installs the Python package, creates a dedicated user, and sets the entrypoint. &gt; ⚠️ **Important:** The Docker build context must be the **project root** (where `setup.py` lives), because `COPY .` copies everything from there — including the `llm_router_services/` package directory and `setup.py`. The Docker file itself is passed via `-f`: ```shell script # Run this from the project ROOT (not inside docker/) docker build \\ -f docker/Dockerfile \\ --build-arg version=prod \\ # optional: override the image label --build-arg USER_ID=5000 \\ # optional: custom UID for the runtime user --build-arg GROUP_ID=5000 \\ # optional: custom GID for the runtime group --build-arg BASE_IMAGE=gpu-base:cuda-12.2.2-ubuntu22.04 \\ # optional: custom base image -t llm-router-services:prod . Common build‑time arguments # Argument Default Description version prod Image label used for documentation / versioning. USER_ID / GROUP_ID 5000 UID/GID for the non‑root llm-router user inside the container. BASE_IMAGE gpu-base:cuda-12.2.2-ubuntu22.04 Base image for the application; can be any compatible CUDA image. 3. Build the No-CUDA Application Image # When you don't need GPU support (CPU-only deployment) or can't use the CUDA base image, use Dockerfile.nocuda . It is a fully standalone image based on python:3.14.6-trixie with Python 3, PyTorch, UTF-8 locale settings, and all required utilities baked in. ⚠️ Same as above — run from the project root with -f docker/Dockerfile.nocuda : ```shell script Run this from the project …","r":"llm-router-services"}]}