{"version":"0.4.3","pages":[{"k":"overview.html","t":"Overview","s":"Getting started","x":"# LLM Router - Open-Source AI Gateway for Local and Cloud LLM Infrastructure [**LLM Router**](https://llm-router.cloud) is a service that can be deployed on‑premises or in the cloud. It adds a layer between any application and the LLM prov…","h":["🧩 Boilerplates","✨ Key Features","📦 Quick Start","1️⃣ Create &amp; activate a virtual environment","2️⃣ Minimum required environment variable","🛠️ Configuration (via environment)","⚖️ Load Balancing Strategies","🛣️ Endpoints Overview","⚙️ Configuration Details","🔧 Development","📜 License","📚 Changelog"],"b":"LLM Router - Open-Source AI Gateway for Local and Cloud LLM Infrastructure # LLM Router is a service that can be deployed on‑premises or in the cloud. It adds a layer between any application and the LLM provider. In real time it controls traffic, distributes a load among providers of a specific LLM, and enables analysis of outgoing requests from a security perspective (masking, anonymization, prohibited content). It is an open‑source solution (Apache 2.0) that can be launched instantly by running a ready‑made image in your own infrastructure. llm_router_api provides a unified REST proxy that can route requests to any supported LLM backend ( OpenAI‑compatible, Ollama, vLLM, LM Studio, etc.), with built‑in load‑balancing, health checks, streaming responses and optional Prometheus metrics. llm_router_lib is a Python SDK that wraps the API with typed request/response models, automatic retries, token handling and a rich exception hierarchy, letting developers focus on application logic rather than raw HTTP calls. llm_router_web offers ready‑to‑use Flask UIs – an anonymizer UI that masks sensitive data and a configuration manager for model/user settings – demonstrating how to consume the router from a browser. llm_router_plugins (e.g., the fast_masker plugin) deliver a rule‑based text anonymisation engine with a comprehensive set of Polish‑specific masking rules (emails, IPs, URLs, phone numbers, PESEL, NIP, KRS, REGON, monetary amounts, dates, etc.) and an extensible architecture for custom rules and validators. llm_router_services provides HTTP services that implement the core functionality used by the LLM‑Router’s plugin system. The services expose guardrail and masking capabilities through Flask applications. All components run on Python 3.10+ using virtualenv and require only the listed dependencies, making the suite easy to install, extend, and deploy in both development and production environments. 🧩 Boilerplates # For a detailed explanation of each example’s purpose, structure, and how the boilerplates are organized, see the main project README: Main README – Boilerplate Overview – examples LlamaIndex Boilerplate Details – README ✨ Key Features # Feature Description Unified REST interface One endpoint schema works for OpenAI‑compatible, Ollama, vLLM and any future provider. Provider‑agnostic streaming The stream flag (default true ) controls whether the proxy forwards chunked responses as they arrive or returns a single aggregated payload. Built‑in prompt library Language‑aware system prompts stored under resources/prompts can be referenced automatically. Dynamic model configuration JSON file ( models-config.json ) defines providers, model name, default options and per‑model overrides. Request validation Pydantic models guarantee correct payloads; errors are returned with clear messages. Structured logging Configurable log level, filename, and optional JSON formatting. Health &amp; metadata endpoints /ping (simple 200 OK) and /tags (available model tags/metadata). Simple deployment One‑liner run script or python -m llm_proxy_rest.rest_api . Extensible conversation formats Basic chat, conversation with system prompt, and extended conversation with richer options (e.g., temperature, top‑k, custom system prompt). Multi‑provider model support Each model can be backed by multiple providers (VLLM, Ollama, OpenAI) defined in models-config.json . Provider selection abstraction ProviderChooser delegates to a configurable strategy, enabling easy swapping of load‑balancing, round‑robin, weighted‑random, etc. Load‑balanced default strategy LoadBalancedStrategy distributes requests evenly across providers using in‑memory usage counters. Dynamic model handling ModelHandler loads model definitions at runtime and resolves the appropriate provider per request. Pluggable endpoint architecture Automatic discovery and registration of all concrete EndpointI implementations via EndpointAutoLoader . Prometheus metrics integration Optional /metric…"},{"k":"llm-router-api/core/auditor/index.html","t":"Auditing subsystem – llm-router","s":"Security & auditing","x":"# Auditing subsystem – `llm-router` The **auditor** package provides a pluggable, tamper‑evident audit‑log system for the LLM‑router. All audit entries are written as JSON, encrypted with GPG and stored under `logs/auditor`. The subsystem …","h":["📁 Directory layout","🛠️ How the auditor works","🔐 GPG key management","1️⃣ Generate a new key pair","2️⃣ Place the public key where the router expects it","3️⃣ Decrypt audit logs","📚 Example: Auditing a request guard‑rail decision","🧩 Extending the auditor","📖 Further reading"],"b":"Auditing subsystem – llm-router # The auditor package provides a pluggable, tamper‑evident audit‑log system for the LLM‑router. All audit entries are written as JSON, encrypted with GPG and stored under logs/auditor . The subsystem is used by the router to record: request guard‑rail decisions payload masking operations custom audit events emitted by the application (e.g. business‑logic logs) The implementation is deliberately lightweight so it can be swapped out for a different storage backend (database, cloud bucket, …) without touching the rest of the code base. 📁 Directory layout # llm_router_api/ └─ core/ └─ auditor/ ├─ __init__.py # package marker ├─ auditor.py # public API – AnyRequestAuditor └─ log_storage/ ├─ __init__.py ├─ log_storage_interface.py # abstract storage contract └─ gpg.py # GPG‑backed storage implementation auditor.py – high‑level helper that forwards audit records to a storage backend. The default backend is GPGAuditorLogStorage . log_storage_interface.py – defines the AuditorLogStorageInterface protocol ( store_log(audit_log, audit_type) ). gpg.py – concrete implementation that encrypts each log entry with the public GPG key located at resources/keys/llm-router-auditor-pub.asc and writes the encrypted payload to a timestamped file logs/auditor/&lt;audit_type&gt;__&lt;timestamp&gt;.audit . 🛠️ How the auditor works # Endpoint code (e.g. endpoint_i.py ) creates an AnyRequestAuditor instance with the router’s logger. When an auditable event occurs, the endpoint builds a dictionary that contains at least the keys audit_type and payload . AnyRequestAuditor.add_log() forwards the dictionary to the configured storage backend. GPGAuditorLogStorage.store_log() JSON‑serialises the dictionary (pretty‑printed). Encrypts the JSON string with the imported public key. Writes the encrypted ASCII‑armored data to logs/auditor/ . The resulting files have the extension .audit . They are confidential and tamper‑evident – any modification breaks the GPG decryption. 🔐 GPG key management # The repository ships two helper scripts under scripts/ : Script Purpose gen_and_export_gpg.sh Generates a 4096‑bit RSA key pair (no interactive prompts) and exports the public ( *.asc ) and private ( *-priv.asc ) keys. decrypt_auditor_logs.sh Decrypts all *.audit files in logs/auditor/ and writes the resulting JSON to *.json . 1️⃣ Generate a new key pair # cd scripts ./gen_and_export_gpg.sh The script will: Prompt for an email address (used as the GPG user ID). Prompt for a passphrase (protects the private key). Create a key pair in the local GPG keyring. Export the public key to llm-router-auditor-pub.asc . Export the private key to llm-router-auditor-priv.asc . Important: Keep the private key ( *-priv.asc ) and its passphrase safe. Only the public key is required by the router at runtime. 2️⃣ Place the public key where the router expects it # mkdir -p resources/keys cp llm-router-auditor-pub.asc resources/keys/ The GPGAuditorLogStorage class automatically imports the key from this location when the application starts. 3️⃣ Decrypt audit logs # cd scripts ./decrypt_auditor_logs.sh For each file logs/auditor/&lt;type&gt;__&lt;timestamp&gt;.audit the script produces a human‑readable *.json file next to it: logs/auditor/request__20231129_123456.789012.audit → request__20231129_123456.789012.json You will be prompted for the passphrase of the private key if it is encrypted. 📚 Example: Auditing a request guard‑rail decision # from llm_router_api.core.auditor.auditor import AnyRequestAuditor import logging logger = logging . getLogger ( \"router\" ) auditor = AnyRequestAuditor ( logger ) # Somewhere inside an endpoint, after a guard‑rail check: audit_record = { \"audit_type\" : \"guardrail_request\" , \"payload\" : { \"user_id\" : \"12345\" , \"input\" : \"…\" , \"decision\" : \"blocked\" , \"reason\" : \"PII detected\" } } auditor . add_log ( audit_record ) The record is encrypted and persisted as e.g.: logs/auditor/guardrail_request__20231129_141530.123456.audit 🧩 Exte…"},{"k":"llm-router-api/index.html","t":"REST API reference","s":"REST API","x":"# llm‑router‑api **llm‑router‑api** is a lightweight Python library that provides a flexible, extensible proxy for Large Language Model ( LLM) back‑ends. It abstracts the details of multiple model providers (OpenAI‑compatible, Ollama, vLLM…","h":["Features","Installation","Running the Server","REST API Overview","Load‑Balancing Strategies","Keep‑Alive Mechanism","Extending the Router","Adding a New Provider Type","Adding a New Endpoint","Prompt Files","Monitoring &amp; Metrics","License"],"b":"llm‑router‑api # llm‑router‑api is a lightweight Python library that provides a flexible, extensible proxy for Large Language Model ( LLM) back‑ends. It abstracts the details of multiple model providers (OpenAI‑compatible, Ollama, vLLM, LM Studio, etc.) and offers a unified REST interface with built‑in load‑balancing, health‑checking, and monitoring. Repository: https://github.com/radlab-dev-group/llm-router Features # Unified API – One REST surface ( /api/... ) that proxies calls to any supported LLM back‑end. Provider Selection – Choose a provider per request using pluggable strategies (balanced, weighted, adaptive, first‑available). Prompt Management – System prompts are stored as files and can be dynamically injected with placeholder substitution. Streaming Support – Transparent streaming for both OpenAI‑compatible and Ollama endpoints. Health Checks – Built‑in ping endpoint and Redis‑based provider health monitoring. Prometheus Metrics – Optional instrumentation for request counts, latencies, and error rates. Auto‑Discovery – Endpoints are automatically discovered and instantiated at startup. Extensible – Add new providers, strategies, or custom endpoints with minimal boilerplate. Installation # The project uses Python 3.10.6 and a virtualenv ‑based workflow. ```shell script Clone the repository # git clone https://github.com/radlab-dev-group/llm-router.git cd llm-router Create a virtual environment # python3 -m venv venv source venv/bin/activate Install the package (including optional extras) # pip install -e .[metrics] # installs Prometheus support All required third‑party libraries are listed in `requirements.txt` (e.g., Flask, requests, redis, rdl‑ml‑utils, etc.). --- ## Configuration Configuration is driven primarily by environment variables and a JSON model‑config file. ### Environment Variables | Variable | Description | Default | |-------------------------------------------------------|--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|----------------------------------------| | `LLM_ROUTER_PROMPTS_DIR` | Directory containing predefined system prompts. | `resources/prompts` | | `LLM_ROUTER_MODELS_CONFIG` | Path to the models configuration JSON file. | `resources/configs/models-config.json` | | `LLM_ROUTER_DEFAULT_EP_LANGUAGE` | Default language for endpoint prompts. | `pl` | | `LLM_ROUTER_TIMEOUT` | Timeout (seconds) for llm-router API calls. | `0` | | `LLM_ROUTER_EXTERNAL_TIMEOUT` | Timeout (seconds) for external model API calls. | `300` | | `LLM_ROUTER_LOG_FILENAME` | Name of the log file. | `llm-router.log` | | `LLM_ROUTER_LOG_LEVEL` | Logging level (e.g., INFO, DEBUG). | `INFO` | | `LLM_ROUTER_EP_PREFIX` | Prefix for all API endpoints. | `/api` | | `LLM_ROUTER_MINIMUM` | Run service in proxy‑only mode (boolean). | `False` | | `LLM_ROUTER_IN_DEBUG` | Run server in debug mode (boolean). | `False` | | `LLM_ROUTER_BALANCE_STRATEGY` | Strategy used to balance routing between LLM providers. Allowed values are `balanced`, `weighted`, `dynamic_weighted` (beta), `first_available` and `first_available_optim` as defined in `constants_base.py`. | `balanced` | | `LLM_ROUTER_REDIS_HOST` | Redis host for load‑balancing when a multi‑provider model is available. | `&lt;empty string&gt;` | | `LLM_ROUTER_REDIS_PORT` | Redis port for load‑balancing when a multi‑provider model is available. | `6379` | | `LLM_ROUTER_REDIS_PASSWORD` | Password for Redis connection. | `&lt;not set&gt;` | | `LLM_ROUTER_REDIS_DB` | Redis database number. | `0` | | `LLM_ROUTER_SERVER_TYPE` | Server implementation to use (`flask`, `gunicorn`, `waitress`). | `flask` | | `LLM_ROUTER_SERVER_PORT` | Port on which the server listens. | `8080` | | `LLM_ROUTER_SERVER_HOST` | Host address for the server. | `0.0.0.0` | | `LLM_ROUTER_SERVER_WORKERS_CO…"},{"k":"llm-router-api/endpoints/index.html","t":"Endpoints Overview","s":"REST API","x":"## Endpoints Overview All endpoints are exposed under the REST API service. Unless stated otherwise, methods are POST and consume/produce JSON. ### Health & Info - **GET** `LLM_ROUTER_EP_PREFIX/ping` – Simple health‑check, returns `\"pong\"`…","h":["Endpoints Overview","Health &amp; Info","Provider‑Specific","Chat &amp; Completions (Built‑in)","Utility Endpoints (Built‑in)","Streaming vs. Non‑Streaming Responses","Payload format"],"b":"Endpoints Overview # All endpoints are exposed under the REST API service. Unless stated otherwise, methods are POST and consume/produce JSON. Health &amp; Info # GET LLM_ROUTER_EP_PREFIX/ping – Simple health‑check, returns \"pong\" . GET LLM_ROUTER_EP_PREFIX/ – Ollama health endpoint. Provider‑Specific # GET LLM_ROUTER_EP_PREFIX/tags – List available Ollama model tags. GET LLM_ROUTER_EP_PREFIX/models – List OpenAI‑compatible models. POST LLM_ROUTER_EP_PREFIX/api/v0/models – List LM Studio models. POST LLM_ROUTER_EP_PREFIX/api/chat – Ollama‑style chat completion. POST LLM_ROUTER_EP_PREFIX/api/chat/completions – OpenAI‑style chat completion. POST LLM_ROUTER_EP_PREFIX/chat/completions – OpenAI‑style chat completion (alternative path). POST LLM_ROUTER_EP_PREFIX/v1/chat/completions – vLLM‑like chat completion. Chat &amp; Completions (Built‑in) # POST LLM_ROUTER_EP_PREFIX/api/conversation_with_model – Standard chat endpoint (OpenAI‑compatible payload). POST LLM_ROUTER_EP_PREFIX/api/extended_conversation_with_model – Chat with extended fields support. POST LLM_ROUTER_EP_PREFIX/api/generative_answer – Answer a question using provided context. Utility Endpoints (Built‑in) # POST LLM_ROUTER_EP_PREFIX/api/generate_questions – Generate questions from input texts. POST LLM_ROUTER_EP_PREFIX/api/translate – Translate a list of texts. POST LLM_ROUTER_EP_PREFIX/api/simplify_text – Simplify input texts. POST LLM_ROUTER_EP_PREFIX/api/generate_article_from_text – Generate a short article from a single text. POST LLM_ROUTER_EP_PREFIX/api/create_full_article_from_texts – Generate a full article from multiple texts. Streaming vs. Non‑Streaming Responses # Streaming ( stream: true – default) The proxy opens an HTTP chunked connection and forwards each token/segment from the upstream LLM as soon as it arrives. Clients can process partial output in real time (e.g., live UI updates). Non‑Streaming ( stream: false ) The proxy collects the full response from the provider, then returns a single JSON object containing the complete text. Use this mode when you need the whole answer before proceeding. Both modes are supported for every provider that implements the streaming interface (OpenAI, Ollama, vLLM). The stream flag lives in the request schema ( OpenAIChatModel and analogous models) and is honoured automatically by the proxy. Payload format # Payload format follows the OpenAI schema ( model , messages , optional stream , etc.) unless a custom endpoint overrides it. All endpoints automatically: Validate required arguments (via REQUIRED_ARGS ). Resolve the appropriate provider using the configured load‑balancing strategy . Inject system prompts when SYSTEM_PROMPT_NAME is defined. Return a JSON response with { \"status\": true, \"body\": … } or an error payload."},{"k":"llm-router-api/endpoints/readme-pl.html","t":"Jak tworzyć endpointy w llm-proxy-api – przewodnik","s":"REST API","x":"## Jak tworzyć endpointy w llm-proxy-api – przewodnik Poniżej zebrano kluczowe informacje o tym, jak definiować i konfigurować endpointy (EP) na podstawie klas z `endpoints.*`, z odniesieniem do logiki wykonywania w `endpoint_i.EndpointWit…","h":["Jak tworzyć endpointy w llm-proxy-api – przewodnik","Propozycja EP: BatchFileSummaries – podsumowania plików z listy"],"b":"Jak tworzyć endpointy w llm-proxy-api – przewodnik # Poniżej zebrano kluczowe informacje o tym, jak definiować i konfigurować endpointy (EP) na podstawie klas z endpoints.* , z odniesieniem do logiki wykonywania w endpoint_i.EndpointWithHttpRequestI.run_ep(...) . Uwzględniono też role atrybutów/stałych takich jak self._map_prompt , self._prompt_str_postfix , _prepare_response_function , _prompt_str_force , SYSTEM_PROMPT_NAME , REQUIRED_ARGS , OPTIONAL_ARGS oraz parametry konstruktora. 1) Hierarchia i warianty bazowe - EndpointI: baza dla EP (gdy serwis nie działa jako proxy). Definiuje ogólne API i walidację argumentów, ale nie implementuje run_ep. - EndpointWithHttpRequestI: rozszerza EndpointI o wysyłkę żądań HTTP do zewnętrznego LLM. Ma pełną implementację run_ep, obsługę streamingu i wstrzykiwania promptu systemowego. - PassthroughI: dziedziczy z EndpointWithHttpRequestI i domyślnie “przepuszcza” payload (prepare_payload zwraca parametry bez zmian). Użyteczne dla OpenAI‑kompatybilnych EP, gdzie chcemy prosto forwardować żądania. Dlaczego w openai.py dziedziczymy z PassthroughI? - Bo endpointy OpenAI‑kompatybilne często wymagają minimalnej logiki – wystarczy przekazać dalej to, co przyszło. PassthroughI upraszcza implementację (brak wymuszonych argumentów, brak system promptu, gotowy run_ep proxy). 2) Cykl wykonania – co robi run_ep w EndpointWithHttpRequestI W dużym skrócie: - Inicjalizacja zegara i wyzerowanie atrybutów promptu: _map_prompt , _prompt_str_force , _prompt_str_postfix . - Wywołanie prepare_payload(params): tu podklasa ma przekształcić wejście do formatu, jaki rozumie backend (np. ułożyć messages, przepisać model_name → model, ustawić stream itp.). Jeśli zwróci strukturę z \"status\": False , run_ep zwróci ją bez dalszego przetwarzania. - Jeśli ustawiono direct_return=True, zwracany jest wynik prepare_payload bez proxy. - Tryb “simple proxy”: jeżeli klasa nie definiuje REQUIRED_ARGS (pusta lista) – traktujemy EP jako bezpośredni proxy do odpowiednika po stronie modelu. Wtedy: - _set_model wybiera model na podstawie pól z MODEL_NAME_PARAMS. - Jeżeli typ API modelu jest zgodny z typami EP ( api_types ), payload jest przekazywany dalej do odpowiedniego URL (z opcjonalnym stream). - Jeżeli to nie simple proxy: - _resolve_prompt_name(...) przygotowuje system prompt (opisane w pkt 3). - __dispatch_external_api_model(params) ustawia _api_model na podstawie nazwy modelu. - Wyznaczamy docelowy URL przez ApiTypesDispatcher (np. chat_ep dla danego api_type ). - Obsługa stream=False/True (w tym wariancie streaming może być ograniczony – komunikat o braku wsparcia). - _call_http_request(...) wykonuje POST/GET do hosta modelu, składając finalny payload (w tym system message, jeśli jest). Dodatkowe ścieżki: - call_for_each_user_msg=True : dla zadań wielotekstowych – wysyłamy osobne żądanie dla każdej wiadomości użytkownika, a wynik agregujemy przez _prepare_response_function . 3) System prompt i modyfikacje treści – jak działają pola - SYSTEM_PROMPT_NAME: słownik { \"pl\": prompt_id, \"en\": prompt_id }. W prepare_payload ustawiasz wymagania EP, a run_ep w _resolve_prompt_name : - wybiera język z parametru LANGUAGE_PARAM (z defaultem DEFAULT_EP_LANGUAGE), - pobiera treść promptu systemowego przez PromptHandler jeśli zdefiniowano nazwę, - stosuje _map_prompt – słownik zamian {placeholder: tekst}, np. wstrzyknięcie liczby pytań, treści zapytania użytkownika, - dokleja _prompt_str_postfix na końcu system promptu (np. dodatkowa instrukcja), - jeśli _prompt_str_force jest ustawione – nadpisuje całą treść system promptu (pomija nazwę/system prompt z plików). Efekt: jeśli _prompt_str ostatecznie jest zbudowany, to zostaje dodany do messages jako pierwszy element: {\"role\": \"system\", \"content\": self._prompt_str}. Kiedy to ustawiać? - W prepare_payload: - self._map_prompt: gdy chcesz w promptach z zasobów podmienić znaczniki (np. ##QUESTION_NUM_STR##). - self._prompt_str_postfix: gdy EP potrzebuje dokleić końcową uwagę/regułę do system pr…"},{"k":"llm-router-api/keepalive.html","t":"Keep‑Alive Utility Overview","s":"REST API","x":"# Keep‑Alive Utility Overview The **keep‑alive** subsystem is responsible for periodically “pinging” model endpoints so that they stay warm and ready to serve requests with minimal latency. It consists of two main components: | Component |…","h":["How It Works","Integration Points","Configuring Keep‑Alive","Example Usage","Logging"],"b":"Keep‑Alive Utility Overview # The keep‑alive subsystem is responsible for periodically “pinging” model endpoints so that they stay warm and ready to serve requests with minimal latency. It consists of two main components: Component Purpose Key Methods KeepAlive Sends a single HTTP request to a model provider. It resolves the correct provider configuration (API type, host, token, model name) and builds the request payload. send(model_name, host, prompt=None) – performs the HTTP call. KeepAliveMonitor Schedules repeated keep‑alive calls for each (model_name, host) pair. It stores scheduling data in Redis, checks host availability, and triggers KeepAlive.send when a provider is due. record_usage(model_name, host, keep_alive) – registers a provider for periodic pinging. start() / stop() – control the background thread. How It Works # Provider discovery – KeepAlive._find_provider looks up the provider configuration for a given model name and host inside the global models_configs dictionary. Endpoint resolution – Depending on the provider’s api_type ( vllm , openai , ollama ), _endpoint_for builds the correct URL ( /v1/chat/completions or /api/chat ). HTTP request – A JSON payload containing a short “keep‑alive” prompt (default: “Send an empty message.” ) is posted to the endpoint. Scheduling – KeepAliveMonitor stores metadata in Redis: A hash key ( keepalive:provider:&lt;model&gt;:&lt;host&gt; ) with keep_alive_seconds . A sorted‑set ( keepalive:providers:next_wakeup ) that orders providers by the next scheduled wake‑up timestamp. Background loop – The monitor thread wakes up every check_interval seconds, fetches due providers, verifies that the host is free (via the optional is_host_free_callback ), and invokes KeepAlive.send . After a successful ping, the next wake‑up time is recomputed. Integration Points # Strategy implementations (e.g., FirstAvailableOptimStrategy ) create a KeepAlive instance and pass it to a KeepAliveMonitor . When a provider is selected, the strategy calls keep_alive_monitor.record_usage(model_name, host, keep_alive) so the monitor knows to ping that endpoint. The monitor runs automatically in the background once start() is called (typically during strategy initialization). Configuring Keep‑Alive # Provider configurations live in the global models_configs JSON (see resources/configs/models-config*.json ). To enable keep‑alive for a specific provider, add a keep_alive field with a duration string: { \"model_name\" : \"gpt‑4\" , \"providers\" : [ { \"api_type\" : \"openai\" , \"api_host\" : \"http://localhost:8000\" , \"api_token\" : \"YOUR_TOKEN\" , \"keep_alive\" : \"2m\" // ping every 2 minutes } ] } Supported duration units: Unit Suffix Meaning seconds s e.g., \"30s\" minutes m e.g., \"5m\" hours h e.g., \"1h\" If the keep_alive field is omitted or falsy, the provider will not be scheduled for periodic pings. Example Usage # from llm_router_api.core.monitor.keep_alive import KeepAlive from llm_router_api.core.monitor.keep_alive_monitor import KeepAliveMonitor # Assume `models_configs` has been loaded from the JSON config files. keep_alive = KeepAlive ( models_configs = models_configs ) monitor = KeepAliveMonitor ( redis_client = redis_client , keep_alive = keep_alive , check_interval = 10.0 , # check every 10 seconds is_host_free_callback = my_is_host_free , clear_buffers = True , # clean old keys on start ) monitor . start () # When a provider is selected somewhere in the routing logic: monitor . record_usage ( model_name = \"gpt‑4\" , host = \"http://localhost:8000\" , keep_alive = \"2m\" ) The monitor will now ping the gpt‑4 endpoint every two minutes, provided the host is not busy with another model. Logging # Both KeepAlive and KeepAliveMonitor emit detailed logs at the DEBUG and INFO levels, prefixed with [keep-alive] . Adjust your logger configuration to capture these messages for troubleshooting. Keep‑alive helps maintain low‑latency responses by preventing model containers from idling out. Proper configuration and integration wi…"},{"k":"llm-router-api/lb-strategies.html","t":"Load Balancing Strategies","s":"REST API","x":"## Load Balancing Strategies The `llm-router` supports various strategies for selecting the most suitable provider when multiple options exist for a given model. This ensures efficient and reliable routing of requests. The available strate…","h":["Load Balancing Strategies","1. balanced (Default)","2. weighted","3. dynamic_weighted (beta)","4. first_available","5. first_available_optim","Environments and Redis installation","Extending with Custom Strategies"],"b":"Load Balancing Strategies # The llm-router supports various strategies for selecting the most suitable provider when multiple options exist for a given model. This ensures efficient and reliable routing of requests. The available strategies are: 1. balanced (Default) # Description: This is the default strategy. It aims to distribute requests evenly across available providers by keeping track of how many times each provider has been used for a specific model. It selects the provider that has been used the least. When to use: Ideal for scenarios where all providers are considered equal in terms of capacity and performance. It provides a simple and effective way to balance the load. Implementation: Implemented in llm_router_api.core.lb.balanced.LoadBalancedStrategy . 2. weighted # Description: This strategy allows you to assign static weights to providers. Providers with higher weights are more likely to be selected. The selection is deterministic, ensuring that over time, the request distribution closely matches the configured weights. When to use: Useful when you have providers with different capacities or performance characteristics, and you want to prioritize certain providers without needing dynamic adjustments. Implementation: Implemented in llm_router_api.core.lb.weighted.WeightedStrategy . 3. dynamic_weighted (beta) # Description: An extension of the weighted strategy. It not only uses weights but also tracks the latency between successive selections of the same provider. This allows for more adaptive routing, as providers with consistently high latency might be de-prioritized over time. You can also dynamically update provider weights. When to use: Recommended for dynamic environments where provider performance can fluctuate. It offers more sophisticated load balancing by considering both configured weights and real-time performance metrics (latency). Implementation: Implemented in llm_router_api.core.lb.weighted.DynamicWeightedStrategy . 4. first_available # Description: This strategy selects the very first provider that is available. It uses Redis to coordinate across multiple workers, ensuring that only one worker can use a specific provider at a time. When to use: Suitable for critical applications where you need the fastest possible response and want to ensure that a request is immediately handled by any available provider, without complex load distribution logic. It guarantees that a provider, once taken, is exclusive until released. Implementation: Implemented in llm_router_api.core.lb.first_available.FirstAvailableStrategy . When using the first_available load balancing strategy, a Redis server is required for coordinating provider availability across multiple workers. 5. first_available_optim # What it is first_available_optim is an enhanced version of the plain first‑available load‑balancing strategy. It uses Redis to coordinate across multiple workers and tries to reuse a host that has already been used for the requested model before falling back to the classic “pick the first free provider” logic. How it works Step Purpose Behaviour 1️⃣ Re‑use the last host If the model was previously run on a specific host and that host is currently free, select it. The host identifier is stored in a Redis key :last_host . The strategy checks that the host is not occupied by another model and attempts an atomic acquisition of a provider on that host. 2️⃣ Re‑use any known host Prefer any host that already has the model loaded. A Redis set :hosts tracks all hosts where the model is currently loaded. The strategy scans the provider list, picks a free provider on one of those hosts, and locks it atomically. 3️⃣ Pick an unused host Spread the load to a fresh host when no suitable “known” host is available. It looks for a provider whose host is not present in the :hosts set and is not occupied, then acquires it. 4️⃣ Fallback to plain first‑available Guarantees a result even if the optimisation steps fail. If none of the previous …"},{"k":"llm-router-api/models-config.html","t":"Models configuration description (JSON)","s":"REST API","x":"# Models configuration description (JSON) ## 📄 Purpose This document explains the **model configuration** used by the LLM Router. It describes the JSON schema that drives **`ModelHandler`** and **`ApiModelConfig`**, clarifies each field, a…","h":["📄 Purpose","🏗️ High‑level structure","🧩 How ModelHandler uses the config","📦 Sample configuration (models-config.json)","🎉 Summary"],"b":"Models configuration description (JSON) # 📄 Purpose # This document explains the model configuration used by the LLM Router. It describes the JSON schema that drives ModelHandler and ApiModelConfig , clarifies each field, and provides a ready‑to‑use example ( models-config.json ). Having a single source of truth for model definitions makes it easy to: Add or remove providers for a given model. Switch between cloud (OpenAI, Google) and local (vLLM, Ollama) back‑ends. Control load‑balancing, keep‑alive, and tool‑calling options per provider. Activate only the models you want to expose through the router. 🏗️ High‑level structure # ```plain text { \" \": { # e.g. \"google_models\", \"openai_models\", \"qwen_models\" \" \": { # full identifier used by the router, e.g. \"google/gemma-3-12b-it\" \"providers\": [ … ], # primary providers (used for normal traffic) \"providers_sleep\": [ … ] # optional low‑priority providers (used when others are busy) }, … }, \"active_models\": { # required – tells the router which models are enabled \" \": [ \" \", … ], … } } * **Model type** – a top‑level key grouping models that share the same provider‑type logic. * **Model name** – the identifier that appears in API calls (`model` field). * **`providers`** – a list of dictionaries, each describing a concrete endpoint. * **`providers_sleep`** (optional) – “sleeping” providers that are only used when all primary providers are unavailable or overloaded. * **`active_models`** – the only place where a model is marked as *active*. If a model is missing here, the router will ignore it even if it is present in the rest of the file. --- ## 🔎 Detailed field description ### Provider dictionary (items in `providers` / `providers_sleep`) | Field | Type | Description | Example | |----------------|---------------------------|-------------------------------------------------------------------------------------------------------------------------------------|---------------------------------| | `id` | `str` | Unique identifier for the provider instance (used for logging &amp; selection). | `\"gemma3_12b-vllm-71:7000\"` | | `api_host` | `str` | Base URL of the provider API (must include protocol, may contain trailing slash). | `\"http://192.168.100.71:7000/\"` | | `api_token` | `str` | Authentication token; empty string if not required. | `\"\"` | | `api_type` | `str` | Type of the backend – determines which concrete `BaseProvider` class is used (`openai`, `vllm`, `ollama`, …). | `\"vllm\"` | | `input_size` | `int` (or numeric string) | Maximum context length the provider accepts. The `ApiModel.from_config` helper converts it to `int`. | `4096` | | `model_path` | `str` | Path or name of the model on the provider side (used by Ollama, vLLM, etc.). May be empty for providers that infer it from the URL. | `\"gpt-3.5-turbo-0125\"` | | `weight` | `float` | Relative weight for **weighted‑random** load‑balancing strategies. Default `1.0`. | `0.1` | | `keep_alive` | `str` | Optional keep‑alive duration (e.g. `\"35m\"`). Empty or `null` means the provider is not kept alive. | `\"35m\"` | | `tool_calling` | `bool` | Whether the provider supports tool‑calling (function calling). | `true` | ### `active_models` section ```json { (...) \"active_models\": { \"google_models\": [ \"google/gemma-3-12b-it\", \"google/gemini-2.5-flash-lite\" ], \"openai_models\": [ \"openai/gpt-3.5-turbo-0125\", \"gpt-oss:20b\", \"gpt-oss:120b\" ], \"qwen_models\": [ \"qwen3-coder:30b\" ] } } The key must match a top‑level model type defined elsewhere in the file. The list contains the exact model names that appear under that type. Only the models listed here are loaded by ApiModelConfig._read_active_models() and later exposed by ModelHandler . 🧩 How ModelHandler uses the config # Construction handler = ModelHandler ( models_config_path = \"/path/to/models-config.json\" , provider_chooser = my_provider_strategy ) ApiModelConfig reads the file, extracts active_models , and builds models_configs – a dict that maps each active model name to its full configurati…"},{"k":"llm-router-lib/index.html","t":"llmrouterlib","s":"Python library","x":"# llm_router_lib ## Overview `llm_router_lib` is ** a collection of data‑model definitions**. It supplies the **foundation** for request/response structures used by the `llm_router_api` package **and** provides a **thin, opinionated client…","h":["Overview","Installation","Quick start","Data models","Conversation models","Thin client wrapper (LLMRouterClient)"],"b":"llm_router_lib # Overview # llm_router_lib is ** a collection of data‑model definitions . It supplies the foundation for request/response structures used by the llm_router_api package and provides a thin, opinionated client wrapper** that makes interacting with the LLM Router service straightforward. Key components: Package Purpose data_models pydantic models that define the shape of payloads sent to the router (e.g. GenerativeConversationModel , ExtendedGenerativeConversationModel , utility models for question generation, translation, etc.). These models are shared with the API side, ensuring both client and server speak the same contract. client.py LLMRouterClient – a lightweight wrapper around the router’s HTTP API. It offers high‑level methods ( conversation_with_model , extended_conversation_with_model ) that accept either plain dictionaries or the aforementioned data‑model instances. The client handles payload validation, provider selection, error mapping, and response parsing. services Low‑level service classes ( ConversationService , ExtendedConversationService ) that perform the actual HTTP calls via HttpRequester . They are used internally by the client but can be reused directly if finer‑grained control is needed. exceptions.py Custom exception hierarchy ( LLMRouterError , AuthenticationError , RateLimitError , ValidationError ) that mirrors the router’s error semantics, making error handling in user code clean and explicit. utils/http.py HttpRequester – a small wrapper around requests providing retries, time‑outs and logging. It is the networking backbone for the client wrapper. In short, llm_router_lib provides both the data contract (the “schema”) and a convenient Pythonic client to consume the router service. Installation # The library targets Python 3.10.6 and uses a virtualenv . Install it in editable mode for development: # Clone the repository (if you haven't already) git clone https://github.com/radlab-dev-group/llm-router.git cd llm-router/llm_router_lib # Create and activate a virtual environment python3 -m venv .venv source .venv/bin/activate # Install the package and its dependencies pip install -e . All runtime dependencies ( requests , pydantic , rdl_ml_utils ) are declared in the project’s requirements.txt . Quick start # from llm_router_lib import LLMRouterClient # Initialise the client – point it at the router’s base URL client = LLMRouterClient ( api = \"http://localhost:8080/api\" , # router base URL token = \"YOUR_ROUTER_TOKEN\" , # optional, if router requires auth ) # Build a payload using the provided data model (validation is automatic) payload = { \"model_name\" : \"google/gemma-3-12b-it\" , \"user_last_statement\" : \"Hello, how are you?\" , \"temperature\" : 0.7 , \"max_new_tokens\" : 128 , } # Call the standard conversation endpoint response = client . conversation_with_model ( payload ) print ( response ) # → {'status': True, 'body': {...}} You can also pass a pydantic model instance directly: python from llm_router_lib.data_models.builtin_chat import GenerativeConversationModel model = GenerativeConversationModel( model_name=\"google/gemma-3-12b-it\", user_last_statement=\"Hello, how are you?\", temperature=0.7, max_new_tokens=128, ) response = client.conversation_with_model(model) Data models # All request payloads are defined in llm_router_lib/data_models . Common base: class BaseModelOptions ( BaseModel ): \"\"\"Options shared across many endpoint models.\"\"\" mask_payload : bool = False masker_pipeline : Optional [ List [ str ]] = None Conversation models # Model Required fields Optional / extra fields GenerativeConversationModel model_name , user_last_statement temperature , max_new_tokens , historical_messages , … ExtendedGenerativeConversationModel All of the above + system_prompt – Utility models for other built‑in endpoints (question generation, translation, article creation, context‑based answering, etc.) follow the same pattern and inherit from BaseModelOptions . Thin client wrapper ( LLMRouterClie…"},{"k":"helm-charts/index.html","t":"LLM‑Router Helm Chart","s":"Deployment","x":"# LLM‑Router Helm Chart Quick summary – This Helm chart deploys the LLM‑Router application with its optional Redis dependency. It works on any Kubernetes 1.19+ cluster and can be customized via values.yaml, environment variables, or the --…","h":["Table of Contents","Prerequisites","Dependencies","Installing the chart","Customising the deployment","1️⃣ Using a custom values file"],"b":"LLM‑Router Helm Chart # Quick summary – This Helm chart deploys the LLM‑Router application with its optional Redis dependency. It works on any Kubernetes 1.19+ cluster and can be customized via values.yaml, environment variables, or the --set flag. Table of Contents # Prerequisites Chart layout Dependencies Installing the chart Customising the deployment Prerequisites # Requirement Why we need it Kubernetes (v1.19 or newer) The chart creates Deployments, Services, Ingresses, ConfigMaps, … Helm (v3.x) Used to render and apply the chart Access to a container registry (e.g., Docker Hub, Quay, your private registry) The chart pulls the image defined in llm-router``values.yaml (Optional) cert‑manager If you enable TLS for the Ingress, cert‑manager will provision certificates ## Chart layout helm_charts/ └─ llm-router/ ├─ Chart.yaml # Chart metadata ├─ values.yaml # Default values ├─ values-dev.yaml # Development‑specific overrides ├─ templates/ │ ├─ _helpers.tpl # Helper functions (name, labels, etc.) │ ├─ deployment.yaml # Deployment definition │ ├─ service.yaml # Service definition │ ├─ ingress.yaml # Ingress definition (optional) │ ├─ configmap.yaml # ConfigMap for runtime env vars │ └─ configmap-models.yaml # ConfigMap for `models-config.json` └─ charts/ └─ redis-23.2.12.tgz # Bitnami Redis sub‑chart (dependency) Dependencies # The chart depends on Redis. The dependency is declared in Chart.yaml and pulled automatically when you run helm dependency update. # Chart.yaml dependencies : - name : redis version : \"23.2.12\" repository : \"oci://registry-1.docker.io/bitnamicharts\" condition : redis.enabled Adding / updating dependencies # From the chart root (helm_charts/llm-router) helm dependency update . You can disable the Redis sub‑chart with: --set redis.enabled = false Installing the chart # helm upgrade --install my-llm-router ./helm_charts/llm-router \\ --namespace my-namespace \\ --create-namespace Customising the deployment # You can tailor the LLM‑Router Helm chart to your environment in three different ways. Pick the approach that best fits the task at hand. Method When to use Edit values.yaml and run helm upgrade Long‑term, version‑controlled configuration that lives in source control. Pass --set flags on the command line One‑off tweaks, CI pipelines, quick experiments, or when you need to override just a handful of values. Provide a custom values file ( -f my-values.yaml ) Complex overrides, reusable profiles, or when you prefer to keep the changes in a separate, shareable file. Below are concrete examples for the two most common scenarios: using a custom values file and using --set flags. 1️⃣ Using a custom values file # Create a file (e.g. my-values.yaml ) with the settings you want to override: # my-values.yaml ingress : enabled : true className : traefik hosts : - host : my-custom.cluster.local paths : - path : / pathType : Prefix tls : - hosts : - my-custom.cluster.local secretName : llm-router-tls image : tag : latest # pull the latest container image web_host : my-custom.cluster.local # the hostname used by the app &amp; Ingress Deploy (or upgrade) the chart with this file: ```shell script helm upgrade --install llm-router ./helm_charts/llm-router \\ -f my-values.yaml --- ### 2️⃣ Using `--set` flags If you only need to change a few values, the `--set` syntax is handy: ```shell script helm upgrade --install llm-router ./helm_charts/llm-router \\ --set ingress.enabled=true \\ --set web_host=llm2.k3s.radlab.dev \\ --set image.tag=latest Flag Effect ingress.enabled=true Turns on the Ingress resource (otherwise it’s omitted). llm-1.cluster.local Sets the host name used both in the Ingress rule and the application’s configuration ( LLM_ROUTER_WEB_HOST ). image.tag=latest Pulls the latest image tag instead of the default chart‑version tag."},{"k":"examples/index.html","t":"Integration Examples with LLM Router","s":"Examples","x":"# Integration Examples with LLM Router This directory contains example boilerplates that demonstrate how easy it is to integrate popular LLM libraries with the router by simply switching the host. --- ## Available Examples - **[LlamaIndex]…","h":["Available Examples","Core Principle","Quick Start","Example Structure","Full Stack with Local Models","Additional Information"],"b":"Integration Examples with LLM Router # This directory contains example boilerplates that demonstrate how easy it is to integrate popular LLM libraries with the router by simply switching the host. Available Examples # LlamaIndex – Integration with LlamaIndex (GPT Index) Using LlamaIndex with Local Models – Guide for mapping OpenAI model names to local models via the router. LangChain – Integration with LangChain OpenAI SDK – Direct integration with the OpenAI Python SDK LiteLLM – Integration with LiteLLM Haystack – Integration with Haystack Core Principle # All examples work on the same principle: just change base_url / api_base to the address of your router , and the router will automatically: ✅ Distribute traffic among available providers ✅ Perform load balancing ✅ Provide health checking ✅ Supply monitoring and metrics ✅ Handle streaming and non‑streaming responses Quick Start # Each example can be run directly: ```shell script LlamaIndex # python examples/llamaindex_example.py LangChain # python examples/langchain_example.py OpenAI SDK # python examples/openai_example.py LiteLLM # python examples/litellm_example.py Haystack # python examples/haystack_example.py ``` Example Structure # Each example includes: Basic configuration – how to point the library at the router Streaming – handling streaming responses Non‑streaming – handling full responses Error handling – managing errors Full Stack with Local Models # The quick‑start guides for running the full stack with local models are included in the repository: Gemma 3 12B‑IT – README Bielik 11B‑v2.3‑Instruct – README These guides walk you through: Installing vLLM and the respective model. Setting up LLM‑Router with the provided models-config.json . Testing the end‑to‑end flow (router → vLLM). Follow the linked README files for step‑by‑step instructions to launch a complete stack locally. Additional Information # Learn more about the router: Main README API Documentation Endpoints Overview Load‑Balancing Strategies"},{"k":"examples/quickstart/index.html","t":"Full Stack with Local Models","s":"Examples","x":"## Full Stack with Local Models The quick‑start guides for running the full stack with **local models** are included in the repository: - **Gemma 3 12B‑IT** – [README](google-gemma3-12b-it/README.md) - **Bielik 11B‑v2.3‑Instruct** – [READM…","h":["Full Stack with Local Models"],"b":"Full Stack with Local Models # The quick‑start guides for running the full stack with local models are included in the repository: Gemma 3 12B‑IT – README Bielik 11B‑v2.3‑Instruct – README"},{"k":"examples/readme-llamaindex.html","t":"Using LlamaIndex with Local Models via the LLM‑Router","s":"Examples","x":"## Using LlamaIndex with Local Models via the LLM‑Router LlamaIndex’s `OpenAI` wrapper expects **OpenAI‑style model names** (e.g. `gpt-3.5-turbo`, `gpt-4`). When you want to run *local* models (Gemma, Ollama, vLLM, etc.) behind an LLM‑Rout…","h":["Using LlamaIndex with Local Models via the LLM‑Router","Why the mapping is required","Router configuration","How it works end‑to‑end","Checklist for a working setup","TL;DR"],"b":"Using LlamaIndex with Local Models via the LLM‑Router # LlamaIndex’s OpenAI wrapper expects OpenAI‑style model names (e.g. gpt-3.5-turbo , gpt-4 ). When you want to run local models (Gemma, Ollama, vLLM, etc.) behind an LLM‑Router, the router must translate those OpenAI names to the actual model identifiers used by the local providers. Why the mapping is required # LlamaIndex validates the model name against a known list of OpenAI models to decide whether the call is a chat request, to obtain context‑window size, etc. If the wrapper receives a name it does not recognise (e.g. google/gemma-3-12b-it ), it raises a ValueError like: ValueError: Unknown model 'google/gemma-3-12b-it'. Please provide a valid OpenAI model name … By keeping the OpenAI name in the request and letting the router forward the request to the appropriate backend, you get the best of both worlds: LlamaIndex works unchanged, and the traffic is routed to your local model. Router configuration # The router’s configuration must contain a section that lists OpenAI model names ( openai_models ). Each entry defines one or more providers that actually serve the model. The key points are: Field Meaning id Arbitrary identifier for the provider instance. api_host URL where the provider’s inference server is reachable. api_type The protocol the provider uses ( vllm , ollama , …). model_path The local model identifier (e.g. google/gemma-3-12b-it ). weight Relative load‑balancing weight when several providers are listed. Example snippet # \"openai_models\" : { (...) \"gpt-3.5-turbo\" : { \"providers\" : [ { \"id\" : \"gpt_35_turbo-gemma3_12b-vllm-71:7000\" , \"api_host\" : \"http://192.168.100.71:7000/\" , \"api_token\" : \"\" , \"api_type\" : \"vllm\" , \"input_size\" : 4096 , \"model_path\" : \"google/gemma-3-12b-it\" , \"weight\" : 1.0 }, { \"id\" : \"gpt_35_turbo-gemma3_12b-vllm-71:7001\" , \"api_host\" : \"http://192.168.100.71:7001/\" , \"api_token\" : \"\" , \"api_type\" : \"vllm\" , \"input_size\" : 4096 , \"model_path\" : \"google/gemma-3-12b-it\" , \"weight\" : 1.0 } ] }, \"gpt-4\" : { \"providers\" : [ { \"id\" : \"gpt-4-gpt-oss-20b-ollama-66:11434\" , \"api_host\" : \"http://192.168.100.66:11434\" , \"api_token\" : \"\" , \"api_type\" : \"ollama\" , \"input_size\" : 256000 , \"model_path\" : \"gpt-oss:120b\" } ] } }, ... \"active_models\" : { \"openai_models\" : [ \"gpt-3.5-turbo\" , \"gpt-4\" ... ] } How it works end‑to‑end # LlamaIndex code creates an OpenAI client with a model name that the wrapper knows, e.g.: llm = OpenAI ( model = \"gpt-3.5-turbo\" , api_base = \"http://localhost:8080\" , api_key = \"not-needed\" ) The request is sent to the router ( api_base ). The router looks up \"gpt-3.5-turbo\" in openai_models , selects a provider, and forwards the request to the provider’s api_host . The provider receives the request with model_path=\"google/gemma-3-12b-it\" (or \"gpt-oss:120b\" for gpt-4 ) and runs the local model. The response travels back through the router to LlamaIndex, which treats it exactly like an OpenAI response. Checklist for a working setup # Router is running and reachable at the URL you pass to api_base . openai_models section contains all OpenAI names you intend to use from LlamaIndex. Each OpenAI name maps to at least one provider with the correct model_path . The active_models.openai_models list includes the names you want to expose (otherwise the router will ignore them). Local model servers (vLLM, Ollama, etc.) are up and listening on the api_host URLs specified. TL;DR # *Keep the model names you give to LlamaIndex identical to the OpenAI names defined in the router configuration. The router will translate those names to the real local model identifiers ( model_path ). This mapping is the only thing needed for LlamaIndex to work seamlessly with any self�"},{"k":"examples/quickstart/google-gemma3-12b-it/vllm.html","t":"vLLM + google/gemma‑3‑12b‑it – Quick‑Start Guide (Ubuntu)","s":"Examples","x":"# vLLM + `google/gemma‑3‑12b‑it` – Quick‑Start Guide (Ubuntu) > **Prerequisites** > - Ubuntu 20.04 or newer > - Python 3.10 (our project uses 3.10.6) > - `virtualenv` (installed) > - CUDA 11.8 + GPU **or** a CPU‑only setup --- ## 1️⃣ Creat…","h":["1️⃣ Create &amp; activate a virtual environment","2️⃣ Install vLLM","Verify the installation","3️⃣ Obtain the model google/gemma-3-12b-it","4️⃣ Run the vLLM server","5️⃣ Test the endpoint","6️⃣ Handy tips","🎉 All set!"],"b":"vLLM + google/gemma‑3‑12b‑it – Quick‑Start Guide (Ubuntu) # Prerequisites - Ubuntu 20.04 or newer - Python 3.10 (our project uses 3.10.6) - virtualenv (installed) - CUDA 11.8 + GPU or a CPU‑only setup 1️⃣ Create &amp; activate a virtual environment # mkdir -p ~/vllm-gemma &amp;&amp; cd ~/vllm-gemma python3 -m venv .venv source .venv/bin/activate This creates an optional project directory, sets up a Python virtual environment in .venv , and activates it (you’ll see (.venv) in the prompt). 2️⃣ Install vLLM # pip install --upgrade pip pip install \"vllm[cuda]\" The above installs the latest pip and then installs vLLM with GPU support (the appropriate CUDA libraries are detected automatically). If you have no GPU, install the CPU‑only version instead: pip install vllm[cpu] . Verify the installation # python -c \"import vllm; print(vllm.__version__)\" You should see a version string such as 0.11.2 . 3️⃣ Obtain the model google/gemma-3-12b-it # mkdir -p ./google/gemma-3-12b-it pip install huggingface_hub hf download google/gemma-3-12b-it \\ --local-dir ./google/gemma-3-12b-it This creates a folder for the model, installs the Hugging Face CLI, and downloads the model files into ./google/gemma-3-12b-it . The files are cached under ~/.cache/huggingface/hub by default; you can keep a local copy to avoid re‑downloads. 4️⃣ Run the vLLM server # Copy the ready‑to‑use Bash script ( llm-router/examples/quickstart/google-gemma3-12b-it/run-gemma-3-12b-it-vllm.sh ) to directory wgen the vLLM will be started with the Gemma 3 model. cp path/to/llm-router/examples/quickstart/google-gemma3-12b-it/run-gemma-3-12b-it-vllm.sh . bash run-gemma-3-12b-it-vllm.sh Tip: Run the server inside a tmux or screen session so it stays alive even if you disconnect from the terminal. 5️⃣ Test the endpoint # INFO : curl and jq are system utilities. curl http://localhost:7000/v1/chat/completions \\ -H \"Content-Type: application/json\" \\ -d '{ \"model\": \"google/gemma-3-12b-it\", \"messages\": [{\"role\": \"user\", \"content\": \"Hello, how are you?\"}], \"max_tokens\": 100 }' | jq You should receive a JSON response containing the model’s generated text, for example: { \"id\" : \"chatcmpl-e30bed0db9f9440a8aec14bd287ca63d\" , \"object\" : \"chat.completion\" , \"created\" : 1764516430 , \"model\" : \"google/gemma-3-12b-it\" , \"choices\" : [ { \"index\" : 0 , \"message\" : { \"role\" : \"assistant\" , \"content\" : \"Hello! I'm doing well, thank you for asking! As an AI, I don't experience feelings like humans do, but everything is running smoothly and I'm ready to chat. 😊\\n\\nHow are *you* doing today?\" , \"refusal\" : null , \"annotations\" : null , \"audio\" : null , \"function_call\" : null , \"tool_calls\" : [], \"reasoning\" : null , \"reasoning_content\" : null }, \"logprobs\" : null , \"finish_reason\" : \"stop\" , \"stop_reason\" : 106 , \"token_ids\" : null } ], \"service_tier\" : null , \"system_fingerprint\" : null , \"usage\" : { \"prompt_tokens\" : 15 , \"total_tokens\" : 66 , \"completion_tokens\" : 51 , \"prompt_tokens_details\" : null }, \"prompt_logprobs\" : null , \"prompt_token_ids\" : null , \"kv_transfer_params\" : null } 6️⃣ Handy tips # Topic Recommendation Memory google/gemma‑3‑12b‑it needs ~24GB VRAM. Use --cpu-offload (if supported) for larger models or when GPU memory is limited. Cache location Set HF_HOME=$PWD/.cache/huggingface to keep all model files inside the project directory. Parallelism Export TOKENIZERS_PARALLELISM=false to silence tokenizer warnings. GPU selection export CUDA_VISIBLE_DEVICES=0 (or another index) when multiple GPUs are present. Update pip install -U vllm refreshes the library; the next server start will pull newer model files if available. Deactivate When done, simply run deactivate to leave the virtual environment. 🎉 All set! # You now have a fully functional OpenAI‑compatible API powered by vLLM and the google/gemma‑3‑12b‑it model."},{"k":"examples/quickstart/speakleash-bielik-11b-v2-3-instruct/vllm.html","t":"vLLM + speakleash/Bielik-11B-v2.3-Instruct – Przewodnik Szybkiego Startu (Ubuntu)","s":"Examples","x":"# vLLM + `speakleash/Bielik-11B-v2.3-Instruct` – Przewodnik Szybkiego Startu (Ubuntu) > **Wymagania wstępne** > - Ubuntu 20.04 lub nowszy > - Python 3.10 (w projekcie używamy 3.10.6) > - `virtualenv` (zainstalowany) > - CUDA 11.8 + GPU **l…","h":["1️⃣ Utwórz i aktywuj wirtualne środowisko","2️⃣ Zainstaluj vLLM","Sprawdź instalację","4️⃣ Przygotuj środowisko do pobierania modelu","6️⃣ Pobierz model speakleash/Bielik-11B-v2.3-Instruct","(Opcjonalnie) Ustaw własny katalog cache","7️⃣ Uruchom serwer vLLM","8️⃣ Przetestuj endpoint","9️⃣ Przydatne wskazówki","🎉 Gotowe!"],"b":"vLLM + speakleash/Bielik-11B-v2.3-Instruct – Przewodnik Szybkiego Startu (Ubuntu) # Wymagania wstępne - Ubuntu 20.04 lub nowszy - Python 3.10 (w projekcie używamy 3.10.6) - virtualenv (zainstalowany) - CUDA 11.8 + GPU lub środowisko tylko CPU 1️⃣ Utwórz i aktywuj wirtualne środowisko # mkdir -p ~/bielik &amp;&amp; cd ~/bielik python3 -m venv .venv source .venv/bin/activate Powoduje to utworzenie katalogu projektu, przygotowanie wirtualnego środowiska w folderze .venv oraz jego aktywację (w promptcie pojawi się (.venv) ). 2️⃣ Zainstaluj vLLM # pip install --upgrade pip pip install \"vllm[cuda]\" Instalacja najnowszej wersji vLLM z obsługą GPU (CUDA zostanie wykryte automatycznie). Jeśli nie masz GPU, użyj wersji CPU: pip install vllm[cpu] . Sprawdź instalację # python -c \"import vllm; print(vllm.__version__)\" Powinieneś zobaczyć wersję, np. 0.11.2 . 4️⃣ Przygotuj środowisko do pobierania modelu # pip install huggingface_hub 6️⃣ Pobierz model speakleash/Bielik-11B-v2.3-Instruct # mkdir -p ./speakleash/Bielik-11B-v2.3-Instruct hf download speakleash/Bielik-11B-v2.3-Instruct \\ --local-dir ./speakleash/Bielik-11B-v2.3-Instruct Model zostanie pobrany do wskazanego katalogu. Pliki będą także buforowane domyślnie w ~/.cache/huggingface/hub . (Opcjonalnie) Ustaw własny katalog cache # Jeśli chcesz, aby wszystkie modele były przechowywane wewnątrz projektu, ustaw zmienną przed pobraniem: export HF_HOME=$PWD/.cache/huggingface # np. ./bielik/.cache/huggingface 7️⃣ Uruchom serwer vLLM # Skopiuj gotowy skrypt Bash (przykładowa ścieżka – dostosuj do swojego projektu): cp path/to/llm-router/examples/quickstart/speakleash-bielik-11b-v2_3-Instruct/run-bielik-11b-v2_3-vllm.sh . bash run-bielik-11b-v2_3-vllm.sh Wskazówka: uruchom serwer w sesji tmux lub screen , aby pozostawał aktywny po rozłączeniu się z terminalem. 8️⃣ Przetestuj endpoint # INFO : curl i jq to narzędzia systemowe. curl http://localhost:7000/v1/chat/completions \\ -H \"Content-Type: application/json\" \\ -d '{ \"model\": \"speakleash/Bielik-11B-v2.3-Instruct\", \"messages\": [{\"role\": \"user\", \"content\": \"Cześć, jak się masz?\"}], \"max_tokens\": 100 }' | jq Powinieneś otrzymać odpowiedź w formacie JSON, np.: { \"id\" : \"chatcmpl-xxxx\" , \"object\" : \"chat.completion\" , \"created\" : 1764516430 , \"model\" : \"speakleash/Bielik-11B-v2.3-Instruct\" , \"choices\" : [ { \"index\" : 0 , \"message\" : { \"role\" : \"assistant\" , \"content\" : \"Cześć! Jestem w pełni sprawny i gotowy do rozmowy. Jak mogę Ci pomóc?\" }, \"finish_reason\" : \"stop\" } ], \"usage\" : { \"prompt_tokens\" : 15 , \"total_tokens\" : 66 , \"completion_tokens\" : 51 } } 9️⃣ Przydatne wskazówki # Temat Rekomendacja Pamięć speakleash/Bielik-11B-v2.3-Instruct potrzebuje ok. 24GB VRAM. Użyj --cpu-offload (jeśli wspierane) przy ograniczonej pamięci GPU. Lokalizacja cache Ustaw HF_HOME=$PWD/.cache/huggingface , aby wszystkie pliki modelu znajdowały się w katalogu projektu. Równoległość tokenizera export TOKENIZERS_PARALLELISM=false wyciszy ostrzeżenia tokenizera. Wybór GPU export CUDA_VISIBLE_DEVICES=0 (lub inny indeks) przy wielu kartach GPU. Aktualizacja pip install -U vllm odświeża bibliotekę; przy następnym uruchomieniu serwera zostaną pobrane nowsze pliki modelu, jeśli są dostępne. Dezaktywacja Po zakończeniu pracy wystarczy wpisać deactivate , aby opuścić wirtualne środowisko. 🎉 Gotowe! # Masz już w pełni działające API kompatybilne z OpenAI, oparte na vLLM i modelu speakleash/Bielik-11B-v2.3-Instruct ."},{"k":"examples/quickstart/speakleash-bielik-11b-v2-3-instruct/index.html","t":"🚀 Przewodnik Szybkiego Startu dla speakleash/Bielik-11B-v2.3-Instruct z vLLM & LLM‑Router","s":"Examples","x":"# 🚀 **Przewodnik Szybkiego Startu** dla `speakleash/Bielik-11B-v2.3-Instruct` z **vLLM** & **LLM‑Router** Ten przewodnik prowadzi Cię krok po kroku przez: 1. **Instalację vLLM** i modelu `speakleash/Bielik-11B-v2.3-Instruct`. 2. **Instalac…","h":["📋 Wymagania wstępne","1️⃣ Utworzenie i aktywacja wirtualnego środowiska","6️⃣ Przygotowanie konfiguracji routera","7️⃣ Test pełnego stosu (router → vLLM)","🎉 Co dalej?"],"b":"🚀 Przewodnik Szybkiego Startu dla speakleash/Bielik-11B-v2.3-Instruct z vLLM &amp; LLM‑Router # Ten przewodnik prowadzi Cię krok po kroku przez: Instalację vLLM i modelu speakleash/Bielik-11B-v2.3-Instruct . Instalację LLM‑Router (bramki API). Uruchomienie routera z konfiguracją modeli dostarczoną w models-config.json . Wszystkie polecenia zakładają, że pracujesz na systemie Unix‑like (Linux/macOS) z Python 3.10.6 , virtualenv oraz ( opcjonalnie) kartą GPU obsługującą CUDA 11.8. 📋 Wymagania wstępne # Wymaganie Szczegóły OS Ubuntu 20.04 + (lub dowolna nowsza dystrybucja Linux/macOS) Python 3.10.6 (domyślna wersja projektu) GPU CUDA 11.8 + (minimum 12 GB VRAM) lub środowisko CPU‑only Narzędzia git , curl , jq (opcjonalnie, przydatne do testowania) Sieć Dostęp do PyPI oraz Hugging Face w celu pobrania modelu 1️⃣ Utworzenie i aktywacja wirtualnego środowiska # ```shell script (opcjonalnie) utwórz katalog demo i przejdź do niego # mkdir -p ~/bielik-demo &amp;&amp; cd $_ Inicjalizacja venv # python3 -m venv .venv source .venv/bin/activate Aktualizacja pip (zawsze dobry pomysł) # pip install --upgrade pip --- ## 2️⃣ Instalacja **vLLM** oraz pobranie modelu Bielik &gt; Pełną instrukcję znajdziesz w pliku [`VLLM.md`](./VLLM.md). --- ## 3️⃣ **Uruchomienie serwera vLLM** Skopiuj do bieżącego katalogu dostarczony skrypt Bash (dostosuj ścieżkę, jeśli potrzebujesz) i uruchom go: ```shell script cp path/to/llm-router/examples/quickstart/speakleash-bielik-11b-v2_3-Instruct/run-bielik-11b-v2_3-vllm.sh . chmod +x run-bielik-11b-v2_3-vllm.sh # Uruchom (warto w tmux/screen) ./run-bielik-11b-v2_3-vllm.sh Serwer nasłuchuje na http://0.0.0.0:7000 i udostępnia endpoint zgodny z OpenAI pod /v1/chat/completions . Możesz szybko go przetestować: ```shell script curl http://localhost:7000/v1/chat/completions \\ -H \"Content-Type: application/json\" \\ -d '{ \"model\": \"speakleash/Bielik-11B-v2.3-Instruct\", \"messages\": [{\"role\": \"user\", \"content\": \"Cześć, jak się masz?\"}], \"max_tokens\": 100 }' | jq Powinieneś otrzymać odpowiedź w formacie JSON. --- ## 4️⃣ Instalacja **LLM‑Router** ```shell script # Sklonuj repozytorium (jeśli jeszcze go nie masz) git clone https://github.com/radlab-dev-group/llm-router.git cd llm-router # Instalacja core + API (w tym samym venv) pip install .[api] # (Opcjonalnie) wsparcie dla Prometheus pip install .[api,metrics] Uwaga: Router używa tego samego wirtualnego środowiska, które utworzyłeś wcześniej, więc wszystkie zależności pozostają odizolowane. 6️⃣ Przygotowanie konfiguracji routera # Plik models-config.json znajdujący się w katalogu speakleash‑bielik już zawiera definicję naszego modelu: { \"speakleash_models\" : { \"speakleash/Bielik-11B-v2.3-Instruct\" : { \"providers\" : [ { \"id\" : \"bielik-11B_v2_3-vllm-local:7000\" , \"api_host\" : \"http://localhost:7000/\" , \"api_type\" : \"vllm\" , \"input_size\" : 56000 , \"weight\" : 1.0 } ] } }, \"active_models\" : { \"speakleash_models\" : [ \"speakleash/Bielik-11B-v2.3-Instruct\" ] } } Skopiuj go (lub przenieś) do katalogu resources/configs/ routera: ```shell script mkdir -p resources/configs cp path/to/speakleash-bielik/models-config.json resources/configs/ --- ## 6️⃣ Uruchomienie **LLM‑Router** ### Lokalny Gunicorn W repozytorium znajduje się pomocniczy skrypt `run-rest-api-gunicorn.sh`. Upewnij się, że jest wykonywalny, a następnie go uruchom: ```shell script chmod +x run-rest-api-gunicorn.sh ./run-rest-api-gunicorn.sh Domyślne zmienne środowiskowe (można zmienić w skrypcie): Zmienna Domyślna wartość Opis LLM_ROUTER_SERVER_TYPE gunicorn Backend serwera LLM_ROUTER_SERVER_PORT 8080 Port nasłuchiwania routera LLM_ROUTER_MODELS_CONFIG resources/configs/models-config.json Ścieżka do pliku konfiguracyjnego LLM_ROUTER_USE_PROMETHEUS 1 (jeśli zainstalowano metrics ) Włącza endpoint /api/metrics Router będzie dostępny pod http://0.0.0.0:8080/api . Pełna lista dostępnych zmiennych środowiskowych znajduje się w opisie zmiennych środowiskowych 7️⃣ Test pełnego stosu (router → vLLM) # ```shell script curl http://loc…"},{"k":"examples/quickstart/google-gemma3-12b-it/index.html","t":"🚀 Quick‑Start Guide for google/gemma-3-12b‑it with vLLM & LLM‑Router","s":"Examples","x":"# 🚀 Quick‑Start Guide for `google/gemma-3-12b‑it` with **vLLM** & **LLM‑Router** This guide walks you through: 1. **Installing vLLM** and the `google/gemma‑3‑12b‑it` model. 2. **Installing LLM‑Router** (the API gateway). 3. **Running the r…","h":["📋 Prerequisites","1️⃣ Set up a virtual environment","5️⃣ Prepare the router configuration","7️⃣ Test the full stack (router → vLLM)","🎉 What’s next?"],"b":"🚀 Quick‑Start Guide for google/gemma-3-12b‑it with vLLM &amp; LLM‑Router # This guide walks you through: Installing vLLM and the google/gemma‑3‑12b‑it model. Installing LLM‑Router (the API gateway). Running the router with the model configuration provided in models-config.json . All commands assume you are working on a Unix‑like system (Linux/macOS) with Python 3.10.6 and virtualenv available. 📋 Prerequisites # Requirement Details OS Ubuntu 20.04 + (or any recent Linux/macOS) Python 3.10.6 (project’s default) GPU CUDA 11.8 + (≥ 24 GB VRAM) or CPU‑only setup Tools git , curl , jq (optional but handy for testing) Network Ability to pull Docker images / PyPI packages and download the model from Hugging Face 1️⃣ Set up a virtual environment # ```shell script Create a directory for the whole demo (optional) # mkdir -p ~/gemma3-demo &amp;&amp; cd $_ Initialise the venv # python3 -m venv .venv source .venv/bin/activate Upgrade pip (always a good idea) # pip install --upgrade pip --- ## 2️⃣ Install **vLLM** and download the Gemma 3 model &gt; **See the full step‑by‑step instructions in** [`VLLM.md`](./VLLM.md). --- ## 3️⃣ Run the **vLLM** server Copy the helper script (or run the command manually) inside the demo directory: ```shell script # If you have the script `run-gemma-3-12b-it-vllm.sh` in the repo: cp path/to/llm-router/examples/quickstart/google-gemma3-12b-it/run-gemma-3-12b-it-vllm.sh . chmod +x run-gemma-3-12b-it-vllm.sh # Start the server (you may want to use tmux/screen) ./run-gemma-3-12b-it-vllm.sh The server will listen on http://0.0.0.0:7000 and expose an OpenAI‑compatible endpoint at /v1/chat/completions . You can quickly test it: ```shell script curl http://localhost:7000/v1/chat/completions \\ -H \"Content-Type: application/json\" \\ -d '{ \"model\": \"google/gemma-3-12b-it\", \"messages\": [{\"role\": \"user\", \"content\": \"Hello, how are you?\"}], \"max_tokens\": 100 }' | jq You should receive a JSON payload with the model’s generated text. --- ## 4️⃣ Install **LLM‑Router** ### Local install ```shell script # Clone the router repository (if you haven’t already) git clone https://github.com/radlab-dev-group/llm-router.git cd llm-router # Install the core library + API wrapper (includes the REST server) pip install .[api] # (Optional) Install Prometheus metrics support pip install .[api,metrics] Note: The router uses the same virtual environment you created earlier, so all dependencies stay isolated. 5️⃣ Prepare the router configuration # The example repository already ships a models-config.json that points to the locally running vLLM instance: { \"google_models\" : { \"google/gemma-3-12b-it\" : { \"providers\" : [ { \"id\" : \"gemma3_12b-vllm-local:7000\" , \"api_host\" : \"http://localhost:7000/\" , \"api_type\" : \"vllm\" , \"input_size\" : 56000 , \"weight\" : 1.0 } ] } }, \"active_models\" : { \"google_models\" : [ \"google/gemma-3-12b-it\" ] } } Copy it (or edit the path) to the router’s resources/configs/ directory: ```shell script mkdir -p resources/configs cp path/to/google-gemma3-12b-it/models-config.json resources/configs/ --- ## 6️⃣ Run the **LLM‑Router** ### Local Gunicorn The helper script `run-rest-api-gunicorn.sh` sets a sensible default environment. You can use it directly or export the variables yourself. ```shell script # Make the script executable (if needed) chmod +x path/to/run-rest-api-gunicorn.sh # Run the router ./run-rest-api-gunicorn.sh Key environment variables (already defined in the script) you may want to adjust: Variable Default Meaning LLM_ROUTER_SERVER_TYPE gunicorn Server backend (gunicorn, flask, waitress) LLM_ROUTER_SERVER_PORT 8080 Port on which the router listens LLM_ROUTER_MODELS_CONFIG resources/configs/models-config.json Path to the JSON file above LLM_ROUTER_PROMPTS_DIR resources/prompts Prompt‑template directory (optional) LLM_ROUTER_BALANCE_STRATEGY first_available Load‑balancing strategy LLM_ROUTER_USE_PROMETHEUS 1 (if you installed metrics) Enable /api/metrics endpoint After the script starts, the router will be re…"},{"k":"changelog.html","t":"Changelog","s":"Release notes","x":"## Changelog | Version | Changelog | |---------|-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------…","h":["Changelog"],"b":"Changelog # Version Changelog 0.0.1 Initialization, License, setup, interface for each endpoint and sample ping EP. Autoloader of builtin endpoints and for the future implementations. 0.0.2 Add base models for api call (module llm_proxy_rest.data_models with error.py handling. Decorators to check required params and to measure the response time. 0.0.3 Proper AutoLoading for each found endpoint. Implementation of ApiTypesDispatcher , ApiModelConfig , ModelHandler . Ollama endpoints: / , tags . Added endpoint to full proxy with params. Streaming in case when external api provides stream. 0.0.4 All llama-service endpoints are refactored to llm-proxy-api . Refactoring base ep_run method. Proper handling system message, prompt name, model etc. 0.1.0 Repository name changed from llm-proxy-api to llm-router . Added class HttpRequestExecutor to handle http requests from EndpointWithHttpRequestI . Handled routing between any models: openai -&gt; ollama and ollama -&gt; openai 0.1.1 Prometheus metrics logging. Workers/Threads/Workers class is able to set by environments. Streaming fixes. Multi-providers for single model with default-balanced strategy. 0.2.0 Add balancing strategies: balanced , weighted , dynamic_weighted and first_available which works for streaming and non streaming requests. Included Prometheus metrics logging via /metrics endpoint. First stage of llm_router_lib library, to simply usage of llm-router-api . 0.2.1 Fix stream: OpenAI-&gt;Ollama, Ollama-&gt;OpenAI. Add Redis caching of availability of model providers (when using first_available strategy). Add llm_router_web module with simple flask-based frontend to manage llm-router config files. 0.2.2 Update dockerfile and requirements. Fix routing with vLLM. 0.2.3 New web configurator: Handling projects, configs for each user separately. First Available strategy is more powerful, a lot of improvements to efficiency. 0.2.4 Anonymizer module, integration anonymization with any endpoint (using dynamic payload analysis and full payload anonymisation), dedicated /api/anonymize_text endpoint as memory only anonymization. Whole router may be run in FORCE_ANONYMISATION mode. 0.3.0 Anonymization available with three strategies: fast_masker , genai , prov_masker . 0.3.1 Refactoring lb.strategies to be more flexible modular. Introduced MaskerPipeline and GuardrailPipeline both configured via env. Removed genai-based masking endpoint. 0.4.0 The main repository is divided into dedicated ones: plugins, services, web — separate repositories. Clean up the whole repository. Examples of integration with llamaindex, langchain, openai, litellm and haystack. 0.4.1 Audit log is stored using GPG. Add bash script ( scripts/gen_and_export_gpg.sh to prepare GPG keys and simple scripts/decrypt_auditor_logs.sh to decrypt encrypted audit logs. Moved core functionality from base to module core module. Quickstart. 0.4.2 Fix first_available_optim Strategy. Add KeepAliveMonitor to periodically pings model endpoints to keep them warm. 0.4.3 Add custom Prometheus metrices for logging masker/guardrail inidents. Fix OpenAI compatible v1 /models endpoint. Introduce monitors: services and keep alive models. Fixed guardrail retunr in case when streaming."}]}