{"version":"0.4.0","pages":[{"k":"overview.html","t":"Overview","s":"Getting started","x":"# LLM Router - Open-Source AI Gateway for Local and Cloud LLM Infrastructure [**LLM Router**](https://llm-router.cloud) is a service that can be deployed on‑premises or in the cloud. It adds a layer between any application and the LLM prov…","h":["🧩 Boilerplates","✨ Key Features","📦 Quick Start","1️⃣ Create &amp; activate a virtual environment","2️⃣ Minimum required environment variable","🛠️ Configuration (via environment)","⚖️ Load Balancing Strategies","🛣️ Endpoints Overview","⚙️ Configuration Details","🔧 Development","📜 License","📚 Changelog"],"b":"LLM Router - Open-Source AI Gateway for Local and Cloud LLM Infrastructure # LLM Router is a service that can be deployed on‑premises or in the cloud. It adds a layer between any application and the LLM provider. In real time it controls traffic, distributes a load among providers of a specific LLM, and enables analysis of outgoing requests from a security perspective (masking, anonymization, prohibited content). It is an open‑source solution (Apache 2.0) that can be launched instantly by running a ready‑made image in your own infrastructure. llm_router_api provides a unified REST proxy that can route requests to any supported LLM backend ( OpenAI‑compatible, Ollama, vLLM, LM Studio, etc.), with built‑in load‑balancing, health checks, streaming responses and optional Prometheus metrics. llm_router_lib is a Python SDK that wraps the API with typed request/response models, automatic retries, token handling and a rich exception hierarchy, letting developers focus on application logic rather than raw HTTP calls. llm_router_web offers ready‑to‑use Flask UIs – an anonymizer UI that masks sensitive data and a configuration manager for model/user settings – demonstrating how to consume the router from a browser. llm_router_plugins (e.g., the fast_masker plugin) deliver a rule‑based text anonymisation engine with a comprehensive set of Polish‑specific masking rules (emails, IPs, URLs, phone numbers, PESEL, NIP, KRS, REGON, monetary amounts, dates, etc.) and an extensible architecture for custom rules and validators. llm_router_services provides HTTP services that implement the core functionality used by the LLM‑Router’s plugin system. The services expose guardrail and masking capabilities through Flask applications. All components run on Python 3.10+ using virtualenv and require only the listed dependencies, making the suite easy to install, extend, and deploy in both development and production environments. 🧩 Boilerplates # For a detailed explanation of each example’s purpose, structure, and how the boilerplates are organized, see the main project README: Main README – Boilerplate Overview – examples LlamaIndex Boilerplate Details – README ✨ Key Features # Feature Description Unified REST interface One endpoint schema works for OpenAI‑compatible, Ollama, vLLM and any future provider. Provider‑agnostic streaming The stream flag (default true ) controls whether the proxy forwards chunked responses as they arrive or returns a single aggregated payload. Built‑in prompt library Language‑aware system prompts stored under resources/prompts can be referenced automatically. Dynamic model configuration JSON file ( models-config.json ) defines providers, model name, default options and per‑model overrides. Request validation Pydantic models guarantee correct payloads; errors are returned with clear messages. Structured logging Configurable log level, filename, and optional JSON formatting. Health &amp; metadata endpoints /ping (simple 200 OK) and /tags (available model tags/metadata). Simple deployment One‑liner run script or python -m llm_proxy_rest.rest_api . Extensible conversation formats Basic chat, conversation with system prompt, and extended conversation with richer options (e.g., temperature, top‑k, custom system prompt). Multi‑provider model support Each model can be backed by multiple providers (VLLM, Ollama, OpenAI) defined in models-config.json . Provider selection abstraction ProviderChooser delegates to a configurable strategy, enabling easy swapping of load‑balancing, round‑robin, weighted‑random, etc. Load‑balanced default strategy LoadBalancedStrategy distributes requests evenly across providers using in‑memory usage counters. Dynamic model handling ModelHandler loads model definitions at runtime and resolves the appropriate provider per request. Pluggable endpoint architecture Automatic discovery and registration of all concrete EndpointI implementations via EndpointAutoLoader . Prometheus metrics integration Optional /metric…"},{"k":"llm-router-api/index.html","t":"REST API reference","s":"REST API","x":"# llm‑router‑api **llm‑router‑api** is a lightweight Python library that provides a flexible, extensible proxy for Large Language Model ( LLM) back‑ends. It abstracts the details of multiple model providers (OpenAI‑compatible, Ollama, vLLM…","h":["Features","Installation","Running the Server","REST API Overview","Load‑Balancing Strategies","Extending the Router","Adding a New Provider Type","Adding a New Endpoint","Prompt Files","Monitoring &amp; Metrics","License"],"b":"llm‑router‑api # llm‑router‑api is a lightweight Python library that provides a flexible, extensible proxy for Large Language Model ( LLM) back‑ends. It abstracts the details of multiple model providers (OpenAI‑compatible, Ollama, vLLM, LM Studio, etc.) and offers a unified REST interface with built‑in load‑balancing, health‑checking, and monitoring. Repository: https://github.com/radlab-dev-group/llm-router Features # Unified API – One REST surface ( /api/... ) that proxies calls to any supported LLM back‑end. Provider Selection – Choose a provider per request using pluggable strategies (balanced, weighted, adaptive, first‑available). Prompt Management – System prompts are stored as files and can be dynamically injected with placeholder substitution. Streaming Support – Transparent streaming for both OpenAI‑compatible and Ollama endpoints. Health Checks – Built‑in ping endpoint and Redis‑based provider health monitoring. Prometheus Metrics – Optional instrumentation for request counts, latencies, and error rates. Auto‑Discovery – Endpoints are automatically discovered and instantiated at startup. Extensible – Add new providers, strategies, or custom endpoints with minimal boilerplate. Installation # The project uses Python 3.10.6 and a virtualenv ‑based workflow. ```shell script Clone the repository # git clone https://github.com/radlab-dev-group/llm-router.git cd llm-router Create a virtual environment # python3 -m venv venv source venv/bin/activate Install the package (including optional extras) # pip install -e .[metrics] # installs Prometheus support All required third‑party libraries are listed in `requirements.txt` (e.g., Flask, requests, redis, rdl‑ml‑utils, etc.). --- ## Configuration Configuration is driven primarily by environment variables and a JSON model‑config file. ### Environment Variables | Variable | Description | Default | |---------------------------------------------------|--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|----------------------------------------| | `LLM_ROUTER_PROMPTS_DIR` | Directory containing predefined system prompts. | `resources/prompts` | | `LLM_ROUTER_MODELS_CONFIG` | Path to the models configuration JSON file. | `resources/configs/models-config.json` | | `LLM_ROUTER_DEFAULT_EP_LANGUAGE` | Default language for endpoint prompts. | `pl` | | `LLM_ROUTER_TIMEOUT` | Timeout (seconds) for llm-router API calls. | `0` | | `LLM_ROUTER_EXTERNAL_TIMEOUT` | Timeout (seconds) for external model API calls. | `300` | | `LLM_ROUTER_LOG_FILENAME` | Name of the log file. | `llm-router.log` | | `LLM_ROUTER_LOG_LEVEL` | Logging level (e.g., INFO, DEBUG). | `INFO` | | `LLM_ROUTER_EP_PREFIX` | Prefix for all API endpoints. | `/api` | | `LLM_ROUTER_MINIMUM` | Run service in proxy‑only mode (boolean). | `False` | | `LLM_ROUTER_IN_DEBUG` | Run server in debug mode (boolean). | `False` | | `LLM_ROUTER_BALANCE_STRATEGY` | Strategy used to balance routing between LLM providers. Allowed values are `balanced`, `weighted`, `dynamic_weighted` (beta), `first_available` and `first_available_optim` as defined in `constants_base.py`. | `balanced` | | `LLM_ROUTER_REDIS_HOST` | Redis host for load‑balancing when a multi‑provider model is available. | `&lt;empty string&gt;` | | `LLM_ROUTER_REDIS_PORT` | Redis port for load‑balancing when a multi‑provider model is available. | `6379` | | `LLM_ROUTER_SERVER_TYPE` | Server implementation to use (`flask`, `gunicorn`, `waitress`). | `flask` | | `LLM_ROUTER_SERVER_PORT` | Port on which the server listens. | `8080` | | `LLM_ROUTER_SERVER_HOST` | Host address for the server. | `0.0.0.0` | | `LLM_ROUTER_SERVER_WORKERS_COUNT` | Number of workers (used in case when the selected server type supports multiworkers) | `2` | | `LLM_ROUTER_SERVER_THREADS_COUNT` | Number o…"},{"k":"llm-router-api/endpoints/index.html","t":"Endpoints Overview","s":"REST API","x":"## Endpoints Overview All endpoints are exposed under the REST API service. Unless stated otherwise, methods are POST and consume/produce JSON. ### Health & Info - **GET** `LLM_ROUTER_EP_PREFIX/ping` – Simple health‑check, returns `\"pong\"`…","h":["Endpoints Overview","Health &amp; Info","Provider‑Specific","Chat &amp; Completions (Built‑in)","Utility Endpoints (Built‑in)","Streaming vs. Non‑Streaming Responses","Payload format"],"b":"Endpoints Overview # All endpoints are exposed under the REST API service. Unless stated otherwise, methods are POST and consume/produce JSON. Health &amp; Info # GET LLM_ROUTER_EP_PREFIX/ping – Simple health‑check, returns \"pong\" . GET LLM_ROUTER_EP_PREFIX/ – Ollama health endpoint. Provider‑Specific # GET LLM_ROUTER_EP_PREFIX/tags – List available Ollama model tags. GET LLM_ROUTER_EP_PREFIX/models – List OpenAI‑compatible models. POST LLM_ROUTER_EP_PREFIX/api/v0/models – List LM Studio models. POST LLM_ROUTER_EP_PREFIX/api/chat – Ollama‑style chat completion. POST LLM_ROUTER_EP_PREFIX/api/chat/completions – OpenAI‑style chat completion. POST LLM_ROUTER_EP_PREFIX/chat/completions – OpenAI‑style chat completion (alternative path). POST LLM_ROUTER_EP_PREFIX/v1/chat/completions – vLLM‑like chat completion. Chat &amp; Completions (Built‑in) # POST LLM_ROUTER_EP_PREFIX/api/conversation_with_model – Standard chat endpoint (OpenAI‑compatible payload). POST LLM_ROUTER_EP_PREFIX/api/extended_conversation_with_model – Chat with extended fields support. POST LLM_ROUTER_EP_PREFIX/api/generative_answer – Answer a question using provided context. Utility Endpoints (Built‑in) # POST LLM_ROUTER_EP_PREFIX/api/generate_questions – Generate questions from input texts. POST LLM_ROUTER_EP_PREFIX/api/translate – Translate a list of texts. POST LLM_ROUTER_EP_PREFIX/api/simplify_text – Simplify input texts. POST LLM_ROUTER_EP_PREFIX/api/generate_article_from_text – Generate a short article from a single text. POST LLM_ROUTER_EP_PREFIX/api/create_full_article_from_texts – Generate a full article from multiple texts. Streaming vs. Non‑Streaming Responses # Streaming ( stream: true – default) The proxy opens an HTTP chunked connection and forwards each token/segment from the upstream LLM as soon as it arrives. Clients can process partial output in real time (e.g., live UI updates). Non‑Streaming ( stream: false ) The proxy collects the full response from the provider, then returns a single JSON object containing the complete text. Use this mode when you need the whole answer before proceeding. Both modes are supported for every provider that implements the streaming interface (OpenAI, Ollama, vLLM). The stream flag lives in the request schema ( OpenAIChatModel and analogous models) and is honoured automatically by the proxy. Payload format # Payload format follows the OpenAI schema ( model , messages , optional stream , etc.) unless a custom endpoint overrides it. All endpoints automatically: Validate required arguments (via REQUIRED_ARGS ). Resolve the appropriate provider using the configured load‑balancing strategy . Inject system prompts when SYSTEM_PROMPT_NAME is defined. Return a JSON response with { \"status\": true, \"body\": … } or an error payload."},{"k":"llm-router-api/endpoints/readme-pl.html","t":"Jak tworzyć endpointy w llm-proxy-api – przewodnik","s":"REST API","x":"## Jak tworzyć endpointy w llm-proxy-api – przewodnik Poniżej zebrano kluczowe informacje o tym, jak definiować i konfigurować endpointy (EP) na podstawie klas z `endpoints.*`, z odniesieniem do logiki wykonywania w `endpoint_i.EndpointWit…","h":["Jak tworzyć endpointy w llm-proxy-api – przewodnik","Propozycja EP: BatchFileSummaries – podsumowania plików z listy"],"b":"Jak tworzyć endpointy w llm-proxy-api – przewodnik # Poniżej zebrano kluczowe informacje o tym, jak definiować i konfigurować endpointy (EP) na podstawie klas z endpoints.* , z odniesieniem do logiki wykonywania w endpoint_i.EndpointWithHttpRequestI.run_ep(...) . Uwzględniono też role atrybutów/stałych takich jak self._map_prompt , self._prompt_str_postfix , _prepare_response_function , _prompt_str_force , SYSTEM_PROMPT_NAME , REQUIRED_ARGS , OPTIONAL_ARGS oraz parametry konstruktora. 1) Hierarchia i warianty bazowe - EndpointI: baza dla EP (gdy serwis nie działa jako proxy). Definiuje ogólne API i walidację argumentów, ale nie implementuje run_ep. - EndpointWithHttpRequestI: rozszerza EndpointI o wysyłkę żądań HTTP do zewnętrznego LLM. Ma pełną implementację run_ep, obsługę streamingu i wstrzykiwania promptu systemowego. - PassthroughI: dziedziczy z EndpointWithHttpRequestI i domyślnie “przepuszcza” payload (prepare_payload zwraca parametry bez zmian). Użyteczne dla OpenAI‑kompatybilnych EP, gdzie chcemy prosto forwardować żądania. Dlaczego w openai.py dziedziczymy z PassthroughI? - Bo endpointy OpenAI‑kompatybilne często wymagają minimalnej logiki – wystarczy przekazać dalej to, co przyszło. PassthroughI upraszcza implementację (brak wymuszonych argumentów, brak system promptu, gotowy run_ep proxy). 2) Cykl wykonania – co robi run_ep w EndpointWithHttpRequestI W dużym skrócie: - Inicjalizacja zegara i wyzerowanie atrybutów promptu: _map_prompt , _prompt_str_force , _prompt_str_postfix . - Wywołanie prepare_payload(params): tu podklasa ma przekształcić wejście do formatu, jaki rozumie backend (np. ułożyć messages, przepisać model_name → model, ustawić stream itp.). Jeśli zwróci strukturę z \"status\": False , run_ep zwróci ją bez dalszego przetwarzania. - Jeśli ustawiono direct_return=True, zwracany jest wynik prepare_payload bez proxy. - Tryb “simple proxy”: jeżeli klasa nie definiuje REQUIRED_ARGS (pusta lista) – traktujemy EP jako bezpośredni proxy do odpowiednika po stronie modelu. Wtedy: - _set_model wybiera model na podstawie pól z MODEL_NAME_PARAMS. - Jeżeli typ API modelu jest zgodny z typami EP ( api_types ), payload jest przekazywany dalej do odpowiedniego URL (z opcjonalnym stream). - Jeżeli to nie simple proxy: - _resolve_prompt_name(...) przygotowuje system prompt (opisane w pkt 3). - __dispatch_external_api_model(params) ustawia _api_model na podstawie nazwy modelu. - Wyznaczamy docelowy URL przez ApiTypesDispatcher (np. chat_ep dla danego api_type ). - Obsługa stream=False/True (w tym wariancie streaming może być ograniczony – komunikat o braku wsparcia). - _call_http_request(...) wykonuje POST/GET do hosta modelu, składając finalny payload (w tym system message, jeśli jest). Dodatkowe ścieżki: - call_for_each_user_msg=True : dla zadań wielotekstowych – wysyłamy osobne żądanie dla każdej wiadomości użytkownika, a wynik agregujemy przez _prepare_response_function . 3) System prompt i modyfikacje treści – jak działają pola - SYSTEM_PROMPT_NAME: słownik { \"pl\": prompt_id, \"en\": prompt_id }. W prepare_payload ustawiasz wymagania EP, a run_ep w _resolve_prompt_name : - wybiera język z parametru LANGUAGE_PARAM (z defaultem DEFAULT_EP_LANGUAGE), - pobiera treść promptu systemowego przez PromptHandler jeśli zdefiniowano nazwę, - stosuje _map_prompt – słownik zamian {placeholder: tekst}, np. wstrzyknięcie liczby pytań, treści zapytania użytkownika, - dokleja _prompt_str_postfix na końcu system promptu (np. dodatkowa instrukcja), - jeśli _prompt_str_force jest ustawione – nadpisuje całą treść system promptu (pomija nazwę/system prompt z plików). Efekt: jeśli _prompt_str ostatecznie jest zbudowany, to zostaje dodany do messages jako pierwszy element: {\"role\": \"system\", \"content\": self._prompt_str}. Kiedy to ustawiać? - W prepare_payload: - self._map_prompt: gdy chcesz w promptach z zasobów podmienić znaczniki (np. ##QUESTION_NUM_STR##). - self._prompt_str_postfix: gdy EP potrzebuje dokleić końcową uwagę/regułę do system pr…"},{"k":"llm-router-api/lb-strategies.html","t":"Load Balancing Strategies","s":"REST API","x":"## Load Balancing Strategies The `llm-router` supports various strategies for selecting the most suitable provider when multiple options exist for a given model. This ensures efficient and reliable routing of requests. The available strate…","h":["Load Balancing Strategies","1. balanced (Default)","2. weighted","3. dynamic_weighted (beta)","4. first_available","4. first_available_optim","Extending with Custom Strategies"],"b":"Load Balancing Strategies # The llm-router supports various strategies for selecting the most suitable provider when multiple options exist for a given model. This ensures efficient and reliable routing of requests. The available strategies are: 1. balanced (Default) # Description: This is the default strategy. It aims to distribute requests evenly across available providers by keeping track of how many times each provider has been used for a specific model. It selects the provider that has been used the least. When to use: Ideal for scenarios where all providers are considered equal in terms of capacity and performance. It provides a simple and effective way to balance the load. Implementation: Implemented in llm_router_api.base.lb.balanced.LoadBalancedStrategy . 2. weighted # Description: This strategy allows you to assign static weights to providers. Providers with higher weights are more likely to be selected. The selection is deterministic, ensuring that over time, the request distribution closely matches the configured weights. When to use: Useful when you have providers with different capacities or performance characteristics, and you want to prioritize certain providers without needing dynamic adjustments. Implementation: Implemented in llm_router_api.base.lb.weighted.WeightedStrategy . 3. dynamic_weighted (beta) # Description: An extension of the weighted strategy. It not only uses weights but also tracks the latency between successive selections of the same provider. This allows for more adaptive routing, as providers with consistently high latency might be de-prioritized over time. You can also dynamically update provider weights. When to use: Recommended for dynamic environments where provider performance can fluctuate. It offers more sophisticated load balancing by considering both configured weights and real-time performance metrics (latency). Implementation: Implemented in llm_router_api.base.lb.weighted.DynamicWeightedStrategy . 4. first_available # Description: This strategy selects the very first provider that is available. It uses Redis to coordinate across multiple workers, ensuring that only one worker can use a specific provider at a time. When to use: Suitable for critical applications where you need the fastest possible response and want to ensure that a request is immediately handled by any available provider, without complex load distribution logic. It guarantees that a provider, once taken, is exclusive until released. Implementation: Implemented in llm_router_api.base.lb.first_available.FirstAvailableStrategy . When using the first_available load balancing strategy, a Redis server is required for coordinating provider availability across multiple workers. 4. first_available_optim # UNDER DEVELOPMENT, DESCRIPTION WILL BE SOON The connection details for Redis can be configured using environment variables: LLM_ROUTER_BALANCE_STRATEGY = \"first_available\" \\ LLM_ROUTER_REDIS_HOST = \"your.machine.redis.host\" \\ LLM_ROUTER_REDIS_PORT = redis_port \\ Installing Redis on Ubuntu To install Redis on an Ubuntu system, follow these steps: Update package list: sudo apt update Install Redis server: sudo apt install redis-server Start and enable Redis service: The Redis service should start automatically after installation. To ensure it's running and starts on system boot, you can use the following commands: sudo systemctl status redis-server sudo systemctl enable redis-server Configure Redis (optional): The default Redis configuration ( /etc/redis/redis.conf ) is usually sufficient to get started. If you need to adjust settings (e.g., address, port), edit this file. After making configuration changes, restart the Redis server: sudo systemctl restart redis-server Extending with Custom Strategies # To use a different strategy (e.g., round‑robin, random weighted, latency‑based), implement ChooseProviderStrategyI and pass the instance to ProviderChooser : from llm_router_api.base.lb.chooser import ProviderChooser from m…"},{"k":"llm-router-lib/index.html","t":"llmrouterlib","s":"Python library","x":"# llm_router_lib ## Overview `llm_router_lib` is ** a collection of data‑model definitions**. It supplies the **foundation** for request/response structures used by the `llm_router_api` package **and** provides a **thin, opinionated client…","h":["Overview","Installation","Quick start","Data models","Conversation models","Thin client wrapper (LLMRouterClient)"],"b":"llm_router_lib # Overview # llm_router_lib is ** a collection of data‑model definitions . It supplies the foundation for request/response structures used by the llm_router_api package and provides a thin, opinionated client wrapper** that makes interacting with the LLM Router service straightforward. Key components: Package Purpose data_models pydantic models that define the shape of payloads sent to the router (e.g. GenerativeConversationModel , ExtendedGenerativeConversationModel , utility models for question generation, translation, etc.). These models are shared with the API side, ensuring both client and server speak the same contract. client.py LLMRouterClient – a lightweight wrapper around the router’s HTTP API. It offers high‑level methods ( conversation_with_model , extended_conversation_with_model ) that accept either plain dictionaries or the aforementioned data‑model instances. The client handles payload validation, provider selection, error mapping, and response parsing. services Low‑level service classes ( ConversationService , ExtendedConversationService ) that perform the actual HTTP calls via HttpRequester . They are used internally by the client but can be reused directly if finer‑grained control is needed. exceptions.py Custom exception hierarchy ( LLMRouterError , AuthenticationError , RateLimitError , ValidationError ) that mirrors the router’s error semantics, making error handling in user code clean and explicit. utils/http.py HttpRequester – a small wrapper around requests providing retries, time‑outs and logging. It is the networking backbone for the client wrapper. In short, llm_router_lib provides both the data contract (the “schema”) and a convenient Pythonic client to consume the router service. Installation # The library targets Python 3.10.6 and uses a virtualenv . Install it in editable mode for development: # Clone the repository (if you haven't already) git clone https://github.com/radlab-dev-group/llm-router.git cd llm-router/llm_router_lib # Create and activate a virtual environment python3 -m venv .venv source .venv/bin/activate # Install the package and its dependencies pip install -e . All runtime dependencies ( requests , pydantic , rdl_ml_utils ) are declared in the project’s requirements.txt . Quick start # from llm_router_lib import LLMRouterClient # Initialise the client – point it at the router’s base URL client = LLMRouterClient ( api = \"http://localhost:8080/api\" , # router base URL token = \"YOUR_ROUTER_TOKEN\" , # optional, if router requires auth ) # Build a payload using the provided data model (validation is automatic) payload = { \"model_name\" : \"google/gemma-3-12b-it\" , \"user_last_statement\" : \"Hello, how are you?\" , \"temperature\" : 0.7 , \"max_new_tokens\" : 128 , } # Call the standard conversation endpoint response = client . conversation_with_model ( payload ) print ( response ) # → {'status': True, 'body': {...}} You can also pass a pydantic model instance directly: python from llm_router_lib.data_models.builtin_chat import GenerativeConversationModel model = GenerativeConversationModel( model_name=\"google/gemma-3-12b-it\", user_last_statement=\"Hello, how are you?\", temperature=0.7, max_new_tokens=128, ) response = client.conversation_with_model(model) Data models # All request payloads are defined in llm_router_lib/data_models . Common base: class BaseModelOptions ( BaseModel ): \"\"\"Options shared across many endpoint models.\"\"\" mask_payload : bool = False masker_pipeline : Optional [ List [ str ]] = None Conversation models # Model Required fields Optional / extra fields GenerativeConversationModel model_name , user_last_statement temperature , max_new_tokens , historical_messages , … ExtendedGenerativeConversationModel All of the above + system_prompt – Utility models for other built‑in endpoints (question generation, translation, article creation, context‑based answering, etc.) follow the same pattern and inherit from BaseModelOptions . Thin client wrapper ( LLMRouterClie…"},{"k":"examples/index.html","t":"Integration Examples with LLM Router","s":"Examples","x":"# Integration Examples with LLM Router This directory contains example boilerplates that demonstrate how easy it is to integrate popular LLM libraries with the router by simply switching the host. ## Available Examples - **[LlamaIndex](lla…","h":["Available Examples","Core Principle","Quick Start","Example Structure","Additional Information"],"b":"Integration Examples with LLM Router # This directory contains example boilerplates that demonstrate how easy it is to integrate popular LLM libraries with the router by simply switching the host. Available Examples # LlamaIndex – Integration with LlamaIndex (GPT Index) Using LlamaIndex with Local Models – Guide for mapping OpenAI model names to local models via the router. LangChain – Integration with LangChain OpenAI SDK – Direct integration with the OpenAI Python SDK LiteLLM – Integration with LiteLLM Haystack – Integration with Haystack Core Principle # All examples work on the same principle: just change base_url / api_base to the address of your router , and the router will automatically: ✅ Distribute traffic among available providers ✅ Perform load balancing ✅ Provide health checking ✅ Supply monitoring and metrics ✅ Handle streaming and non‑streaming responses Quick Start # Each example can be run directly: ```shell script LlamaIndex # python examples/llamaindex_example.py LangChain # python examples/langchain_example.py OpenAI SDK # python examples/openai_example.py LiteLLM # python examples/litellm_example.py Haystack # python examples/haystack_example.py ``` Example Structure # Each example includes: Basic configuration – how to point the library at the router Streaming – handling streaming responses Non‑streaming – handling full responses Error handling – managing errors Additional Information # Learn more about the router: Main README API Documentation Endpoints Overview Load‑Balancing Strategies"},{"k":"examples/readme-llamaindex.html","t":"Using LlamaIndex with Local Models via the LLM‑Router","s":"Examples","x":"## Using LlamaIndex with Local Models via the LLM‑Router LlamaIndex’s `OpenAI` wrapper expects **OpenAI‑style model names** (e.g. `gpt-3.5-turbo`, `gpt-4`). When you want to run *local* models (Gemma, Ollama, vLLM, etc.) behind an LLM‑Rout…","h":["Using LlamaIndex with Local Models via the LLM‑Router","Why the mapping is required","Router configuration","How it works end‑to‑end","Checklist for a working setup","TL;DR"],"b":"Using LlamaIndex with Local Models via the LLM‑Router # LlamaIndex’s OpenAI wrapper expects OpenAI‑style model names (e.g. gpt-3.5-turbo , gpt-4 ). When you want to run local models (Gemma, Ollama, vLLM, etc.) behind an LLM‑Router, the router must translate those OpenAI names to the actual model identifiers used by the local providers. Why the mapping is required # LlamaIndex validates the model name against a known list of OpenAI models to decide whether the call is a chat request, to obtain context‑window size, etc. If the wrapper receives a name it does not recognise (e.g. google/gemma-3-12b-it ), it raises a ValueError like: ValueError: Unknown model 'google/gemma-3-12b-it'. Please provide a valid OpenAI model name … By keeping the OpenAI name in the request and letting the router forward the request to the appropriate backend, you get the best of both worlds: LlamaIndex works unchanged, and the traffic is routed to your local model. Router configuration # The router’s configuration must contain a section that lists OpenAI model names ( openai_models ). Each entry defines one or more providers that actually serve the model. The key points are: Field Meaning id Arbitrary identifier for the provider instance. api_host URL where the provider’s inference server is reachable. api_type The protocol the provider uses ( vllm , ollama , …). model_path The local model identifier (e.g. google/gemma-3-12b-it ). weight Relative load‑balancing weight when several providers are listed. Example snippet # \"openai_models\" : { (...) \"gpt-3.5-turbo\" : { \"providers\" : [ { \"id\" : \"gpt_35_turbo-gemma3_12b-vllm-71:7000\" , \"api_host\" : \"http://192.168.100.71:7000/\" , \"api_token\" : \"\" , \"api_type\" : \"vllm\" , \"input_size\" : 4096 , \"model_path\" : \"google/gemma-3-12b-it\" , \"weight\" : 1.0 }, { \"id\" : \"gpt_35_turbo-gemma3_12b-vllm-71:7001\" , \"api_host\" : \"http://192.168.100.71:7001/\" , \"api_token\" : \"\" , \"api_type\" : \"vllm\" , \"input_size\" : 4096 , \"model_path\" : \"google/gemma-3-12b-it\" , \"weight\" : 1.0 } ] }, \"gpt-4\" : { \"providers\" : [ { \"id\" : \"gpt-4-gpt-oss-20b-ollama-66:11434\" , \"api_host\" : \"http://192.168.100.66:11434\" , \"api_token\" : \"\" , \"api_type\" : \"ollama\" , \"input_size\" : 256000 , \"model_path\" : \"gpt-oss:120b\" } ] } }, ... \"active_models\" : { \"openai_models\" : [ \"gpt-3.5-turbo\" , \"gpt-4\" ... ] } How it works end‑to‑end # LlamaIndex code creates an OpenAI client with a model name that the wrapper knows, e.g.: llm = OpenAI ( model = \"gpt-3.5-turbo\" , api_base = \"http://localhost:8080\" , api_key = \"not-needed\" ) The request is sent to the router ( api_base ). The router looks up \"gpt-3.5-turbo\" in openai_models , selects a provider, and forwards the request to the provider’s api_host . The provider receives the request with model_path=\"google/gemma-3-12b-it\" (or \"gpt-oss:120b\" for gpt-4 ) and runs the local model. The response travels back through the router to LlamaIndex, which treats it exactly like an OpenAI response. Checklist for a working setup # Router is running and reachable at the URL you pass to api_base . openai_models section contains all OpenAI names you intend to use from LlamaIndex. Each OpenAI name maps to at least one provider with the correct model_path . The active_models.openai_models list includes the names you want to expose (otherwise the router will ignore them). Local model servers (vLLM, Ollama, etc.) are up and listening on the api_host URLs specified. TL;DR # *Keep the model names you give to LlamaIndex identical to the OpenAI names defined in the router configuration. The router will translate those names to the real local model identifiers ( model_path ). This mapping is the only thing needed for LlamaIndex to work seamlessly with any self�"},{"k":"changelog.html","t":"Changelog","s":"Release notes","x":"## Changelog | Version | Changelog | |---------|-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------…","h":["Changelog"],"b":"Changelog # Version Changelog 0.0.1 Initialization, License, setup, interface for each endpoint and sample ping EP. Autoloader of builtin endpoints and for the future implementations. 0.0.2 Add base models for api call (module llm_proxy_rest.data_models with error.py handling. Decorators to check required params and to measure the response time. 0.0.3 Proper AutoLoading for each found endpoint. Implementation of ApiTypesDispatcher , ApiModelConfig , ModelHandler . Ollama endpoints: / , tags . Added endpoint to full proxy with params. Streaming in case when external api provides stream. 0.0.4 All llama-service endpoints are refactored to llm-proxy-api . Refactoring base ep_run method. Proper handling system message, prompt name, model etc. 0.1.0 Repository name changed from llm-proxy-api to llm-router . Added class HttpRequestExecutor to handle http requests from EndpointWithHttpRequestI . Handled routing between any models: openai -&gt; ollama and ollama -&gt; openai 0.1.1 Prometheus metrics logging. Workers/Threads/Workers class is able to set by environments. Streaming fixes. Multi-providers for single model with default-balanced strategy. 0.2.0 Add balancing strategies: balanced , weighted , dynamic_weighted and first_available which works for streaming and non streaming requests. Included Prometheus metrics logging via /metrics endpoint. First stage of llm_router_lib library, to simply usage of llm-router-api . 0.2.1 Fix stream: OpenAI-&gt;Ollama, Ollama-&gt;OpenAI. Add Redis caching of availability of model providers (when using first_available strategy). Add llm_router_web module with simple flask-based frontend to manage llm-router config files. 0.2.2 Update dockerfile and requirements. Fix routing with vLLM. 0.2.3 New web configurator: Handling projects, configs for each user separately. First Available strategy is more powerful, a lot of improvements to efficiency. 0.2.4 Anonymizer module, integration anonymization with any endpoint (using dynamic payload analysis and full payload anonymisation), dedicated /api/anonymize_text endpoint as memory only anonymization. Whole router may be run in FORCE_ANONYMISATION mode. 0.3.0 Anonymization available with three strategies: fast_masker , genai , prov_masker . 0.3.1 Refactoring lb.strategies to be more flexible modular. Introduced MaskerPipeline and GuardrailPipeline both configured via env. Removed genai-based masking endpoint. 0.4.0 The main repository is divided into dedicated ones: plugins, services, web — separate repositories. Clean up the whole repository. Examples of integration with llamaindex, langchain, openai, litellm and haystack."}]}