{"version":"0.3.1","pages":[{"k":"overview.html","t":"Overview","s":"Getting started","x":"## LLM Router - Open-Source AI Gateway for Local and Cloud LLM Infrastructure [**LLM Router**](https://llm-router.cloud) is a service that can be deployed on‑premises or in the cloud. It adds a layer between any application and the LLM pro…","h":["LLM Router - Open-Source AI Gateway for Local and Cloud LLM Infrastructure","✨ Key Features","📦 Quick Start","1️⃣ Create &amp; activate a virtual environment","2️⃣ Minimum required environment variable","3️⃣ Optional configuration (via environment)","4️⃣ Run the REST API","⚖️ Load Balancing Strategies","🛣️ Endpoints Overview","Health &amp; Info","Provider‑Specific","Chat &amp; Completions (Built‑in)","Utility Endpoints (Built‑in)","Streaming vs. Non‑Streaming Responses","⚙️ Configuration Details","🛠️ Development","📜 License","📚 Changelog"],"b":"LLM Router - Open-Source AI Gateway for Local and Cloud LLM Infrastructure # LLM Router is a service that can be deployed on‑premises or in the cloud. It adds a layer between any application and the LLM provider. In real time it controls traffic, distributes a load among providers of a specific LLM, and enables analysis of outgoing requests from a security perspective (masking, anonymization, prohibited content). It is an open‑source solution (Apache 2.0) that can be launched instantly by running a ready‑made image in your own infrastructure. llm_router_api provides a unified REST proxy that can route requests to any supported LLM backend ( OpenAI‑compatible, Ollama, vLLM, LM Studio, etc.), with built‑in load‑balancing, health checks, streaming responses and optional Prometheus metrics. llm_router_lib is a Python SDK that wraps the API with typed request/response models, automatic retries, token handling and a rich exception hierarchy, letting developers focus on application logic rather than raw HTTP calls. llm_router_web offers ready‑to‑use Flask UIs – an anonymizer UI that masks sensitive data and a configuration manager for model/user settings – demonstrating how to consume the router from a browser. llm_router_plugins (e.g., the fast_masker plugin) deliver a rule‑based text anonymisation engine with a comprehensive set of Polish‑specific masking rules (emails, IPs, URLs, phone numbers, PESEL, NIP, KRS, REGON, monetary amounts, dates, etc.) and an extensible architecture for custom rules and validators. All components run on Python 3.10+ using virtualenv and require only the listed dependencies, making the suite easy to install, extend, and deploy in both development and production environments. ✨ Key Features # Feature Description Unified REST interface One endpoint schema works for OpenAI‑compatible, Ollama, vLLM and any future provider. Provider‑agnostic streaming The stream flag (default true ) controls whether the proxy forwards chunked responses as they arrive or returns a single aggregated payload. Built‑in prompt library Language‑aware system prompts stored under resources/prompts can be referenced automatically. Dynamic model configuration JSON file ( models-config.json ) defines providers, model name, default options and per‑model overrides. Request validation Pydantic models guarantee correct payloads; errors are returned with clear messages. Structured logging Configurable log level, filename, and optional JSON formatting. Health &amp; metadata endpoints /ping (simple 200 OK) and /tags (available model tags/metadata). Simple deployment One‑liner run script or python -m llm_proxy_rest.rest_api . Extensible conversation formats Basic chat, conversation with system prompt, and extended conversation with richer options (e.g., temperature, top‑k, custom system prompt). Multi‑provider model support Each model can be backed by multiple providers (VLLM, Ollama, OpenAI) defined in models-config.json . Provider selection abstraction ProviderChooser delegates to a configurable strategy, enabling easy swapping of load‑balancing, round‑robin, weighted‑random, etc. Load‑balanced default strategy LoadBalancedStrategy distributes requests evenly across providers using in‑memory usage counters. Dynamic model handling ModelHandler loads model definitions at runtime and resolves the appropriate provider per request. Pluggable endpoint architecture Automatic discovery and registration of all concrete EndpointI implementations via EndpointAutoLoader . Prometheus metrics integration Optional /metrics endpoint for latency, error counts, and provider usage statistics. Docker ready Dockerfile and scripts for containerised deployment. 📦 Quick Start # 1️⃣ Create &amp; activate a virtual environment # Base requirements # Prerequisite : radlab-ml-utils This project uses the radlab-ml-utils library for machine learning utilities (e.g., experiment/result logging with Weights &amp; Biases/wandb). Install it before working with ML-related par…"},{"k":"install.html","t":"Instalacja zależności","s":"Getting started","x":"# Instalacja zależności ## ml-utils ```bash git clone https://github.com/radlab-dev-group/ml-utils ``` ## llm-proxy-api ```shell pip install . --break ```","h":["ml-utils","llm-proxy-api"],"b":"Instalacja zależności # ml-utils # git clone https://github.com/radlab-dev-group/ml-utils llm-proxy-api # pip install . --break"},{"k":"llm-router-api/index.html","t":"REST API reference","s":"REST API","x":"# llm‑router‑api **llm‑router‑api** is a lightweight Python library that provides a flexible, extensible proxy for Large Language Model ( LLM) back‑ends. It abstracts the details of multiple model providers (OpenAI‑compatible, Ollama, vLLM…","h":["Features","Installation","Running the Server","REST API Overview","Load‑Balancing Strategies","Extending the Router","Adding a New Provider Type","Adding a New Endpoint","Prompt Files","Monitoring &amp; Metrics","License"],"b":"llm‑router‑api # llm‑router‑api is a lightweight Python library that provides a flexible, extensible proxy for Large Language Model ( LLM) back‑ends. It abstracts the details of multiple model providers (OpenAI‑compatible, Ollama, vLLM, LM Studio, etc.) and offers a unified REST interface with built‑in load‑balancing, health‑checking, and monitoring. Repository: https://github.com/radlab-dev-group/llm-router Features # Unified API – One REST surface ( /api/... ) that proxies calls to any supported LLM back‑end. Provider Selection – Choose a provider per request using pluggable strategies (balanced, weighted, adaptive, first‑available). Prompt Management – System prompts are stored as files and can be dynamically injected with placeholder substitution. Streaming Support – Transparent streaming for both OpenAI‑compatible and Ollama endpoints. Health Checks – Built‑in ping endpoint and Redis‑based provider health monitoring. Prometheus Metrics – Optional instrumentation for request counts, latencies, and error rates. Auto‑Discovery – Endpoints are automatically discovered and instantiated at startup. Extensible – Add new providers, strategies, or custom endpoints with minimal boilerplate. Installation # The project uses Python 3.10.6 and a virtualenv ‑based workflow. ```shell script Clone the repository # git clone https://github.com/radlab-dev-group/llm-router.git cd llm-router Create a virtual environment # python3 -m venv venv source venv/bin/activate Install the package (including optional extras) # pip install -e .[metrics] # installs Prometheus support All required third‑party libraries are listed in `requirements.txt` (e.g., Flask, requests, redis, rdl‑ml‑utils, etc.). --- ## Configuration Configuration is driven primarily by environment variables and a JSON model‑config file. ### Environment Variables | Variable | Description | Default | |---------------------------------------------------|--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|----------------------------------------| | `LLM_ROUTER_PROMPTS_DIR` | Directory containing predefined system prompts. | `resources/prompts` | | `LLM_ROUTER_MODELS_CONFIG` | Path to the models configuration JSON file. | `resources/configs/models-config.json` | | `LLM_ROUTER_DEFAULT_EP_LANGUAGE` | Default language for endpoint prompts. | `pl` | | `LLM_ROUTER_TIMEOUT` | Timeout (seconds) for llm-router API calls. | `0` | | `LLM_ROUTER_EXTERNAL_TIMEOUT` | Timeout (seconds) for external model API calls. | `300` | | `LLM_ROUTER_LOG_FILENAME` | Name of the log file. | `llm-router.log` | | `LLM_ROUTER_LOG_LEVEL` | Logging level (e.g., INFO, DEBUG). | `INFO` | | `LLM_ROUTER_EP_PREFIX` | Prefix for all API endpoints. | `/api` | | `LLM_ROUTER_MINIMUM` | Run service in proxy‑only mode (boolean). | `False` | | `LLM_ROUTER_IN_DEBUG` | Run server in debug mode (boolean). | `False` | | `LLM_ROUTER_BALANCE_STRATEGY` | Strategy used to balance routing between LLM providers. Allowed values are `balanced`, `weighted`, `dynamic_weighted` (beta), `first_available` and `first_available_optim` as defined in `constants_base.py`. | `balanced` | | `LLM_ROUTER_REDIS_HOST` | Redis host for load‑balancing when a multi‑provider model is available. | `&lt;empty string&gt;` | | `LLM_ROUTER_REDIS_PORT` | Redis port for load‑balancing when a multi‑provider model is available. | `6379` | | `LLM_ROUTER_SERVER_TYPE` | Server implementation to use (`flask`, `gunicorn`, `waitress`). | `flask` | | `LLM_ROUTER_SERVER_PORT` | Port on which the server listens. | `8080` | | `LLM_ROUTER_SERVER_HOST` | Host address for the server. | `0.0.0.0` | | `LLM_ROUTER_SERVER_WORKERS_COUNT` | Number of workers (used in case when the selected server type supports multiworkers) | `2` | | `LLM_ROUTER_SERVER_THREADS_COUNT` | Number o…"},{"k":"llm-router-api/endpoints/readme-pl.html","t":"Jak tworzyć endpointy w llm-proxy-api – przewodnik","s":"REST API","x":"## Jak tworzyć endpointy w llm-proxy-api – przewodnik Poniżej zebrano kluczowe informacje o tym, jak definiować i konfigurować endpointy (EP) na podstawie klas z `endpoints.*`, z odniesieniem do logiki wykonywania w `endpoint_i.EndpointWit…","h":["Jak tworzyć endpointy w llm-proxy-api – przewodnik","Propozycja EP: BatchFileSummaries – podsumowania plików z listy"],"b":"Jak tworzyć endpointy w llm-proxy-api – przewodnik # Poniżej zebrano kluczowe informacje o tym, jak definiować i konfigurować endpointy (EP) na podstawie klas z endpoints.* , z odniesieniem do logiki wykonywania w endpoint_i.EndpointWithHttpRequestI.run_ep(...) . Uwzględniono też role atrybutów/stałych takich jak self._map_prompt , self._prompt_str_postfix , _prepare_response_function , _prompt_str_force , SYSTEM_PROMPT_NAME , REQUIRED_ARGS , OPTIONAL_ARGS oraz parametry konstruktora. 1) Hierarchia i warianty bazowe - EndpointI: baza dla EP (gdy serwis nie działa jako proxy). Definiuje ogólne API i walidację argumentów, ale nie implementuje run_ep. - EndpointWithHttpRequestI: rozszerza EndpointI o wysyłkę żądań HTTP do zewnętrznego LLM. Ma pełną implementację run_ep, obsługę streamingu i wstrzykiwania promptu systemowego. - PassthroughI: dziedziczy z EndpointWithHttpRequestI i domyślnie “przepuszcza” payload (prepare_payload zwraca parametry bez zmian). Użyteczne dla OpenAI‑kompatybilnych EP, gdzie chcemy prosto forwardować żądania. Dlaczego w openai.py dziedziczymy z PassthroughI? - Bo endpointy OpenAI‑kompatybilne często wymagają minimalnej logiki – wystarczy przekazać dalej to, co przyszło. PassthroughI upraszcza implementację (brak wymuszonych argumentów, brak system promptu, gotowy run_ep proxy). 2) Cykl wykonania – co robi run_ep w EndpointWithHttpRequestI W dużym skrócie: - Inicjalizacja zegara i wyzerowanie atrybutów promptu: _map_prompt , _prompt_str_force , _prompt_str_postfix . - Wywołanie prepare_payload(params): tu podklasa ma przekształcić wejście do formatu, jaki rozumie backend (np. ułożyć messages, przepisać model_name → model, ustawić stream itp.). Jeśli zwróci strukturę z \"status\": False , run_ep zwróci ją bez dalszego przetwarzania. - Jeśli ustawiono direct_return=True, zwracany jest wynik prepare_payload bez proxy. - Tryb “simple proxy”: jeżeli klasa nie definiuje REQUIRED_ARGS (pusta lista) – traktujemy EP jako bezpośredni proxy do odpowiednika po stronie modelu. Wtedy: - _set_model wybiera model na podstawie pól z MODEL_NAME_PARAMS. - Jeżeli typ API modelu jest zgodny z typami EP ( api_types ), payload jest przekazywany dalej do odpowiedniego URL (z opcjonalnym stream). - Jeżeli to nie simple proxy: - _resolve_prompt_name(...) przygotowuje system prompt (opisane w pkt 3). - __dispatch_external_api_model(params) ustawia _api_model na podstawie nazwy modelu. - Wyznaczamy docelowy URL przez ApiTypesDispatcher (np. chat_ep dla danego api_type ). - Obsługa stream=False/True (w tym wariancie streaming może być ograniczony – komunikat o braku wsparcia). - _call_http_request(...) wykonuje POST/GET do hosta modelu, składając finalny payload (w tym system message, jeśli jest). Dodatkowe ścieżki: - call_for_each_user_msg=True : dla zadań wielotekstowych – wysyłamy osobne żądanie dla każdej wiadomości użytkownika, a wynik agregujemy przez _prepare_response_function . 3) System prompt i modyfikacje treści – jak działają pola - SYSTEM_PROMPT_NAME: słownik { \"pl\": prompt_id, \"en\": prompt_id }. W prepare_payload ustawiasz wymagania EP, a run_ep w _resolve_prompt_name : - wybiera język z parametru LANGUAGE_PARAM (z defaultem DEFAULT_EP_LANGUAGE), - pobiera treść promptu systemowego przez PromptHandler jeśli zdefiniowano nazwę, - stosuje _map_prompt – słownik zamian {placeholder: tekst}, np. wstrzyknięcie liczby pytań, treści zapytania użytkownika, - dokleja _prompt_str_postfix na końcu system promptu (np. dodatkowa instrukcja), - jeśli _prompt_str_force jest ustawione – nadpisuje całą treść system promptu (pomija nazwę/system prompt z plików). Efekt: jeśli _prompt_str ostatecznie jest zbudowany, to zostaje dodany do messages jako pierwszy element: {\"role\": \"system\", \"content\": self._prompt_str}. Kiedy to ustawiać? - W prepare_payload: - self._map_prompt: gdy chcesz w promptach z zasobów podmienić znaczniki (np. ##QUESTION_NUM_STR##). - self._prompt_str_postfix: gdy EP potrzebuje dokleić końcową uwagę/regułę do system pr…"},{"k":"llm-router-api/lb-strategies.html","t":"Load Balancing Strategies","s":"REST API","x":"## Load Balancing Strategies The `llm-router` supports various strategies for selecting the most suitable provider when multiple options exist for a given model. This ensures efficient and reliable routing of requests. The available strate…","h":["Load Balancing Strategies","1. balanced (Default)","2. weighted","3. dynamic_weighted (beta)","4. first_available","4. first_available_optim","Extending with Custom Strategies"],"b":"Load Balancing Strategies # The llm-router supports various strategies for selecting the most suitable provider when multiple options exist for a given model. This ensures efficient and reliable routing of requests. The available strategies are: 1. balanced (Default) # Description: This is the default strategy. It aims to distribute requests evenly across available providers by keeping track of how many times each provider has been used for a specific model. It selects the provider that has been used the least. When to use: Ideal for scenarios where all providers are considered equal in terms of capacity and performance. It provides a simple and effective way to balance the load. Implementation: Implemented in llm_router_api.base.lb.balanced.LoadBalancedStrategy . 2. weighted # Description: This strategy allows you to assign static weights to providers. Providers with higher weights are more likely to be selected. The selection is deterministic, ensuring that over time, the request distribution closely matches the configured weights. When to use: Useful when you have providers with different capacities or performance characteristics, and you want to prioritize certain providers without needing dynamic adjustments. Implementation: Implemented in llm_router_api.base.lb.weighted.WeightedStrategy . 3. dynamic_weighted (beta) # Description: An extension of the weighted strategy. It not only uses weights but also tracks the latency between successive selections of the same provider. This allows for more adaptive routing, as providers with consistently high latency might be de-prioritized over time. You can also dynamically update provider weights. When to use: Recommended for dynamic environments where provider performance can fluctuate. It offers more sophisticated load balancing by considering both configured weights and real-time performance metrics (latency). Implementation: Implemented in llm_router_api.base.lb.weighted.DynamicWeightedStrategy . 4. first_available # Description: This strategy selects the very first provider that is available. It uses Redis to coordinate across multiple workers, ensuring that only one worker can use a specific provider at a time. When to use: Suitable for critical applications where you need the fastest possible response and want to ensure that a request is immediately handled by any available provider, without complex load distribution logic. It guarantees that a provider, once taken, is exclusive until released. Implementation: Implemented in llm_router_api.base.lb.first_available.FirstAvailableStrategy . When using the first_available load balancing strategy, a Redis server is required for coordinating provider availability across multiple workers. 4. first_available_optim # UNDER DEVELOPMENT, DESCRIPTION WILL BE SOON The connection details for Redis can be configured using environment variables: LLM_ROUTER_BALANCE_STRATEGY = \"first_available\" \\ LLM_ROUTER_REDIS_HOST = \"your.machine.redis.host\" \\ LLM_ROUTER_REDIS_PORT = redis_port \\ Installing Redis on Ubuntu To install Redis on an Ubuntu system, follow these steps: Update package list: sudo apt update Install Redis server: sudo apt install redis-server Start and enable Redis service: The Redis service should start automatically after installation. To ensure it's running and starts on system boot, you can use the following commands: sudo systemctl status redis-server sudo systemctl enable redis-server Configure Redis (optional): The default Redis configuration ( /etc/redis/redis.conf ) is usually sufficient to get started. If you need to adjust settings (e.g., address, port), edit this file. After making configuration changes, restart the Redis server: sudo systemctl restart redis-server Extending with Custom Strategies # To use a different strategy (e.g., round‑robin, random weighted, latency‑based), implement ChooseProviderStrategyI and pass the instance to ProviderChooser : from llm_router_api.base.lb.chooser import ProviderChooser from m…"},{"k":"llm-router-lib/index.html","t":"llm‑router-LIB — Python client library","s":"Python library","x":"# llm‑router-LIB — Python client library **llm‑router** is a lightweight Python client for interacting with the LLM‑Router API. It provides typed request models, convenient service wrappers, and robust error handling so you can focus on bu…","h":["Overview","Features","Installation","Quick start","Extended conversation","Core concepts","Client","Data models","Services","Utilities","Error handling","License"],"b":"llm‑router-LIB — Python client library # llm‑router is a lightweight Python client for interacting with the LLM‑Router API. It provides typed request models, convenient service wrappers, and robust error handling so you can focus on building LLM‑driven applications rather than dealing with raw HTTP calls. Overview # llm_router_lib is the official Python SDK for the LLM‑Router project https://github.com/radlab-dev-group/llm-router . It abstracts the HTTP layer behind a small, well‑typed API: Typed payloads built with pydantic (e.g., GenerativeConversationModel ). Service objects that know the endpoint URL and the model class they expect. Automatic token handling , request retries, and exponential back‑off. Rich exception hierarchy ( LLMRouterError , AuthenticationError , RateLimitError , ValidationError ). Features # Feature Description Typed request/response models Guarantees payload correctness at runtime using Pydantic. Built‑in conversation services Simple conversation_with_model and extended_conversation_with_model calls. Retry &amp; timeout Configurable request timeout and automatic retries with exponential back‑off. Authentication Bearer‑token support; raises AuthenticationError on 401/403. Rate‑limit handling Detects HTTP 429 and raises RateLimitError . Extensible Add custom services or models by extending the base classes. Test suite Ready‑to‑run unit tests in llm_router_lib/tests . Installation # The library is pure Python and works with Python 3.10+ . ```shell script Create a virtualenv (recommended) # python -m venv .venv source .venv/bin/activate Install from the repository (editable mode) # pip install -e . If you prefer a regular installation from a wheel or source distribution, use: ```shell script pip install . Note – The project relies only on the packages listed in the repository’s requirements.txt (pydantic, requests, etc.), all of which are installed automatically by pip . Quick start # from llm_router_lib.client import LLMRouterClient from llm_router_lib.data_models.builtin_chat import GenerativeConversationModel # Initialise the client (replace with your own endpoint and token) client = LLMRouterClient ( api = \"https://api.your-llm-router.com\" , token = \"YOUR_ACCESS_TOKEN\" ) # Build a request payload payload = GenerativeConversationModel ( model_name = \"google/gemma-3-12b-it\" , user_last_statement = \"Hello, how are you?\" , historical_messages = [{ \"user\" : \"Hi\" }], temperature = 0.7 , max_new_tokens = 128 , ) # Call the API response = client . conversation_with_model ( payload ) print ( response ) # → dict with the model's answer and metadata Extended conversation # from llm_router_lib.data_models.builtin_chat import ExtendedGenerativeConversationModel payload = ExtendedGenerativeConversationModel ( model_name = \"google/gemma-3-12b-it\" , user_last_statement = \"Explain quantum entanglement.\" , system_prompt = \"Answer as a friendly professor.\" , temperature = 0.6 , max_new_tokens = 256 , ) response = client . extended_conversation_with_model ( payload ) print ( response ) Core concepts # Client # LLMRouterClient is the entry point. It handles: Base URL normalization. Optional bearer token injection. Construction of the internal HttpRequester . All public methods accept either a dict or a pydantic model ; models are automatically serialized with .model_dump() . Data models # Located in llm_router_lib/data_models/ . Key models: Model Purpose GenerativeConversationModel Simple chat payload (model name, user message, optional history). ExtendedGenerativeConversationModel Same as above, plus a system_prompt . GenerateQuestionFromTextsModel Generate questions from a list of texts. TranslateTextModel , SimplifyTextModel , … Various utility models for text transformation. OpenAIChatModel Payload for direct OpenAI‑compatible chat calls. All models inherit from a common _GenerativeOptions base that defines temperature, token limits, language, etc. Services # Implemented in llm_router_lib/services/ . Each service ext…"},{"k":"llm-router-web/web/anonymizer/index.html","t":"llmrouterweb.anonymizer","s":"Web UI (archived)","x":"# llm_router_web.anonymizer **llm_router_web: anonymizer** is a lightweight Flask web interface that provides a simple UI for text anonymization. It forwards the supplied text to an external LLM‑router anonymization service (`/api/fast_tex…","h":["Features","Installation","Production (Gunicorn)","Adding new features","License"],"b":"llm_router_web.anonymizer # llm_router_web: anonymizer is a lightweight Flask web interface that provides a simple UI for text anonymization. It forwards the supplied text to an external LLM‑router anonymization service ( /api/fast_text_mask ) and displays the masked result, with on‑the‑fly highlighting of detected tags ( {{…}} ). llm_router_web.anonymizer is the web interface component of the llm-router library. Features # Web form – Paste text and trigger anonymization with a single click. HTMX‑powered UI – Asynchronous request/response without a full page reload. Result highlighting – Detected placeholders are wrapped in a colored span for easy spotting. Spinner indicator – Visual feedback while the request is in progress. Error handling – Returns clear messages for missing input or communication problems. Configurable backend – Target anonymization service URL is supplied via LLM_ROUTER_HOST environment variable. Installation # The project follows the same conventions as the rest of the llm‑router repository and uses Python 3.10.6 with virtualenv . ```shell script Clone the repository (if not already done) # git clone https://github.com/radlab-dev-group/llm-router.git cd llm-router/llm_router_web Create and activate a virtual environment # python3 -m venv venv source venv/bin/activate Install dependencies (the base requirements already contain Flask) # pip install -r requirements.txt &gt; **Note:** The only additional packages required for the anonymizer are already listed in `requirements.txt` (`flask`, `requests`, `htmx`, `alpine.js` are loaded from CDN). ## Running the Application ### Development server ```shell script # Start the Flask development server for the anonymizer module python -m web.anonymizer The UI will be reachable at http://localhost:5000/anonymize . The root path ( / ) redirects to the form page. Production (Gunicorn) # A small launch script is included in the main repository. Example: ```shell script LLM_ROUTER_HOST=http://localhost:8000 \\ FLASK_SECRET_KEY=super-secret \\ gunicorn -w 4 -b 0.0.0.0:8082 \"web.anonymizer:create_anonymize_app()\" - `LLM_ROUTER_HOST` – base URL of the external anonymization service (default: `http://localhost:8000`). - `FLASK_SECRET_KEY` – secret key for session signing (default: `change-me-anonymizer`). Adjust the number of workers (`-w`) as needed. ## Configuration All configuration is performed via environment variables: | Variable | Description | Default | |--------------------|-------------------------------------------|-------------------------| | `FLASK_SECRET_KEY` | Secret key for Flask session signing | `change-me-anonymizer` | | `LLM_ROUTER_HOST` | URL of the external anonymization service | `http://localhost:8000` | The variables are read in `web/anonymizer/__init__.py` when `create_anonymize_app()` is called. ## Endpoints Overview | URL | Methods | Description | |------------------|---------|------------------------------------------------------------------------------------------------------------------------------------| | `/` (root) | GET | Redirects to `/anonymize/`. | | `/anonymize/` | GET | Renders the anonymization form (`anonymize.html`). | | `/anonymize/` | POST | Accepts `text` form field, forwards it to the external service, and returns the rendered result (`anonymize_result_partial.html`). | | *Error handlers* | – | Returns JSON payloads for 400, 404, and 500 errors. | ### Request flow (POST `/anonymize/`) 1. The form posts the `text` field via HTMX. 2. The server builds the target URL: `\"{LLM_ROUTER_HOST.rstrip('/')}/api/fast_text_mask\"`. 3. It sends a JSON payload `{ \"text\": \"&lt;raw text&gt;\" }` to the external service. 4. On success the response text (or the `text` field from JSON) is injected back into the page, where JavaScript highlights any `{{…}}` tags. ## Development ### Project layout web/ └─ anonymizer/ ├─ templates/ │ ├─ anonymize.html # Main form page (HTMX enabled) │ ├─ anonymize_result_partial.html # Partial used to render the result │ …"},{"k":"llm-router-web/web/configs-manager/index.html","t":"llmrouterweb.configs_manager","s":"Web UI (archived)","x":"# llm_router_web.configs_manager **llm_router_web: configs manager** is the web interface component of the **[llm-router](https://github.com/radlab-dev-group/llm-router) ** project. It provides a Flask‑based UI for managing LLM model confi…","h":["Features","Installation","Running with gunicorn (recommended for production)","Additional Flask environment variables","Configuration","Endpoints Overview","Development","Code structure","Adding new features","License"],"b":"llm_router_web.configs_manager # llm_router_web: configs manager is the web interface component of the ** llm-router ** project. It provides a Flask‑based UI for managing LLM model configurations, users, and version history. Features # User Management – Admins can create, edit, block/unblock users and assign roles ( admin / user ). Configuration CRUD – Create, import, edit, view, export, activate and delete model configurations. Model &amp; Provider Management – Add/remove models, manage multiple providers per model, reorder providers via drag‑and‑drop. Versioning – Automatic snapshot of each change; view history and restore previous versions. Theme Switching – Light / dark UI themes toggled client‑side. Responsive UI – Built with HTML, CSS, HTMX and Alpine.js for a smooth, single‑page‑like experience. Installation # The project uses Python 3.10.6 and virtualenv . ```shell script Clone the repository # git clone https://github.com/radlab-dev-group/llm-router.git cd llm-router/llm_router_web Create and activate a virtual environment # python3 -m venv venv source venv/bin/activate Install dependencies # pip install -r requirements.txt &gt; **Note:** The only required packages are listed in `requirements.txt` (`flask` and `flask_sqlalchemy`). All other &gt; libraries are part of the broader `llm-router` repository. --- ## Running the Application ``` shell script # Start the Flask development server python app.py The UI will be available at http://localhost:8081 . The first run will redirect you to a setup page where you must create an initial admin user. Running with gunicorn (recommended for production) # The project now includes a small launch script that reads the host, port and debug mode from environment variables prefixed with LLM_ROUTER_WEB_ . LLM_ROUTER_WEB_CFG_HOST = 0 .0.0.0 \\ LLM_ROUTER_WEB_CFG_PORT = 8081 \\ LLM_ROUTER_WEB_CFG_DEBUG = true \\ ./run-llm-router-web.sh LLM_ROUTER_WEB_CFG_HOST – address to bind (default: 0.0.0.0 ). LLM_ROUTER_WEB_CFG_PORT – numeric port (default: 8081 ). LLM_ROUTER_WEB_CFG_DEBUG – any truthy value ( true , 1 , yes , on ) enables Flask debug mode (default: true ). The script will automatically start the application with four gunicorn workers. Adjust the number of workers or other gunicorn options inside run.sh as needed. Additional Flask environment variables # Variable Description Default FLASK_SECRET_KEY Secret key for session signing change-me-local DATABASE_URL SQLAlchemy database URL (SQLite by default) sqlite:///configs.db FLASK_ENV Set to production for HTTPS scheme handling – Configuration # All Flask configuration is performed in web/__init__.py via create_app() . Key settings: SQLALCHEMY_DATABASE_URI – points to the SQLite DB ( configs.db ) unless overridden. SQLALCHEMY_TRACK_MODIFICATIONS – disabled for performance. PREFERRED_URL_SCHEME – set to https when FLASK_ENV=production . The database schema is automatically created on first launch. If the order column is missing from the provider table (e.g., after a schema change), the helper _ensure_provider_order_column() adds it on startup. Endpoints Overview # URL Methods Description /setup GET, POST One‑time admin creation (first run). /login GET, POST User authentication. /logout GET End session. /admin/users GET, POST List users / add new user (admin only). /admin/users/&lt;id&gt;/edit POST Edit role or password (admin only). /admin/users/&lt;id&gt;/toggle_block POST Block / unblock a user (admin only). / GET Dashboard – list of user’s configs. /configs GET Same as dashboard (alternative view). /configs/new GET, POST Create a new empty configuration. /configs/import GET, POST Import configuration from JSON file or text. /configs/&lt;id&gt; GET Preview configuration (JSON view). /configs/&lt;id&gt;/export GET Download configuration as models-config.json . /configs/&lt;id&gt;/edit GET, POST Edit active models, rename config, add providers, etc. /configs/&lt;id&gt;/models/add POST Add a new model to a configuration. /models/&lt;id&gt…"},{"k":"llm-router-plugins/maskers/fast-masker/index.html","t":"Overview","s":"Masker plugins (archived)","x":"## Overview The **fast_masker** plugin provides a simple, rule‑based engine that scans a piece of text and replaces sensitive data ( e‑mail addresses, IPs, URLs, phone numbers, Polish PESEL identifiers, etc.) with clearly marked placeholde…","h":["Overview","Masking Rules","Utility Validators"],"b":"Overview # The fast_masker plugin provides a simple, rule‑based engine that scans a piece of text and replaces sensitive data ( e‑mail addresses, IPs, URLs, phone numbers, Polish PESEL identifiers, etc.) with clearly marked placeholders. The core component is the :class: ~llm_router_plugins.plugins.fast_masker.core.masker.FastMasker , which receives an ordered list of rule objects and applies each rule sequentially to the input text. Because the rules are applied in the order they are supplied, you can control precedence (e.g., replace URLs before e‑mails if needed). Masking Rules # Rule Placeholder What it Detects Notes EmailRule {{EMAIL}} E‑mail addresses (e.g., user@example.com ). Uses a permissive regex that matches the local‑part, @ , and a domain with a TLD. IpRule {{IP}} IPv4, IPv6 addresses and the hostname localhost . Also masks ports as {{IP}}:{{PORT}} when a port follows the address. Light validation of octet ranges; port is captured separately. UrlRule {{URL}} HTTP/HTTPS URLs and plain domain names (e.g., https://example.com , www.wp.pl ). Optional scheme, optional path/query/fragment. PhoneRule {{PHONE}} Various phone number formats, with optional country/area codes and separators ( +48 123 456 789 , 123-456-789 , (123) 456 7890 , 1234567890 ). Very permissive pattern. PeselRule {{PESEL}} Polish PESEL numbers (11‑digit personal identifiers). Validates checksum via is_valid_pesel . BankAccountRule {{BANK_ACCOUNT}} Polish IBAN numbers (full 28‑character form) and partially masked accounts where any group may contain X . Matches exact length; no further validation required. DateNumberRule {{DATE_NUM}} Numeric dates in forms YYYY.MM.DD , DD.MM.YYYY (also with - , / or whitespace separators). Handles surrounding whitespace; replaces with {{DATE_NUM}} . DateWordRule {{DATE_STR}} Textual dates in Polish and English (e.g., 12 stycznia 2023 , January 12, 2023 ). Supports month names, abbreviations, optional ordinal suffixes and commas. KrsRule {{KRS}} Polish KRS numbers (plain or hyphen‑separated) that pass checksum validation. Uses is_valid_krs for verification. MoneyRule {{MONEY}} Monetary amounts that contain a currency identifier (symbols, ISO codes, or Polish words) together with a number. Discards surrounding markdown emphasis; replaces whole match. NipRule {{NIP}} Polish NIP numbers (plain, hyphen‑separated, embedded in letters, or wrapped in markdown). Validates checksum via _is_valid_nip . PostalCodeRule {{POSTAL_CODE}} Polish postal codes in forms dd-ddd or ddddd , optionally wrapped in markdown emphasis. No checksum validation; format‑based detection only. RegonRule {{REGON}} Polish REGON numbers (9 or 14 digits, optionally split by single spaces) that pass checksum validation. Uses is_valid_regon after stripping spaces. StreetNameRule (beta) {{STREET}} Polish street names with optional house numbers (e.g., ul. Mickiewicza 12 , aleja Jana Pawła II ). Recognises common street type prefixes and abbreviations; case‑insensitive, diacritics supported. SimplePersonalDataRule (beta) {{MASKED}} Polish surnames loaded from CSV resources; matches whole words that start with an uppercase letter. Uses a pre‑loaded set of surnames with heuristic inflection handling. BaseRule (abstract) — Provides common behaviour for rules that only need a compiled regular expression and a placeholder. Concrete rules inherit from this class. Each rule implements the :class: ~llm_router_plugins.plugins.fast_masker.core.rule_interface.MaskerRuleI interface, exposing an apply(text: str) -&gt; str method that returns the transformed string. Utility Validators # The plugin provides a set of helper functions used by several masking rules to verify the correctness of identified identifiers: is_valid_pesel(pesel: str) -&gt; bool – validates Polish PESEL numbers by checking length, date components and checksum. is_valid_krs(krs: str) -&gt; bool – checks the KRS checksum for both plain and hyphen‑separated forms. is_valid_regon(regon: str) -&gt; bool…"},{"k":"changelog.html","t":"Changelog","s":"Release notes","x":"## Changelog | Version | Changelog | |---------|-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------…","h":["Changelog"],"b":"Changelog # Version Changelog 0.0.1 Initialization, License, setup, interface for each endpoint and sample ping EP. Autoloader of builtin endpoints and for the future implementations. 0.0.2 Add base models for api call (module llm_proxy_rest.data_models with error.py handling. Decorators to check required params and to measure the response time. 0.0.3 Proper AutoLoading for each found endpoint. Implementation of ApiTypesDispatcher , ApiModelConfig , ModelHandler . Ollama endpoints: / , tags . Added endpoint to full proxy with params. Streaming in case when external api provides stream. 0.0.4 All llama-service endpoints are refactored to llm-proxy-api . Refactoring base ep_run method. Proper handling system message, prompt name, model etc. 0.1.0 Repository name changed from llm-proxy-api to llm-router . Added class HttpRequestExecutor to handle http requests from EndpointWithHttpRequestI . Handled routing between any models: openai -&gt; ollama and ollama -&gt; openai 0.1.1 Prometheus metrics logging. Workers/Threads/Workers class is able to set by environments. Streaming fixes. Multi-providers for single model with default-balanced strategy. 0.2.0 Add balancing strategies: balanced , weighted , dynamic_weighted and first_available which works for streaming and non streaming requests. Included Prometheus metrics logging via /metrics endpoint. First stage of llm_router_lib library, to simply usage of llm-router-api . 0.2.1 Fix stream: OpenAI-&gt;Ollama, Ollama-&gt;OpenAI. Add Redis caching of availability of model providers (when using first_available strategy). Add llm_router_web module with simple flask-based frontend to manage llm-router config files. 0.2.2 Update dockerfile and requirements. Fix routing with vLLM. 0.2.3 New web configurator: Handling projects, configs for each user separately. First Available strategy is more powerful, a lot of improvements to efficiency. 0.2.4 Anonymizer module, integration anonymization with any endpoint (using dynamic payload analysis and full payload anonymisation), dedicated /api/anonymize_text endpoint as memory only anonymization. Whole router may be run in FORCE_ANONYMISATION mode. 0.3.0 Anonymization available with three strategies: fast_masker , genai , prov_masker . 0.3.1 Refactoring lb.strategies to be more flexible modular. Introduced MaskerPipeline and GuardrailPipeline both configured via env. Removed genai-based masking endpoint."}]}