{"version":"0.2.1","pages":[{"k":"overview.html","t":"Overview","s":"Getting started","x":"## llm‑router A lightweight, extensible gateway that exposes a clean **REST** API for interacting with multiple Large Language Model (LLM) providers (OpenAI, Ollama, vLLM, etc.). It centralises request validation, prompt management, model …","h":["llm‑router","✨ Key Features","📦 Quick Start","1️⃣ Create &amp; activate a virtual environment","2️⃣ Minimum required environment variable","3️⃣ Optional configuration (via environment)","4️⃣ Run the REST API","Extending with Custom Strategies","🛣️ Endpoints Overview","Built-in Text Utilities","Content Generation","Context QA (RAG-like)","Streaming vs. Non‑Streaming Responses","⚙️ Configuration Details","🛠️ Development","📜 License","📚 Changelog"],"b":"llm‑router # A lightweight, extensible gateway that exposes a clean REST API for interacting with multiple Large Language Model (LLM) providers (OpenAI, Ollama, vLLM, etc.). It centralises request validation, prompt management, model configuration and logging, allowing your application to talk to any supported LLM through a single, consistent interface. This project provides a robust solution for managing and routing requests to various LLM backends. It simplifies the integration of LLMs into your applications by offering a unified API and advanced features like load balancing strategies. ✨ Key Features # Feature Description Unified REST interface One endpoint schema works for OpenAI‑compatible, Ollama, vLLM and any future provider. Provider‑agnostic streaming The stream flag (default true ) controls whether the proxy forwards chunked responses as they arrive or returns a single aggregated payload. Built‑in prompt library Language‑aware system prompts stored under resources/prompts can be referenced automatically. Dynamic model configuration JSON file ( models-config.json ) defines providers, model name, default options and per‑model overrides. Request validation Pydantic models guarantee correct payloads; errors are returned with clear messages. Structured logging Configurable log level, filename, and optional JSON formatting. Health &amp; metadata endpoints /ping (simple 200 OK) and /tags (available model tags/metadata). Simple deployment One‑liner run script or python -m llm_proxy_rest.rest_api . Extensible conversation formats Basic chat, conversation with system prompt, and extended conversation with richer options (e.g., temperature, top‑k, custom system prompt). Multi‑provider model support Each model can be backed by multiple providers (VLLM, Ollama, OpenAI) defined in models-config.json . Provider selection abstraction ProviderChooser delegates to a configurable strategy, enabling easy swapping of load‑balancing, round‑robin, weighted‑random, etc. Load‑balanced default strategy LoadBalancedStrategy distributes requests evenly across providers using in‑memory usage counters. Dynamic model handling ModelHandler loads model definitions at runtime and resolves the appropriate provider per request. Pluggable endpoint architecture Automatic discovery and registration of all concrete EndpointI implementations via EndpointAutoLoader . Prometheus metrics integration Optional /metrics endpoint for latency, error counts, and provider usage statistics. Docker ready Dockerfile and scripts for containerised deployment. 📦 Quick Start # 1️⃣ Create &amp; activate a virtual environment # Base requirements # Prerequisite : radlab-ml-utils This project uses the radlab-ml-utils library for machine learning utilities (e.g., experiment/result logging with Weights &amp; Biases/wandb). Install it before working with ML-related parts: bash pip install git+https://github.com/radlab-dev-group/ml-utils.git For more options and details, see the library README: https://github.com/radlab-dev-group/ml-utils ```shell script python3 -m venv .venv source .venv/bin/activate Only the core library (llm-router-lib). # pip install . Core library + API wrapper (llm-router-api). # pip install .[api] #### Prometheus Metrics To enable Prometheus metrics collection you must install the optional metrics dependencies: ``` bash pip install .[api,metrics] Then start the application with the environment variable set: export LLM_ROUTER_USE_PROMETHEUS = 1 When LLM_ROUTER_USE_PROMETHEUS is enabled, the router automatically registers a /metrics endpoint (under the API prefix, e.g. /api/metrics ). This endpoint exposes Prometheus‑compatible metrics such as request counts, latencies, and any custom counters defined by the application. Prometheus servers can scrape this URL to collect runtime metrics for monitoring and alerting. 2️⃣ Minimum required environment variable # ``` shell script ./run-rest-api.sh or # LLM_ROUTER_MINIMUM=1 python3 -m llm_router_api.rest_api ### 📦 D…"},{"k":"install.html","t":"Instalacja zależności","s":"Getting started","x":"# Instalacja zależności ## ml-utils ```bash git clone https://github.com/radlab-dev-group/ml-utils ``` ## llm-proxy-api ```shell pip install . --break ```","h":["ml-utils","llm-proxy-api"],"b":"Instalacja zależności # ml-utils # git clone https://github.com/radlab-dev-group/ml-utils llm-proxy-api # pip install . --break"},{"k":"llm-router-api/index.html","t":"REST API reference","s":"REST API","x":"# llm‑router‑api **llm‑router‑api** is a lightweight Python library that provides a flexible, extensible proxy for Large Language Model ( LLM) back‑ends. It abstracts the details of multiple model providers (OpenAI‑compatible, Ollama, vLLM…","h":["Table of Contents","Features","Installation","Running the Server","REST API Overview","Load‑Balancing Strategies","Extending the Router","Adding a New Provider Type","Adding a New Endpoint","Prompt Files","Monitoring &amp; Metrics","License"],"b":"llm‑router‑api # llm‑router‑api is a lightweight Python library that provides a flexible, extensible proxy for Large Language Model ( LLM) back‑ends. It abstracts the details of multiple model providers (OpenAI‑compatible, Ollama, vLLM, LM Studio, etc.) and offers a unified REST interface with built‑in load‑balancing, health‑checking, and monitoring. Repository: https://github.com/radlab-dev-group/llm-router Table of Contents # Features Installation Configuration Running the Server REST API Overview Load‑Balancing Strategies Extending the Router Monitoring &amp; Metrics Development &amp; Testing License Features # Unified API – One REST surface ( /api/... ) that proxies calls to any supported LLM back‑end. Provider Selection – Choose a provider per request using pluggable strategies (balanced, weighted, adaptive, first‑available). Prompt Management – System prompts are stored as files and can be dynamically injected with placeholder substitution. Streaming Support – Transparent streaming for both OpenAI‑compatible and Ollama endpoints. Health Checks – Built‑in ping endpoint and Redis‑based provider health monitoring. Prometheus Metrics – Optional instrumentation for request counts, latencies, and error rates. Auto‑Discovery – Endpoints are automatically discovered and instantiated at startup. Extensible – Add new providers, strategies, or custom endpoints with minimal boilerplate. Installation # The project uses Python 3.10.6 and a virtualenv ‑based workflow. ```shell script Clone the repository # git clone https://github.com/radlab-dev-group/llm-router.git cd llm-router Create a virtual environment # python3 -m venv venv source venv/bin/activate Install the package (including optional extras) # pip install -e .[metrics] # installs Prometheus support All required third‑party libraries are listed in `requirements.txt` (e.g., Flask, requests, redis, rdl‑ml‑utils, etc.). --- ## Configuration Configuration is driven primarily by environment variables and a JSON model‑config file. ### Environment Variables | Variable | Description | Default | |---------------------------------------------------|------------------------------------------------------------------------------------------|----------------------------------------| | `LLM_ROUTER_PROMPTS_DIR` | Directory containing system prompt files. | `resources/prompts` | | `LLM_ROUTER_MODELS_CONFIG` | Path to the JSON file defining models and providers. | `resources/configs/models-config.json` | | `LLM_ROUTER_EXTERNAL_TIMEOUT` | HTTP timeout (seconds) for outbound LLM calls. | `300` | | `LLM_ROUTER_TIMEOUT` | Timeout for the proxy server itself. | `0` (no timeout) | | `LLM_ROUTER_LOG_FILENAME` | Log file name for the router. | `llm-router.log` | | `LLM_ROUTER_LOG_LEVEL` | Logging level (`DEBUG`, `INFO`, …). | `INFO` | | `LLM_ROUTER_EP_PREFIX` | Global URL prefix (e.g., `/api`). | `/api` | | `LLM_ROUTER_MINIMUM` | Must be set to enable proxy mode (`1`/`true`). | *required* | | `LLM_ROUTER_BALANCE_STRATEGY` | Load‑balancing strategy (`balanced`, `weighted`, `dynamic_weighted`, `first_available`). | `balanced` | | `LLM_ROUTER_REDIS_HOST` / `LLM_ROUTER_REDIS_PORT` | Redis connection details for provider locking/monitoring. | `\"\"` / `6379` | | `LLM_ROUTER_USE_PROMETHEUS` | Enable Prometheus metrics (`1`/`true`). | `False` | | `LLM_ROUTER_SERVER_TYPE` | Server backend (`flask`, `gunicorn`, `waitress`). | `flask` | | `LLM_ROUTER_SERVER_PORT` | Port the server listens on. | `8080` | | `LLM_ROUTER_SERVER_HOST` | Host/interface to bind. | `0.0.0.0` | | `LLM_ROUTER_SERVER_WORKERS_COUNT` | Number of workers (Gunicorn/Waitress). | `2` | | `LLM_ROUTER_SERVER_THREADS_COUNT` | Number of threads per worker. | `8` | | `LLM_ROUTER_SERVER_WORKER_CLASS` | Gunicorn worker class (e.g., `gevent`). | *empty* | ### Model Configuration `models-config.json` follows the schema: ```json { \"active_models\": { \"openai_models\": [ \"gpt-4\", \"gpt-3.5-turbo\" ], \"ollama_models\": [ \"llama2\" ] }, \"openai_models\": { \"gpt-4\":…"},{"k":"llm-router-api/endpoints/readme-pl.html","t":"Jak tworzyć endpointy w llm-proxy-api – przewodnik","s":"REST API","x":"## Jak tworzyć endpointy w llm-proxy-api – przewodnik Poniżej zebrano kluczowe informacje o tym, jak definiować i konfigurować endpointy (EP) na podstawie klas z `endpoints.*`, z odniesieniem do logiki wykonywania w `endpoint_i.EndpointWit…","h":["Jak tworzyć endpointy w llm-proxy-api – przewodnik","Propozycja EP: BatchFileSummaries – podsumowania plików z listy"],"b":"Jak tworzyć endpointy w llm-proxy-api – przewodnik # Poniżej zebrano kluczowe informacje o tym, jak definiować i konfigurować endpointy (EP) na podstawie klas z endpoints.* , z odniesieniem do logiki wykonywania w endpoint_i.EndpointWithHttpRequestI.run_ep(...) . Uwzględniono też role atrybutów/stałych takich jak self._map_prompt , self._prompt_str_postfix , _prepare_response_function , _prompt_str_force , SYSTEM_PROMPT_NAME , REQUIRED_ARGS , OPTIONAL_ARGS oraz parametry konstruktora. 1) Hierarchia i warianty bazowe - EndpointI: baza dla EP (gdy serwis nie działa jako proxy). Definiuje ogólne API i walidację argumentów, ale nie implementuje run_ep. - EndpointWithHttpRequestI: rozszerza EndpointI o wysyłkę żądań HTTP do zewnętrznego LLM. Ma pełną implementację run_ep, obsługę streamingu i wstrzykiwania promptu systemowego. - PassthroughI: dziedziczy z EndpointWithHttpRequestI i domyślnie “przepuszcza” payload (prepare_payload zwraca parametry bez zmian). Użyteczne dla OpenAI‑kompatybilnych EP, gdzie chcemy prosto forwardować żądania. Dlaczego w openai.py dziedziczymy z PassthroughI? - Bo endpointy OpenAI‑kompatybilne często wymagają minimalnej logiki – wystarczy przekazać dalej to, co przyszło. PassthroughI upraszcza implementację (brak wymuszonych argumentów, brak system promptu, gotowy run_ep proxy). 2) Cykl wykonania – co robi run_ep w EndpointWithHttpRequestI W dużym skrócie: - Inicjalizacja zegara i wyzerowanie atrybutów promptu: _map_prompt , _prompt_str_force , _prompt_str_postfix . - Wywołanie prepare_payload(params): tu podklasa ma przekształcić wejście do formatu, jaki rozumie backend (np. ułożyć messages, przepisać model_name → model, ustawić stream itp.). Jeśli zwróci strukturę z \"status\": False , run_ep zwróci ją bez dalszego przetwarzania. - Jeśli ustawiono direct_return=True, zwracany jest wynik prepare_payload bez proxy. - Tryb “simple proxy”: jeżeli klasa nie definiuje REQUIRED_ARGS (pusta lista) – traktujemy EP jako bezpośredni proxy do odpowiednika po stronie modelu. Wtedy: - _set_model wybiera model na podstawie pól z MODEL_NAME_PARAMS. - Jeżeli typ API modelu jest zgodny z typami EP ( api_types ), payload jest przekazywany dalej do odpowiedniego URL (z opcjonalnym stream). - Jeżeli to nie simple proxy: - _resolve_prompt_name(...) przygotowuje system prompt (opisane w pkt 3). - __dispatch_external_api_model(params) ustawia _api_model na podstawie nazwy modelu. - Wyznaczamy docelowy URL przez ApiTypesDispatcher (np. chat_ep dla danego api_type ). - Obsługa stream=False/True (w tym wariancie streaming może być ograniczony – komunikat o braku wsparcia). - _call_http_request(...) wykonuje POST/GET do hosta modelu, składając finalny payload (w tym system message, jeśli jest). Dodatkowe ścieżki: - call_for_each_user_msg=True : dla zadań wielotekstowych – wysyłamy osobne żądanie dla każdej wiadomości użytkownika, a wynik agregujemy przez _prepare_response_function . 3) System prompt i modyfikacje treści – jak działają pola - SYSTEM_PROMPT_NAME: słownik { \"pl\": prompt_id, \"en\": prompt_id }. W prepare_payload ustawiasz wymagania EP, a run_ep w _resolve_prompt_name : - wybiera język z parametru LANGUAGE_PARAM (z defaultem DEFAULT_EP_LANGUAGE), - pobiera treść promptu systemowego przez PromptHandler jeśli zdefiniowano nazwę, - stosuje _map_prompt – słownik zamian {placeholder: tekst}, np. wstrzyknięcie liczby pytań, treści zapytania użytkownika, - dokleja _prompt_str_postfix na końcu system promptu (np. dodatkowa instrukcja), - jeśli _prompt_str_force jest ustawione – nadpisuje całą treść system promptu (pomija nazwę/system prompt z plików). Efekt: jeśli _prompt_str ostatecznie jest zbudowany, to zostaje dodany do messages jako pierwszy element: {\"role\": \"system\", \"content\": self._prompt_str}. Kiedy to ustawiać? - W prepare_payload: - self._map_prompt: gdy chcesz w promptach z zasobów podmienić znaczniki (np. ##QUESTION_NUM_STR##). - self._prompt_str_postfix: gdy EP potrzebuje dokleić końcową uwagę/regułę do system pr…"},{"k":"llm-router-lib/index.html","t":"llm‑router — Python client library","s":"Python library","x":"# llm‑router — Python client library **llm‑router** is a lightweight Python client for interacting with the LLM‑Router API. It provides typed request models, convenient service wrappers, and robust error handling so you can focus on buildi…","h":["Table of Contents","Overview","Features","Installation","Quick start","Extended conversation","Core concepts","Client","Data models","Services","Utilities","Error handling","License"],"b":"llm‑router — Python client library # llm‑router is a lightweight Python client for interacting with the LLM‑Router API. It provides typed request models, convenient service wrappers, and robust error handling so you can focus on building LLM‑driven applications rather than dealing with raw HTTP calls. Table of Contents # Overview Features Installation Quick start Core concepts Client Data models Services Utilities Error handling Testing Contributing License Overview # llm_router_lib is the official Python SDK for the LLM‑Router project https://github.com/radlab-dev-group/llm-router . It abstracts the HTTP layer behind a small, well‑typed API: Typed payloads built with pydantic (e.g., GenerativeConversationModel ). Service objects that know the endpoint URL and the model class they expect. Automatic token handling , request retries, and exponential back‑off. Rich exception hierarchy ( LLMRouterError , AuthenticationError , RateLimitError , ValidationError ). Features # Feature Description Typed request/response models Guarantees payload correctness at runtime using Pydantic. Built‑in conversation services Simple conversation_with_model and extended_conversation_with_model calls. Retry &amp; timeout Configurable request timeout and automatic retries with exponential back‑off. Authentication Bearer‑token support; raises AuthenticationError on 401/403. Rate‑limit handling Detects HTTP 429 and raises RateLimitError . Extensible Add custom services or models by extending the base classes. Test suite Ready‑to‑run unit tests in llm_router_lib/tests . Installation # The library is pure Python and works with Python 3.10+ . ```shell script Create a virtualenv (recommended) # python -m venv .venv source .venv/bin/activate Install from the repository (editable mode) # pip install -e . If you prefer a regular installation from a wheel or source distribution, use: ```shell script pip install . Note – The project relies only on the packages listed in the repository’s requirements.txt (pydantic, requests, etc.), all of which are installed automatically by pip . Quick start # from llm_router_lib.client import LLMRouterClient from llm_router_lib.data_models.builtin_chat import GenerativeConversationModel # Initialise the client (replace with your own endpoint and token) client = LLMRouterClient ( api = \"https://api.your-llm-router.com\" , token = \"YOUR_ACCESS_TOKEN\" ) # Build a request payload payload = GenerativeConversationModel ( model_name = \"google/gemma-3-12b-it\" , user_last_statement = \"Hello, how are you?\" , historical_messages = [{ \"user\" : \"Hi\" }], temperature = 0.7 , max_new_tokens = 128 , ) # Call the API response = client . conversation_with_model ( payload ) print ( response ) # → dict with the model's answer and metadata Extended conversation # from llm_router_lib.data_models.builtin_chat import ExtendedGenerativeConversationModel payload = ExtendedGenerativeConversationModel ( model_name = \"google/gemma-3-12b-it\" , user_last_statement = \"Explain quantum entanglement.\" , system_prompt = \"Answer as a friendly professor.\" , temperature = 0.6 , max_new_tokens = 256 , ) response = client . extended_conversation_with_model ( payload ) print ( response ) Core concepts # Client # LLMRouterClient is the entry point. It handles: Base URL normalization. Optional bearer token injection. Construction of the internal HttpRequester . All public methods accept either a dict or a pydantic model ; models are automatically serialized with .model_dump() . Data models # Located in llm_router_lib/data_models/ . Key models: Model Purpose GenerativeConversationModel Simple chat payload (model name, user message, optional history). ExtendedGenerativeConversationModel Same as above, plus a system_prompt . GenerateQuestionFromTextsModel Generate questions from a list of texts. TranslateTextModel , SimplifyTextModel , … Various utility models for text transformation. OpenAIChatModel Payload for direct OpenAI‑compatible chat calls. All models inherit from a …"},{"k":"llm-router-web/index.html","t":"llmrouterweb","s":"Web UI (archived)","x":"# llm_router_web **llm_router_web** is the web interface component of the **[llm-router](https://github.com/radlab-dev-group/llm-router) ** project. It provides a Flask‑based UI for managing LLM model configurations, users, and version his…","h":["Table of Contents","Features","Installation","Environment variables","Configuration","Endpoints Overview","Development","Code structure","Adding new features","Testing","License"],"b":"llm_router_web # llm_router_web is the web interface component of the ** llm-router ** project. It provides a Flask‑based UI for managing LLM model configurations, users, and version history. Table of Contents # Features Installation Running the Application Configuration Endpoints Overview Development License Features # User Management – Admins can create, edit, block/unblock users and assign roles ( admin / user ). Configuration CRUD – Create, import, edit, view, export, activate and delete model configurations. Model &amp; Provider Management – Add/remove models, manage multiple providers per model, reorder providers via drag‑and‑drop. Versioning – Automatic snapshot of each change; view history and restore previous versions. Theme Switching – Light / dark UI themes toggled client‑side. Responsive UI – Built with HTML, CSS, HTMX and Alpine.js for a smooth, single‑page‑like experience. Installation # The project uses Python 3.10.6 and virtualenv . ```shell script Clone the repository # git clone https://github.com/radlab-dev-group/llm-router.git cd llm-router/llm_router_web Create and activate a virtual environment # python3 -m venv venv source venv/bin/activate Install dependencies # pip install -r requirements.txt &gt; **Note:** The only required packages are listed in `requirements.txt` (`flask` and `flask_sqlalchemy`). All other &gt; libraries are part of the broader `llm-router` repository. --- ## Running the Application ```shell script # Start the Flask development server python app.py The UI will be available at http://localhost:8081 . The first run will redirect you to a setup page where you must create an initial admin user. Environment variables # Variable Description Default FLASK_SECRET_KEY Secret key for session signing change-me-local DATABASE_URL SQLAlchemy database URL (SQLite by default) sqlite:///configs.db FLASK_ENV Set to production for HTTPS scheme handling – Configuration # All Flask configuration is performed in web/__init__.py via create_app() . Key settings: SQLALCHEMY_DATABASE_URI – points to the SQLite DB ( configs.db ) unless overridden. SQLALCHEMY_TRACK_MODIFICATIONS – disabled for performance. PREFERRED_URL_SCHEME – set to https when FLASK_ENV=production . The database schema is automatically created on first launch. If the order column is missing from the provider table (e.g., after a schema change), the helper _ensure_provider_order_column() adds it on startup. Endpoints Overview # URL Methods Description /setup GET, POST One‑time admin creation (first run). /login GET, POST User authentication. /logout GET End session. /admin/users GET, POST List users / add new user (admin only). /admin/users/&lt;id&gt;/edit POST Edit role or password (admin only). /admin/users/&lt;id&gt;/toggle_block POST Block / unblock a user (admin only). / GET Dashboard – list of user’s configs. /configs GET Same as dashboard (alternative view). /configs/new GET, POST Create a new empty configuration. /configs/import GET, POST Import configuration from JSON file or text. /configs/&lt;id&gt; GET Preview configuration (JSON view). /configs/&lt;id&gt;/export GET Download configuration as models-config.json . /configs/&lt;id&gt;/edit GET, POST Edit active models, rename config, add providers, etc. /configs/&lt;id&gt;/models/add POST Add a new model to a configuration. /models/&lt;id&gt;/delete POST Delete a model. /models/&lt;id&gt;/providers/add POST Add a provider (JSON payload). /models/&lt;id&gt;/providers/reorder POST Reorder providers (drag‑and‑drop). /providers/&lt;id&gt;/update POST Update provider fields (JSON payload). /providers/&lt;id&gt;/delete POST Delete a provider. /configs/&lt;id&gt;/activate POST Mark a configuration as the default for the user. /configs/&lt;id&gt;/delete POST Delete a configuration. /configs/&lt;id&gt;/versions GET List version history (JSON). /configs/&lt;id&gt;/versions/&lt;ver&gt;/restore POST Restore a previous version. /check_host POST Verify reachability of an API host (used by the …"},{"k":"changelog.html","t":"Changelog","s":"Release notes","x":"## Changelog | Version | Changelog | |---------|-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------…","h":["Changelog"],"b":"Changelog # Version Changelog 0.0.1 Initialization, License, setup, interface for each endpoint and sample ping EP. Autoloader of builtin endpoints and for the future implementations. 0.0.2 Add base models for api call (module llm_proxy_rest.data_models with error.py handling. Decorators to check required params and to measure the response time. 0.0.3 Proper AutoLoading for each found endpoint. Implementation of ApiTypesDispatcher , ApiModelConfig , ModelHandler . Ollama endpoints: / , tags . Added endpoint to full proxy with params. Streaming in case when external api provides stream. 0.0.4 All llama-service endpoints are refactored to llm-proxy-api . Refactoring base ep_run method. Proper handling system message, prompt name, model etc. 0.1.0 Repository name changed from llm-proxy-api to llm-router . Added class HttpRequestExecutor to handle http requests from EndpointWithHttpRequestI . Handled routing between any models: openai -&gt; ollama and ollama -&gt; openai 0.1.1 Prometheus metrics logging. Workers/Threads/Workers class is able to set by environments. Streaming fixes. Multi-providers for single model with default-balanced strategy. 0.2.0 Add balancing strategies: balanced , weighted , dynamic_weighted and first_available which works for streaming and non streaming requests. Included Prometheus metrics logging via /metrics endpoint. First stage of llm_router_lib library, to simply usage of llm-router-api . 0.2.1 Fix stream: OpenAI-&gt;Ollama, Ollama-&gt;OpenAI. Add Redis caching of availability of model providers (when using first_available strategy). Add llm_router_web module with simple flask-based frontend to manage llm-router config files."}]}