{"version":"0.2.0","pages":[{"k":"overview.html","t":"Overview","s":"Getting started","x":"## llm‑router A lightweight, extensible gateway that exposes a clean **REST** API for interacting with multiple Large Language Model (LLM) providers (OpenAI, Ollama, vLLM, etc.). It centralises request validation, prompt management, model …","h":["llm‑router","✨ Key Features","📦 Quick Start","1️⃣ Create &amp; activate a virtual environment","2️⃣ Minimum required environment variable","3️⃣ Optional configuration (via environment)","4️⃣ Run the REST API","Extending with Custom Strategies","🛣️ Endpoints Overview","Built-in Text Utilities","Content Generation","Context QA (RAG-like)","Streaming vs. Non‑Streaming Responses","⚙️ Configuration Details","🛠️ Development","📜 License","📚 Changelog"],"b":"llm‑router # A lightweight, extensible gateway that exposes a clean REST API for interacting with multiple Large Language Model (LLM) providers (OpenAI, Ollama, vLLM, etc.). It centralises request validation, prompt management, model configuration and logging, allowing your application to talk to any supported LLM through a single, consistent interface. This project provides a robust solution for managing and routing requests to various LLM backends. It simplifies the integration of LLMs into your applications by offering a unified API and advanced features like load balancing strategies. ✨ Key Features # Feature Description Unified REST interface One endpoint schema works for OpenAI‑compatible, Ollama, vLLM and any future provider. Provider‑agnostic streaming The stream flag (default true ) controls whether the proxy forwards chunked responses as they arrive or returns a single aggregated payload. Built‑in prompt library Language‑aware system prompts stored under resources/prompts can be referenced automatically. Dynamic model configuration JSON file ( models-config.json ) defines providers, model name, default options and per‑model overrides. Request validation Pydantic models guarantee correct payloads; errors are returned with clear messages. Structured logging Configurable log level, filename, and optional JSON formatting. Health &amp; metadata endpoints /ping (simple 200 OK) and /tags (available model tags/metadata). Simple deployment One‑liner run script or python -m llm_proxy_rest.rest_api . Extensible conversation formats Basic chat, conversation with system prompt, and extended conversation with richer options (e.g., temperature, top‑k, custom system prompt). Multi‑provider model support Each model can be backed by multiple providers (VLLM, Ollama, OpenAI) defined in models-config.json . Provider selection abstraction ProviderChooser delegates to a configurable strategy, enabling easy swapping of load‑balancing, round‑robin, weighted‑random, etc. Load‑balanced default strategy LoadBalancedStrategy distributes requests evenly across providers using in‑memory usage counters. Dynamic model handling ModelHandler loads model definitions at runtime and resolves the appropriate provider per request. Pluggable endpoint architecture Automatic discovery and registration of all concrete EndpointI implementations via EndpointAutoLoader . Prometheus metrics integration Optional /metrics endpoint for latency, error counts, and provider usage statistics. Docker ready Dockerfile and scripts for containerised deployment. 📦 Quick Start # 1️⃣ Create &amp; activate a virtual environment # Base requirements # Prerequisite : radlab-ml-utils This project uses the radlab-ml-utils library for machine learning utilities (e.g., experiment/result logging with Weights &amp; Biases/wandb). Install it before working with ML-related parts: bash pip install git+https://github.com/radlab-dev-group/ml-utils.git For more options and details, see the library README: https://github.com/radlab-dev-group/ml-utils ```shell script python3 -m venv .venv source .venv/bin/activate Only the core library (llm-router-lib). # pip install . Core library + API wrapper (llm-router-api). # pip install .[api] #### Prometheus Metrics To enable Prometheus metrics collection you must install the optional metrics dependencies: ``` bash pip install .[api,metrics] Then start the application with the environment variable set: export LLM_ROUTER_USE_PROMETHEUS = 1 When LLM_ROUTER_USE_PROMETHEUS is enabled, the router automatically registers a /metrics endpoint (under the API prefix, e.g. /api/metrics ). This endpoint exposes Prometheus‑compatible metrics such as request counts, latencies, and any custom counters defined by the application. Prometheus servers can scrape this URL to collect runtime metrics for monitoring and alerting. 2️⃣ Minimum required environment variable # ``` shell script ./run-rest-api.sh or # LLM_ROUTER_MINIMUM=1 python3 -m llm_router_api.rest_api ### 📦 D…"},{"k":"install.html","t":"Instalacja zależności","s":"Getting started","x":"# Instalacja zależności ## ml-utils ```bash git clone https://github.com/radlab-dev-group/ml-utils ``` ## llm-proxy-api ```shell pip install . --break ```","h":["ml-utils","llm-proxy-api"],"b":"Instalacja zależności # ml-utils # git clone https://github.com/radlab-dev-group/ml-utils llm-proxy-api # pip install . --break"},{"k":"llm-router-api/endpoints/readme-pl.html","t":"Jak tworzyć endpointy w llm-proxy-api – przewodnik","s":"REST API","x":"## Jak tworzyć endpointy w llm-proxy-api – przewodnik Poniżej zebrano kluczowe informacje o tym, jak definiować i konfigurować endpointy (EP) na podstawie klas z `endpoints.*`, z odniesieniem do logiki wykonywania w `endpoint_i.EndpointWit…","h":["Jak tworzyć endpointy w llm-proxy-api – przewodnik","Propozycja EP: BatchFileSummaries – podsumowania plików z listy"],"b":"Jak tworzyć endpointy w llm-proxy-api – przewodnik # Poniżej zebrano kluczowe informacje o tym, jak definiować i konfigurować endpointy (EP) na podstawie klas z endpoints.* , z odniesieniem do logiki wykonywania w endpoint_i.EndpointWithHttpRequestI.run_ep(...) . Uwzględniono też role atrybutów/stałych takich jak self._map_prompt , self._prompt_str_postfix , _prepare_response_function , _prompt_str_force , SYSTEM_PROMPT_NAME , REQUIRED_ARGS , OPTIONAL_ARGS oraz parametry konstruktora. 1) Hierarchia i warianty bazowe - EndpointI: baza dla EP (gdy serwis nie działa jako proxy). Definiuje ogólne API i walidację argumentów, ale nie implementuje run_ep. - EndpointWithHttpRequestI: rozszerza EndpointI o wysyłkę żądań HTTP do zewnętrznego LLM. Ma pełną implementację run_ep, obsługę streamingu i wstrzykiwania promptu systemowego. - PassthroughI: dziedziczy z EndpointWithHttpRequestI i domyślnie “przepuszcza” payload (prepare_payload zwraca parametry bez zmian). Użyteczne dla OpenAI‑kompatybilnych EP, gdzie chcemy prosto forwardować żądania. Dlaczego w openai.py dziedziczymy z PassthroughI? - Bo endpointy OpenAI‑kompatybilne często wymagają minimalnej logiki – wystarczy przekazać dalej to, co przyszło. PassthroughI upraszcza implementację (brak wymuszonych argumentów, brak system promptu, gotowy run_ep proxy). 2) Cykl wykonania – co robi run_ep w EndpointWithHttpRequestI W dużym skrócie: - Inicjalizacja zegara i wyzerowanie atrybutów promptu: _map_prompt , _prompt_str_force , _prompt_str_postfix . - Wywołanie prepare_payload(params): tu podklasa ma przekształcić wejście do formatu, jaki rozumie backend (np. ułożyć messages, przepisać model_name → model, ustawić stream itp.). Jeśli zwróci strukturę z \"status\": False , run_ep zwróci ją bez dalszego przetwarzania. - Jeśli ustawiono direct_return=True, zwracany jest wynik prepare_payload bez proxy. - Tryb “simple proxy”: jeżeli klasa nie definiuje REQUIRED_ARGS (pusta lista) – traktujemy EP jako bezpośredni proxy do odpowiednika po stronie modelu. Wtedy: - _set_model wybiera model na podstawie pól z MODEL_NAME_PARAMS. - Jeżeli typ API modelu jest zgodny z typami EP ( api_types ), payload jest przekazywany dalej do odpowiedniego URL (z opcjonalnym stream). - Jeżeli to nie simple proxy: - _resolve_prompt_name(...) przygotowuje system prompt (opisane w pkt 3). - __dispatch_external_api_model(params) ustawia _api_model na podstawie nazwy modelu. - Wyznaczamy docelowy URL przez ApiTypesDispatcher (np. chat_ep dla danego api_type ). - Obsługa stream=False/True (w tym wariancie streaming może być ograniczony – komunikat o braku wsparcia). - _call_http_request(...) wykonuje POST/GET do hosta modelu, składając finalny payload (w tym system message, jeśli jest). Dodatkowe ścieżki: - call_for_each_user_msg=True : dla zadań wielotekstowych – wysyłamy osobne żądanie dla każdej wiadomości użytkownika, a wynik agregujemy przez _prepare_response_function . 3) System prompt i modyfikacje treści – jak działają pola - SYSTEM_PROMPT_NAME: słownik { \"pl\": prompt_id, \"en\": prompt_id }. W prepare_payload ustawiasz wymagania EP, a run_ep w _resolve_prompt_name : - wybiera język z parametru LANGUAGE_PARAM (z defaultem DEFAULT_EP_LANGUAGE), - pobiera treść promptu systemowego przez PromptHandler jeśli zdefiniowano nazwę, - stosuje _map_prompt – słownik zamian {placeholder: tekst}, np. wstrzyknięcie liczby pytań, treści zapytania użytkownika, - dokleja _prompt_str_postfix na końcu system promptu (np. dodatkowa instrukcja), - jeśli _prompt_str_force jest ustawione – nadpisuje całą treść system promptu (pomija nazwę/system prompt z plików). Efekt: jeśli _prompt_str ostatecznie jest zbudowany, to zostaje dodany do messages jako pierwszy element: {\"role\": \"system\", \"content\": self._prompt_str}. Kiedy to ustawiać? - W prepare_payload: - self._map_prompt: gdy chcesz w promptach z zasobów podmienić znaczniki (np. ##QUESTION_NUM_STR##). - self._prompt_str_postfix: gdy EP potrzebuje dokleić końcową uwagę/regułę do system pr…"},{"k":"changelog.html","t":"Changelog","s":"Release notes","x":"## Changelog | Version | Changelog | |---------|-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------…","h":["Changelog"],"b":"Changelog # Version Changelog 0.0.1 Initialization, License, setup, interface for each endpoint and sample ping EP. Autoloader of builtin endpoints and for the future implementations. 0.0.2 Add base models for api call (module llm_proxy_rest.data_models with error.py handling. Decorators to check required params and to measure the response time. 0.0.3 Proper AutoLoading for each found endpoint. Implementation of ApiTypesDispatcher , ApiModelConfig , ModelHandler . Ollama endpoints: / , tags . Added endpoint to full proxy with params. Streaming in case when external api provides stream. 0.0.4 All llama-service endpoints are refactored to llm-proxy-api . Refactoring base ep_run method. Proper handling system message, prompt name, model etc. 0.1.0 Repository name changed from llm-proxy-api to llm-router . Added class HttpRequestExecutor to handle http requests from EndpointWithHttpRequestI . Handled routing between any models: openai -&gt; ollama and ollama -&gt; openai 0.1.1 Prometheus metrics logging. Workers/Threads/Workers class is able to set by environments. Streaming fixes. Multi-providers for single model with default-balanced strategy. 0.2.0 Add balancing strategies: balanced , weighted , dynamic_weighted and first_available which works for streaming and non streaming requests. Included Prometheus metrics logging via /metrics endpoint. First stage of llm_router_lib library, to simply usage of llm-router-api ."}]}