llm-router/docs

Changelog#

Version Changelog
0.0.1 Initialization, License, setup, interface for each endpoint and sample ping EP. Autoloader of builtin endpoints and for the future implementations.
0.0.2 Add base models for api call (module llm_proxy_rest.data_models with error.py handling. Decorators to check required params and to measure the response time.
0.0.3 Proper AutoLoading for each found endpoint. Implementation of ApiTypesDispatcher, ApiModelConfig, ModelHandler. Ollama endpoints: /, tags. Added endpoint to full proxy with params. Streaming in case when external api provides stream.
0.0.4 All llama-service endpoints are refactored to llm-proxy-api. Refactoring base ep_run method. Proper handling system message, prompt name, model etc.
0.1.0 Repository name changed from llm-proxy-api to llm-router. Added class HttpRequestExecutor to handle http requests from EndpointWithHttpRequestI. Handled routing between any models: openai -> ollama and ollama -> openai
0.1.1 Prometheus metrics logging. Workers/Threads/Workers class is able to set by environments. Streaming fixes. Multi-providers for single model with default-balanced strategy.
0.2.0 Add balancing strategies: balanced, weighted, dynamic_weighted and first_available which works for streaming and non streaming requests. Included Prometheus metrics logging via /metrics endpoint. First stage of llm_router_lib library, to simply usage of llm-router-api.
0.2.1 Fix stream: OpenAI->Ollama, Ollama->OpenAI. Add Redis caching of availability of model providers (when using first_available strategy). Add llm_router_web module with simple flask-based frontend to manage llm-router config files.
0.2.2 Update dockerfile and requirements. Fix routing with vLLM.
0.2.3 New web configurator: Handling projects, configs for each user separately. First Available strategy is more powerful, a lot of improvements to efficiency.
0.2.4 Anonymizer module, integration anonymization with any endpoint (using dynamic payload analysis and full payload anonymisation), dedicated /api/anonymize_text endpoint as memory only anonymization. Whole router may be run in FORCE_ANONYMISATION mode.
0.3.0 Anonymization available with three strategies: fast_masker, genai, prov_masker.
0.3.1 Refactoring lb.strategies to be more flexible modular. Introduced MaskerPipeline and GuardrailPipeline both configured via env. Removed genai-based masking endpoint.
0.4.0 The main repository is divided into dedicated ones: plugins, services, web — separate repositories. Clean up the whole repository. Examples of integration with llamaindex, langchain, openai, litellm and haystack.
0.4.1 Audit log is stored using GPG. Add bash script (scripts/gen_and_export_gpg.sh to prepare GPG keys and simple scripts/decrypt_auditor_logs.sh to decrypt encrypted audit logs. Moved core functionality from base to module core module. Quickstart.
0.4.2 Fix first_available_optim Strategy. Add KeepAliveMonitor to periodically pings model endpoints to keep them warm.
0.4.3 Add custom Prometheus metrices for logging masker/guardrail inidents. Fix OpenAI compatible v1 /models endpoint. Introduce monitors: services and keep alive models. Fixed guardrail retunr in case when streaming.
0.4.4 Validate unique provider identifiers. Store all hosts with keep‑alive configured in a Redis. UtilsPlugin pipeline with LangChain based simple RAG plugin (extending context to GenAI with locally built databse). Add handling of v1/response endpoint
0.4.5 Fixed sreaming to LMStudio native. Refactor streaming module.
0.4.6 Added support for embeddings endpoints across all providers. Extended ApiModel and ApiTypesI with is_embedding flag. Added test_embeddings.py utility for verifying embedding models through the API.
0.4.7 Integration with native Anthropic API. Add translate, generative_answer and ping methods to LLMRouterClient (with tests). Refactor LLMRouterCkientServices to use self.model_cls. Add payload converter for vLLM.
0.5.0 Integration with PII masker, code refactoring
0.5.1 Add /v1/messages endpoint (Claude Agent compatibility), streaming Cache‑Control/Pragma/Expires/Vary headers, mandatory Redis (runtime error on missing connection), non‑root Docker startup, remove ml‑utils dependency, add gnupg to requirements, update default plugin to simple_semantic_routing.
0.5.2 Prevent network topology leak in error messages (full details remain in server‑side logs), example models-config.json no longer contains real internal IPs. Introduced LLM_ROUTER_MAX_REQUEST_BODY_SIZE to set the maximum content length. Sanitize all error messages returned to the Client. Local security.
0.6.0 Authentication system: API key-based auth with multi-backend key stores (Memory, Redis, Vault), plaintext and secret-key lookup, enable/disable keys, seed-file persistence. Auth CLI: auth subcommands for managing API keys (create/list/enable/disable/delete) with formatted tabular output and prefix matching. Rate limiting: Per-key rate limiting via token bucket, predefined rate-limiting policies in rate_limiting-policies.json, PolicyEngine accepting dict key records. Anonymizer CLI: Migrated fast_masker to anonymizer CLI with deprecation warning; moved to masker subpackage. Infrastructure: Shared Redis client across stores and cache, dynamic column widths for CLI output, environment variable updates for auth/rate-limiting/audit logging config.
0.6.1 Added config CLI command with discover (auto-discover local Ollama/vLLM/LM Studio providers) and merge (deep-merge multiple models-config.json files).
0.6.2 CLI: fix _RATE_LIMIT_COMMANDS NameError in auth CLI (commit 2d9e593). Core: extract Prometheus multiprocess dir handling into MetricsHandler.prepare_multiproc_dir() with sane default; add LLM_ROUTER_AUTH_MEMORY_SEED_FILE env var for memory store seed path. CLI: restructure config commands, clean up imports across core modules, add llm_router_client.py test script with predefined model tests.
0.6.3 Core: consolidate streaming mode flags into StreamConversion enum, simplify handler dispatch and fix swapped handler type hints in OpenAI/Ollama endpoints. Simplify error handling and response logic in endpoint_i.py.
0.6.4 Refac: Clear REAME files.
0.6.5 CLI discover: added KoboldCpp and TabbyAPI providers to config discover auto-discovery (now covers 6 local providers: Ollama, vLLM, LM Studio, llama.cpp, KoboldCpp, TabbyAPI). Infrastructure: fixed _scan_and_merge passing best_port as host parameter causing empty config; removed deprecated llm-router-fast-masker entry point and masker module.
0.6.6 Auth CLI: restored working key generation by fixing _handle_key instance method call (cls._handle_key → cls()._handle_key). Config CLI: fixed _do_discover/_do_merge calls to use class methods properly (_do_discover(args) → cls._do_discover(args)). Removed all legacy backward-compatibility shim comments and module-level functions from CLI command modules.
0.6.7 Docs: consolidate environment variable documentation into ENV_DEFINITIONS.md, simplify references across README files, add semantic_biencoder_routing env var details, remove LLM_ROUTER_HOST and update default MASKING_STRATEGY_PIPELINE. Chore: add /metrics endpoint to public endpoints list (LLM_ROUTER_AUTH_PUBLIC_ENDPOINTS) across auth and config modules; bump version to 0.6.7.
0.6.8 Feature: Prometheus router metrics, Grafana dashboard, and /metrics public endpoint for improved observability and visualization. Docs: consolidate env var docs into ENV_DEFINITIONS.md, update docstrings and line wrapping across modules. Chore: expand mypy checks, add types-setuptools.
0.6.9 Chore: update prompts to be more specific. Feature: add prepare_response_function to endpoint_i, add GenerateNewsFromTextService, add default_model parameter to Client to provide fallback model configuration
0.7.0 Feature: add polarity_3c endpoint (/api/polarity_3c) to detect 3-class text polarity (ambivalent, positive, negative) with PL/EN system prompts. Lib: add Polarity3cModel, Polarity3cService, and LLMRouterClient.polarity_3c() method to llm_router_lib. Auth & Docs: register endpoint with builtin policy permission, update endpoint documentation and add comprehensive test suite.
0.8.0 Feature: add endpoints: /api/simplify_text, /api/generate_questions, /api/create_full_article_from_texts, /api/generate_article_from_texts, /api/generate_label with models, services and matching LLMRouterClient methods. Auth: register the new endpoints with the builtin policy permission. Fix _ensure_alternating_roles in endpoint_i to build a valid payload by merging consecutive messages of the same role (joining their contents), folding all system messages into one leading message, and prepending/appending empty user placeholders when the dialogue starts with or ends with an assistant turn.
0.9.0 Refactor (dependency direction): llm_router_lib no longer imports from llm_router_api. The shared DEFAULT_EP_LANGUAGE constant (and the LLM_ROUTER_ env prefix) now live in the dependency‑free llm_router_lib.core.constants; llm_router_api.base.constants_base re‑exports them for backward compatibility. BREAKING – structural rename anchored on the endpoint URL (HTTP paths are unchanged). Canonical name = URL without /api/; class/model/service/client‑method/constant names now follow it and legacy aliases are removed: TranslateTexts→Translate, TranslateTextModel→TranslateModel, TranslateTextService→TranslateService; SimplifyTexts→SimplifyText, client simplify_texts()→simplify_text(); GenerateQuestionsFromTexts→GenerateQuestions, GenerateQuestionFromTextsModel→GenerateQuestionsModel, GenerateQuestionsFromTextsService→GenerateQuestionsService (alias GenerateQuestionFromTextsService removed), client generate_questions_from_texts()→generate_questions(), constants GENERATE_Q_*→GENERATE_QUESTIONS_*; GenerateNewsFromTextHandler→GenerateArticleFromText, GenerateNewsFromTextService→GenerateArticleFromTextService, client generate_news_from_text()→generate_article_from_text(), constants GENERATE_ART_*→GENERATE_ARTICLE_FROM_TEXT_*; GenerateArticleFromTexts constants GENERATE_ARTICLES_*→GENERATE_ARTICLE_FROM_TEXTS_*; FullArticleFromTexts→CreateFullArticleFromTexts, CreateArticleFromNewsListModel→CreateFullArticleFromTextsModel, constants FULL_ARTICLE_*→CREATE_FULL_ARTICLE_FROM_TEXTS_*; AnswerBasedOnTheContext→GenerativeAnswer, AnswerBasedOnTheContextModel→GenerativeAnswerModel, constants CONTEXT_ANSWER_*→GENERATIVE_ANSWER_*; GenerativeConversationModel→ConversationWithModelRequest, ExtendedGenerativeConversationModel→ExtendedConversationWithModelRequest, ConversationService→ConversationWithModelService, ExtendedConversationService→ExtendedConversationWithModelService. Class inheritance hierarchy and the auth _ENDPOINT_PERMISSION_MAP (URL‑keyed) are unchanged. Refactor (structure only, no behavior change): shared base TextListUtilityEndpoint for the texts utility endpoints (common prepare_payload — one user message per source text, build_map_prompt hook, per‑text _prepare_response helper); message normalisation moved verbatim into pure functions in llm_router_api/endpoints/message_normalizer.py (thin delegating methods kept on the endpoint class); outbound HTTP dispatch + retry orchestration moved verbatim into llm_router_api/endpoints/http_dispatch.py (HttpDispatch; EndpointWithHttpRequestI.RetryResponse kept as a backward‑compatible alias of http_dispatch.RetryPolicy); run_ep slimmed into a readable orchestrator with extracted sequence helpers (endpoint_i.py reduced from 2123 to ~1810 lines). Fix (behavior): per‑text endpoints now actually return one result per source text — the role normalizer (introduced in 0.8.0 to build valid chat payloads) was applied unconditionally in run_ep and merged the consecutive user messages built by the texts endpoints, so polarity_3c, translate, simplify_text and generate_questions silently returned one merged result; run_ep now skips role normalisation for call_for_each_user_msg=True endpoints (the HTTP executor already dispatches each text as its own [system, user] call), while conversation endpoints (conversation_with_model, extended_conversation_with_model) keep it; client code consuming a single merged entry for those four endpoints must iterate the response list. Tests: regression grid extended to all remaining builtin endpoints; new test_http_dispatch.py (dispatch/retry, late‑binding, RetryResponse alias) and end‑to‑end test_per_text_dispatch.py (per‑text ⇒ one model call per text ⇒ results paired per text; chat endpoint ⇒ consecutive user messages still merged). Feature: Add log rotation. Feature: add method models() to client.py, add public /health endpoint
0.9.1 BREAKING – unified LLMRouterClient calling contract (pre‑1.0): all 12 endpoint methods (conversation_with_model, extended_conversation_with_model, polarity_3c, translate, simplify_text, generative_answer, generate_article_from_text, generate_article_from_texts, create_full_article_from_texts, generate_questions, generate_label, …) now share one keyword‑only signature: payload (a Pydantic request model instance) or named domain arguments + model + optional generation options; argument order is fixed (payload → domain fields → model → temperature → max_new_tokens). Raw dict payloads are removed – passing a dict as payload raises TypeError (build the matching model explicitly, e.g. MyModel(**payload)). Hard‑coded defaults in signatures (0.2/0.75/256/512/1024/64, number_of_questions=1) are removed – generation defaults come exclusively from the Pydantic models (GenerativeOptions); optional parameters default to None so None‑valued kwargs no longer break Pydantic validation (fixes the texts=None bug). conversation_with_model / extended_conversation_with_model gain the same kwargs contract (user_last_statement, historical_messages, system_prompt, model, temperature, max_new_tokens); calling them with neither payload nor user_last_statement now raises NoArgsAndNoPayloadError (previously TypeError). Partial kwargs (required field missing, or no resolvable model name) raise NoArgsAndNoPayloadError instead of a raw Pydantic ValidationError. Type aliases _ConvPayload / _ExtConvPayload removed; Complete numpydoc docstrings added to every endpoint method (previously missing on generative_answer and generate_article_from_text). Tests: new llm_router_lib/tests/test_client_unified_api.py (89 unit tests, python -m unittest); api tests for dict payloads rewritten to assert TypeError, plus new kwargs‑based conversation client tests. Style: unify all type annotations across llm_router_lib, llm_router_api and llm_router_cli to the typing module (Optional[...] instead of X | None, List/Dict/Tuple instead of the built-in list/dict/tuple generics, Union[...] for pure unions) so the codebase is consistent; runtime behaviour is unchanged. Docs: llm_router_lib/README.md and RESPONSE_MODELS.md updated to the keyword‑only calling contract (raw‑dict payloads rejected, previously missing methods documented).
0.9.2 Refactor and enhance codebase with type hints, tests, and fixes
0.9.3 Tests: Added unit tests for GPGAuditorLogStorage, AuthMiddleware, and PermissionEngine to cover encryption errors, IP whitelists, token budgets, and custom policies. Refactors: Removed auth key reveal functionality and updated related CLI commands and documentation to align with the enhanced key security model. Introduced LLM_ROUTER_TRUSTED_PROXIES and LLM_ROUTER_AUTH_FAILURE_LIMIT for better handling of proxy trust and authentication failures. Enhanced request authorization logic with a default-deny model, IP whitelists, token budgets, and more refined endpoint-level policies. Improved Lua scripts by ensuring explicit type coercion. Added SHA-256 reverse indexing for O(1) key lookups and improved key handling across memory, Redis, and Vault backends. Fixes: Enforced fatal errors on GPG encryption failures to ensure audit records' integrity. Modified auth_429_response to return an improved JSON error response containing a retry_after field for better clarity.
0.9.4 Feature: add Redis protocol (RESP2/RESP3) support across all Redis-backed modules: new env vars LLM_ROUTER_REDIS_PROTOCOL (default 3) and LLM_ROUTER_AUTH_REDIS_PROTOCOL (default 3); protocol= is now passed to every redis.Redis client (auth key store, shared client factory, rate limiter, load-balancing strategies, CLI auth command, scripts/check_redis.py); new CLI flag --auth-redis-protocol (2/3, default 2) on all llm-router auth subcommands with --store redis. Docs: new protocol env vars and CLI flag documented in ENV_DEFINITIONS.md, AUTHENTICATION.md, RATE_LIMITING.md and llm_router_cli/README.md. Refactor: move llm_router_api docs into llm_router_api/docs/. Changelog: bump version to 0.9.4 (.version, Helm chart appVersion).
0.9.5 Fix (correctness): automatic retry actually works now — non‑OK provider responses (429/500/502/503/504) are retried on a different provider with exponential backoff + jitter (capped, configurable policy in llm_router_api/endpoints/http_dispatch.py); when the retry budget is exhausted the provider's last status code is returned to the client instead of a masked 500; transport errors (connection refused/timeout) are retried the same way; replay safety documented (only reproducible JSON bodies are retried). Fix (correctness): client input errors (missing/invalid params, ValueError/pydantic ValidationError) now return 400 with {"error": {"message": …}} — the Flask registrar is the single owner of the exception→status mapping; real server‑side failures still return 500. Fix (correctness): provider health‑check now treats only 2xx/3xx as available (401/403 → auth_error, 404 → not_found, other 4xx/5xx diagnosed separately); ping path per api_type (vLLM /health, Ollama /api/version, OpenAI‑compatible /v1/models) with fallback probing and auth token forwarding; missing api_type no longer crashes the monitor; new hysteresis — a provider is marked down only after LLM_ROUTER_PROVIDER_MONITOR_MAX_CONSECUTIVE_FAILURES (default 2) consecutive failed pings, one success restores it (a busy host that is slow to accept a new connection but serves keep‑alive traffic no longer drops out of the pool); ping timeout configurable via LLM_ROUTER_PROVIDER_MONITOR_PING_TIMEOUT_SECONDS (default 5.0 s). Resilience: balanced/weighted load‑balancing counters moved to Redis (shared across gunicorn workers, atomic pick‑and‑increment, deterministic weighted sequence) with in‑memory fallback when Redis is unavailable (llm_router_api/core/lb/lb_counters.py). Lifecycle: FlaskEngine exposes an explicit, idempotent start()/stop() + context manager + atexit hook and a SIGTERM/SIGINT shutdown hook in rest_api — gunicorn workers no longer leak the services‑monitor thread (monitor threads are daemon=True). Data quality: /models no longer returns fabricated constants ("type":"vllm", "state":"not-loaded", "quantization":"4bit") — provider‑specific fields come from real per‑provider data or are null. Cleanup: removed dead LLM_ROUTER_AUTH_RATE_LIMIT_ENABLED switch and the broken, unregistered beta/AdaptiveStrategy; unified the Vault key‑store backend on hvac (was non‑installable hvault); setup.py no longer ships test sub‑packages in the wheel and pins the git dependencies to commits; single source of truth for the version (.version) enforced in CI against git tag and Helm appVersion; Helm deployment checksum annotation fixed to reference the real files/models-config.json; requirements*.txt pinned, [build-system] added to pyproject.toml, requirements.lock generated; print() and bare except: pass in production code replaced with logging or intentional comments; llm-router auth CLI rewritten on a single argparse pass (no argv re‑scanning) with working --store vault wiring. Tests: new suites for dispatch retry, 400‑input handling, provider‑monitor (status mapping, fallbacks, hysteresis), engine lifecycle, LB global counters, /models fields, hvac import and the auth CLI — 559 passed, 4 skipped. Changelog: bump version to 0.9.5 (.version, Helm chart appVersion).
0.9.6 Fix LMStudio /api/v0/model returns proper response format.
1.0.0-rc1 Chore Bump version to 1.0.0-rc1; Change python version from 3.13.9-slim-trixie to 3.14-slim-trixie, Fix: remove unused orig_params and reconnect_number from endpoint_i dispatcher. Docs: add ENDPOINT_DEV.md endpoint‑development guide (English version of endpoints/README-pl.md) and register it in site/docs.html; cross‑link it from root and llm_router_api READMEs.
1.0.0 Docs: add "Endpoint Development Guide" to llm_router_api/docs with comprehensive explanations for creating, configuring, and extending REST endpoints, remove old Polish version. Chore: remove unused Docker entrypoint script and Supervisord configuration, update Dockerfile to simplify CMD setup.
1.0.1 Chore: update _ENDPOINT_PERMISSION_MAP to adjust permissions for /api/ping and /api/version endpoints
1.0.2 Feat: add the util subcommands to llm-router — translate, genai-classifier and genai-data-augmentation. These utilities are ported over from the standalone llm-router-utils repository and now run directly through the LLMRouter client. Feat: add opt-in --verbose (DEBUG) logging across the util, auth and config commands, with secrets always masked (***) in the logs. Refactor: extract the shared setup_logging/shorten helpers into llm_router_cli/log_utils.py and reuse the --verbose flag via BaseCommand.add_verbose (no per-command duplication).
1.0.3 Features: Introduced persistent custom authentication policies with JSON support. Enabled file-based caching and improved policy management via CLI. Fixes: Resolved missing package error during __version__ lookup. Updated README with clarified flag descriptions and safer examples. Refactoring: Modularized provider discovery and configuration generation. Enhanced flexibility and error handling for CLI policy management. Improved record processing and added warnings for missing fields during tasks. Default prefix for the API Auth keys changed to sk-llmr-live. Testing: Increased test coverage for custom policy handling in CLI. Documentation: Added details on using custom auth policies and environment configurations like LLM_ROUTER_AUTH_CUSTOM_POLICIES_FILE.
1.0.4 Tests: Add comprehensive unit tests and coverage configuration. Added unit tests for: (1) LLMRouterClient, CLI commands and helpers; (2) llm_router_api.core.server helper functions; (3) llm_router_api.rest_api CLI arguments and server dispatch logic; (4) Prometheus metrics, auth metrics, and stream converters; (5) Strategies, caching mechanisms, and monitoring functionality; (6) Endpoint handling and API conversions; (7) Constants, error handling, and environment-specific behavior; (8) Multiple endpoints, including Health, Ping, and version utilities. Add Manifest.in and author_email in setup.py. Prepared to PyPi package.
llm-router · docs are generated from the repository by tools/build_docs.py 1.0.4-prod @ 677231d