| 0.0.1 |
Initialization, License, setup, interface for each endpoint and sample ping EP. Autoloader of builtin endpoints and for the future implementations. |
| 0.0.2 |
Add base models for api call (module llm_proxy_rest.data_models with error.py handling. Decorators to check required params and to measure the response time. |
| 0.0.3 |
Proper AutoLoading for each found endpoint. Implementation of ApiTypesDispatcher, ApiModelConfig, ModelHandler. Ollama endpoints: /, tags. Added endpoint to full proxy with params. Streaming in case when external api provides stream. |
| 0.0.4 |
All llama-service endpoints are refactored to llm-proxy-api. Refactoring base ep_run method. Proper handling system message, prompt name, model etc. |
| 0.1.0 |
Repository name changed from llm-proxy-api to llm-router. Added class HttpRequestExecutor to handle http requests from EndpointWithHttpRequestI. Handled routing between any models: openai -> ollama and ollama -> openai |
| 0.1.1 |
Prometheus metrics logging. Workers/Threads/Workers class is able to set by environments. Streaming fixes. Multi-providers for single model with default-balanced strategy. |
| 0.2.0 |
Add balancing strategies: balanced, weighted, dynamic_weighted and first_available which works for streaming and non streaming requests. Included Prometheus metrics logging via /metrics endpoint. First stage of llm_router_lib library, to simply usage of llm-router-api. |
| 0.2.1 |
Fix stream: OpenAI->Ollama, Ollama->OpenAI. Add Redis caching of availability of model providers (when using first_available strategy). Add llm_router_web module with simple flask-based frontend to manage llm-router config files. |
| 0.2.2 |
Update dockerfile and requirements. Fix routing with vLLM. |
| 0.2.3 |
New web configurator: Handling projects, configs for each user separately. First Available strategy is more powerful, a lot of improvements to efficiency. |
| 0.2.4 |
Anonymizer module, integration anonymization with any endpoint (using dynamic payload analysis and full payload anonymisation), dedicated /api/anonymize_text endpoint as memory only anonymization. Whole router may be run in FORCE_ANONYMISATION mode. |
| 0.3.0 |
Anonymization available with three strategies: fast_masker, genai, prov_masker. |
| 0.3.1 |
Refactoring lb.strategies to be more flexible modular. Introduced MaskerPipeline and GuardrailPipeline both configured via env. Removed genai-based masking endpoint. |
| 0.4.0 |
The main repository is divided into dedicated ones: plugins, services, web — separate repositories. Clean up the whole repository. Examples of integration with llamaindex, langchain, openai, litellm and haystack. |
| 0.4.1 |
Audit log is stored using GPG. Add bash script (scripts/gen_and_export_gpg.sh to prepare GPG keys and simple scripts/decrypt_auditor_logs.sh to decrypt encrypted audit logs. Moved core functionality from base to module core module. Quickstart. |
| 0.4.2 |
Fix first_available_optim Strategy. Add KeepAliveMonitor to periodically pings model endpoints to keep them warm. |
| 0.4.3 |
Add custom Prometheus metrices for logging masker/guardrail inidents. Fix OpenAI compatible v1 /models endpoint. Introduce monitors: services and keep alive models. Fixed guardrail retunr in case when streaming. |
| 0.4.4 |
Validate unique provider identifiers. Store all hosts with keep‑alive configured in a Redis. UtilsPlugin pipeline with LangChain based simple RAG plugin (extending context to GenAI with locally built databse). Add handling of v1/response endpoint |
| 0.4.5 |
Fixed sreaming to LMStudio native. Refactor streaming module. |
| 0.4.6 |
Added support for embeddings endpoints across all providers. Extended ApiModel and ApiTypesI with is_embedding flag. Added test_embeddings.py utility for verifying embedding models through the API. |
| 0.4.7 |
Integration with native Anthropic API. Add translate, generative_answer and ping methods to LLMRouterClient (with tests). Refactor LLMRouterCkientServices to use self.model_cls. Add payload converter for vLLM. |
| 0.5.0 |
Integration with PII masker, code refactoring |
| 0.5.1 |
Add /v1/messages endpoint (Claude Agent compatibility), streaming Cache‑Control/Pragma/Expires/Vary headers, mandatory Redis (runtime error on missing connection), non‑root Docker startup, remove ml‑utils dependency, add gnupg to requirements, update default plugin to simple_semantic_routing. |
| 0.5.2 |
Prevent network topology leak in error messages (full details remain in server‑side logs), example models-config.json no longer contains real internal IPs. Introduced LLM_ROUTER_MAX_REQUEST_BODY_SIZE to set the maximum content length. Sanitize all error messages returned to the Client. Local security. |
| 0.6.0 |
Authentication system: API key-based auth with multi-backend key stores (Memory, Redis, Vault), plaintext and secret-key lookup, enable/disable keys, seed-file persistence. Auth CLI: auth subcommands for managing API keys (create/list/enable/disable/delete) with formatted tabular output and prefix matching. Rate limiting: Per-key rate limiting via token bucket, predefined rate-limiting policies in rate_limiting-policies.json, PolicyEngine accepting dict key records. Anonymizer CLI: Migrated fast_masker to anonymizer CLI with deprecation warning; moved to masker subpackage. Infrastructure: Shared Redis client across stores and cache, dynamic column widths for CLI output, environment variable updates for auth/rate-limiting/audit logging config. |
| 0.6.1 |
Added config CLI command with discover (auto-discover local Ollama/vLLM/LM Studio providers) and merge (deep-merge multiple models-config.json files). |
| 0.6.2 |
CLI: fix _RATE_LIMIT_COMMANDS NameError in auth CLI (commit 2d9e593). Core: extract Prometheus multiprocess dir handling into MetricsHandler.prepare_multiproc_dir() with sane default; add LLM_ROUTER_AUTH_MEMORY_SEED_FILE env var for memory store seed path. CLI: restructure config commands, clean up imports across core modules, add llm_router_client.py test script with predefined model tests. |
| 0.6.3 |
Core: consolidate streaming mode flags into StreamConversion enum, simplify handler dispatch and fix swapped handler type hints in OpenAI/Ollama endpoints. Simplify error handling and response logic in endpoint_i.py. |
| 0.6.4 |
Refac: Clear REAME files. |
| 0.6.5 |
CLI discover: added KoboldCpp and TabbyAPI providers to config discover auto-discovery (now covers 6 local providers: Ollama, vLLM, LM Studio, llama.cpp, KoboldCpp, TabbyAPI). Infrastructure: fixed _scan_and_merge passing best_port as host parameter causing empty config; removed deprecated llm-router-fast-masker entry point and masker module. |
| 0.6.6 |
Auth CLI: restored working key generation by fixing _handle_key instance method call (cls._handle_key → cls()._handle_key). Config CLI: fixed _do_discover/_do_merge calls to use class methods properly (_do_discover(args) → cls._do_discover(args)). Removed all legacy backward-compatibility shim comments and module-level functions from CLI command modules. |
| 0.6.7 |
Docs: consolidate environment variable documentation into ENV_DEFINITIONS.md, simplify references across README files, add semantic_biencoder_routing env var details, remove LLM_ROUTER_HOST and update default MASKING_STRATEGY_PIPELINE. Chore: add /metrics endpoint to public endpoints list (LLM_ROUTER_AUTH_PUBLIC_ENDPOINTS) across auth and config modules; bump version to 0.6.7. |
| 0.6.8 |
Feature: Prometheus router metrics, Grafana dashboard, and /metrics public endpoint for improved observability and visualization. Docs: consolidate env var docs into ENV_DEFINITIONS.md, update docstrings and line wrapping across modules. Chore: expand mypy checks, add types-setuptools. |
| 0.6.9 |
Chore: update prompts to be more specific. Feature: add prepare_response_function to endpoint_i, add GenerateNewsFromTextService, add default_model parameter to Client to provide fallback model configuration |
| 0.7.0 |
Feature: add polarity_3c endpoint (/api/polarity_3c) to detect 3-class text polarity (ambivalent, positive, negative) with PL/EN system prompts. Lib: add Polarity3cModel, Polarity3cService, and LLMRouterClient.polarity_3c() method to llm_router_lib. Auth & Docs: register endpoint with builtin policy permission, update endpoint documentation and add comprehensive test suite. |
| 0.8.0 |
Feature: add endpoints: /api/simplify_text, /api/generate_questions, /api/create_full_article_from_texts, /api/generate_article_from_texts, /api/generate_label with models, services and matching LLMRouterClient methods. Auth: register the new endpoints with the builtin policy permission. Fix _ensure_alternating_roles in endpoint_i to build a valid payload by merging consecutive messages of the same role (joining their contents), folding all system messages into one leading message, and prepending/appending empty user placeholders when the dialogue starts with or ends with an assistant turn. |
| 0.9.0 |
Refactor (dependency direction): llm_router_lib no longer imports from llm_router_api. The shared DEFAULT_EP_LANGUAGE constant (and the LLM_ROUTER_ env prefix) now live in the dependency‑free llm_router_lib.core.constants; llm_router_api.base.constants_base re‑exports them for backward compatibility. BREAKING – structural rename anchored on the endpoint URL (HTTP paths are unchanged). Canonical name = URL without /api/; class/model/service/client‑method/constant names now follow it and legacy aliases are removed: TranslateTexts→Translate, TranslateTextModel→TranslateModel, TranslateTextService→TranslateService; SimplifyTexts→SimplifyText, client simplify_texts()→simplify_text(); GenerateQuestionsFromTexts→GenerateQuestions, GenerateQuestionFromTextsModel→GenerateQuestionsModel, GenerateQuestionsFromTextsService→GenerateQuestionsService (alias GenerateQuestionFromTextsService removed), client generate_questions_from_texts()→generate_questions(), constants GENERATE_Q_*→GENERATE_QUESTIONS_*; GenerateNewsFromTextHandler→GenerateArticleFromText, GenerateNewsFromTextService→GenerateArticleFromTextService, client generate_news_from_text()→generate_article_from_text(), constants GENERATE_ART_*→GENERATE_ARTICLE_FROM_TEXT_*; GenerateArticleFromTexts constants GENERATE_ARTICLES_*→GENERATE_ARTICLE_FROM_TEXTS_*; FullArticleFromTexts→CreateFullArticleFromTexts, CreateArticleFromNewsListModel→CreateFullArticleFromTextsModel, constants FULL_ARTICLE_*→CREATE_FULL_ARTICLE_FROM_TEXTS_*; AnswerBasedOnTheContext→GenerativeAnswer, AnswerBasedOnTheContextModel→GenerativeAnswerModel, constants CONTEXT_ANSWER_*→GENERATIVE_ANSWER_*; GenerativeConversationModel→ConversationWithModelRequest, ExtendedGenerativeConversationModel→ExtendedConversationWithModelRequest, ConversationService→ConversationWithModelService, ExtendedConversationService→ExtendedConversationWithModelService. Class inheritance hierarchy and the auth _ENDPOINT_PERMISSION_MAP (URL‑keyed) are unchanged. Refactor (structure only, no behavior change): shared base TextListUtilityEndpoint for the texts utility endpoints (common prepare_payload — one user message per source text, build_map_prompt hook, per‑text _prepare_response helper); message normalisation moved verbatim into pure functions in llm_router_api/endpoints/message_normalizer.py (thin delegating methods kept on the endpoint class); outbound HTTP dispatch + retry orchestration moved verbatim into llm_router_api/endpoints/http_dispatch.py (HttpDispatch; EndpointWithHttpRequestI.RetryResponse kept as a backward‑compatible alias of http_dispatch.RetryPolicy); run_ep slimmed into a readable orchestrator with extracted sequence helpers (endpoint_i.py reduced from 2123 to ~1810 lines). Fix (behavior): per‑text endpoints now actually return one result per source text — the role normalizer (introduced in 0.8.0 to build valid chat payloads) was applied unconditionally in run_ep and merged the consecutive user messages built by the texts endpoints, so polarity_3c, translate, simplify_text and generate_questions silently returned one merged result; run_ep now skips role normalisation for call_for_each_user_msg=True endpoints (the HTTP executor already dispatches each text as its own [system, user] call), while conversation endpoints (conversation_with_model, extended_conversation_with_model) keep it; client code consuming a single merged entry for those four endpoints must iterate the response list. Tests: regression grid extended to all remaining builtin endpoints; new test_http_dispatch.py (dispatch/retry, late‑binding, RetryResponse alias) and end‑to‑end test_per_text_dispatch.py (per‑text ⇒ one model call per text ⇒ results paired per text; chat endpoint ⇒ consecutive user messages still merged). Feature: Add log rotation. Feature: add method models() to client.py, add public /health endpoint |
| 0.9.1 |
BREAKING – unified LLMRouterClient calling contract (pre‑1.0): all 12 endpoint methods (conversation_with_model, extended_conversation_with_model, polarity_3c, translate, simplify_text, generative_answer, generate_article_from_text, generate_article_from_texts, create_full_article_from_texts, generate_questions, generate_label, …) now share one keyword‑only signature: payload (a Pydantic request model instance) or named domain arguments + model + optional generation options; argument order is fixed (payload → domain fields → model → temperature → max_new_tokens). Raw dict payloads are removed – passing a dict as payload raises TypeError (build the matching model explicitly, e.g. MyModel(**payload)). Hard‑coded defaults in signatures (0.2/0.75/256/512/1024/64, number_of_questions=1) are removed – generation defaults come exclusively from the Pydantic models (GenerativeOptions); optional parameters default to None so None‑valued kwargs no longer break Pydantic validation (fixes the texts=None bug). conversation_with_model / extended_conversation_with_model gain the same kwargs contract (user_last_statement, historical_messages, system_prompt, model, temperature, max_new_tokens); calling them with neither payload nor user_last_statement now raises NoArgsAndNoPayloadError (previously TypeError). Partial kwargs (required field missing, or no resolvable model name) raise NoArgsAndNoPayloadError instead of a raw Pydantic ValidationError. Type aliases _ConvPayload / _ExtConvPayload removed; Complete numpydoc docstrings added to every endpoint method (previously missing on generative_answer and generate_article_from_text). Tests: new llm_router_lib/tests/test_client_unified_api.py (89 unit tests, python -m unittest); api tests for dict payloads rewritten to assert TypeError, plus new kwargs‑based conversation client tests. Style: unify all type annotations across llm_router_lib, llm_router_api and llm_router_cli to the typing module (Optional[...] instead of X | None, List/Dict/Tuple instead of the built-in list/dict/tuple generics, Union[...] for pure unions) so the codebase is consistent; runtime behaviour is unchanged. Docs: llm_router_lib/README.md and RESPONSE_MODELS.md updated to the keyword‑only calling contract (raw‑dict payloads rejected, previously missing methods documented). |
| 0.9.2 |
Refactor and enhance codebase with type hints, tests, and fixes |
| 0.9.3 |
Tests: Added unit tests for GPGAuditorLogStorage, AuthMiddleware, and PermissionEngine to cover encryption errors, IP whitelists, token budgets, and custom policies. Refactors: Removed auth key reveal functionality and updated related CLI commands and documentation to align with the enhanced key security model. Introduced LLM_ROUTER_TRUSTED_PROXIES and LLM_ROUTER_AUTH_FAILURE_LIMIT for better handling of proxy trust and authentication failures. Enhanced request authorization logic with a default-deny model, IP whitelists, token budgets, and more refined endpoint-level policies. Improved Lua scripts by ensuring explicit type coercion. Added SHA-256 reverse indexing for O(1) key lookups and improved key handling across memory, Redis, and Vault backends. Fixes: Enforced fatal errors on GPG encryption failures to ensure audit records' integrity. Modified auth_429_response to return an improved JSON error response containing a retry_after field for better clarity. |
| 0.9.4 |
Feature: add Redis protocol (RESP2/RESP3) support across all Redis-backed modules: new env vars LLM_ROUTER_REDIS_PROTOCOL (default 3) and LLM_ROUTER_AUTH_REDIS_PROTOCOL (default 3); protocol= is now passed to every redis.Redis client (auth key store, shared client factory, rate limiter, load-balancing strategies, CLI auth command, scripts/check_redis.py); new CLI flag --auth-redis-protocol (2/3, default 2) on all llm-router auth subcommands with --store redis. Docs: new protocol env vars and CLI flag documented in ENV_DEFINITIONS.md, AUTHENTICATION.md, RATE_LIMITING.md and llm_router_cli/README.md. Refactor: move llm_router_api docs into llm_router_api/docs/. Changelog: bump version to 0.9.4 (.version, Helm chart appVersion). |
| 0.9.5 |
Fix (correctness): automatic retry actually works now — non‑OK provider responses (429/500/502/503/504) are retried on a different provider with exponential backoff + jitter (capped, configurable policy in llm_router_api/endpoints/http_dispatch.py); when the retry budget is exhausted the provider's last status code is returned to the client instead of a masked 500; transport errors (connection refused/timeout) are retried the same way; replay safety documented (only reproducible JSON bodies are retried). Fix (correctness): client input errors (missing/invalid params, ValueError/pydantic ValidationError) now return 400 with {"error": {"message": …}} — the Flask registrar is the single owner of the exception→status mapping; real server‑side failures still return 500. Fix (correctness): provider health‑check now treats only 2xx/3xx as available (401/403 → auth_error, 404 → not_found, other 4xx/5xx diagnosed separately); ping path per api_type (vLLM /health, Ollama /api/version, OpenAI‑compatible /v1/models) with fallback probing and auth token forwarding; missing api_type no longer crashes the monitor; new hysteresis — a provider is marked down only after LLM_ROUTER_PROVIDER_MONITOR_MAX_CONSECUTIVE_FAILURES (default 2) consecutive failed pings, one success restores it (a busy host that is slow to accept a new connection but serves keep‑alive traffic no longer drops out of the pool); ping timeout configurable via LLM_ROUTER_PROVIDER_MONITOR_PING_TIMEOUT_SECONDS (default 5.0 s). Resilience: balanced/weighted load‑balancing counters moved to Redis (shared across gunicorn workers, atomic pick‑and‑increment, deterministic weighted sequence) with in‑memory fallback when Redis is unavailable (llm_router_api/core/lb/lb_counters.py). Lifecycle: FlaskEngine exposes an explicit, idempotent start()/stop() + context manager + atexit hook and a SIGTERM/SIGINT shutdown hook in rest_api — gunicorn workers no longer leak the services‑monitor thread (monitor threads are daemon=True). Data quality: /models no longer returns fabricated constants ("type":"vllm", "state":"not-loaded", "quantization":"4bit") — provider‑specific fields come from real per‑provider data or are null. Cleanup: removed dead LLM_ROUTER_AUTH_RATE_LIMIT_ENABLED switch and the broken, unregistered beta/AdaptiveStrategy; unified the Vault key‑store backend on hvac (was non‑installable hvault); setup.py no longer ships test sub‑packages in the wheel and pins the git dependencies to commits; single source of truth for the version (.version) enforced in CI against git tag and Helm appVersion; Helm deployment checksum annotation fixed to reference the real files/models-config.json; requirements*.txt pinned, [build-system] added to pyproject.toml, requirements.lock generated; print() and bare except: pass in production code replaced with logging or intentional comments; llm-router auth CLI rewritten on a single argparse pass (no argv re‑scanning) with working --store vault wiring. Tests: new suites for dispatch retry, 400‑input handling, provider‑monitor (status mapping, fallbacks, hysteresis), engine lifecycle, LB global counters, /models fields, hvac import and the auth CLI — 559 passed, 4 skipped. Changelog: bump version to 0.9.5 (.version, Helm chart appVersion). |