| 0.0.1 |
Initialization, License, setup, interface for each endpoint and sample ping EP. Autoloader of builtin endpoints and for the future implementations. |
| 0.0.2 |
Add base models for api call (module llm_proxy_rest.data_models with error.py handling. Decorators to check required params and to measure the response time. |
| 0.0.3 |
Proper AutoLoading for each found endpoint. Implementation of ApiTypesDispatcher, ApiModelConfig, ModelHandler. Ollama endpoints: /, tags. Added endpoint to full proxy with params. Streaming in case when external api provides stream. |
| 0.0.4 |
All llama-service endpoints are refactored to llm-proxy-api. Refactoring base ep_run method. Proper handling system message, prompt name, model etc. |
| 0.1.0 |
Repository name changed from llm-proxy-api to llm-router. Added class HttpRequestExecutor to handle http requests from EndpointWithHttpRequestI. Handled routing between any models: openai -> ollama and ollama -> openai |
| 0.1.1 |
Prometheus metrics logging. Workers/Threads/Workers class is able to set by environments. Streaming fixes. Multi-providers for single model with default-balanced strategy. |
| 0.2.0 |
Add balancing strategies: balanced, weighted, dynamic_weighted and first_available which works for streaming and non streaming requests. Included Prometheus metrics logging via /metrics endpoint. First stage of llm_router_lib library, to simply usage of llm-router-api. |
| 0.2.1 |
Fix stream: OpenAI->Ollama, Ollama->OpenAI. Add Redis caching of availability of model providers (when using first_available strategy). Add llm_router_web module with simple flask-based frontend to manage llm-router config files. |
| 0.2.2 |
Update dockerfile and requirements. Fix routing with vLLM. |
| 0.2.3 |
New web configurator: Handling projects, configs for each user separately. First Available strategy is more powerful, a lot of improvements to efficiency. |
| 0.2.4 |
Anonymizer module, integration anonymization with any endpoint (using dynamic payload analysis and full payload anonymisation), dedicated /api/anonymize_text endpoint as memory only anonymization. Whole router may be run in FORCE_ANONYMISATION mode. |
| 0.3.0 |
Anonymization available with three strategies: fast_masker, genai, prov_masker. |
| 0.3.1 |
Refactoring lb.strategies to be more flexible modular. Introduced MaskerPipeline and GuardrailPipeline both configured via env. Removed genai-based masking endpoint. |
| 0.4.0 |
The main repository is divided into dedicated ones: plugins, services, web — separate repositories. Clean up the whole repository. Examples of integration with llamaindex, langchain, openai, litellm and haystack. |
| 0.4.1 |
Audit log is stored using GPG. Add bash script (scripts/gen_and_export_gpg.sh to prepare GPG keys and simple scripts/decrypt_auditor_logs.sh to decrypt encrypted audit logs. Moved core functionality from base to module core module. Quickstart. |
| 0.4.2 |
Fix first_available_optim Strategy. Add KeepAliveMonitor to periodically pings model endpoints to keep them warm. |
| 0.4.3 |
Add custom Prometheus metrices for logging masker/guardrail inidents. Fix OpenAI compatible v1 /models endpoint. Introduce monitors: services and keep alive models. Fixed guardrail retunr in case when streaming. |
| 0.4.4 |
Validate unique provider identifiers. Store all hosts with keep‑alive configured in a Redis. UtilsPlugin pipeline with LangChain based simple RAG plugin (extending context to GenAI with locally built databse). Add handling of v1/response endpoint |
| 0.4.5 |
Fixed sreaming to LMStudio native. Refactor streaming module. |
| 0.4.6 |
Added support for embeddings endpoints across all providers. Extended ApiModel and ApiTypesI with is_embedding flag. Added test_embeddings.py utility for verifying embedding models through the API. |
| 0.4.7 |
Integration with native Anthropic API. Add translate, generative_answer and ping methods to LLMRouterClient (with tests). Refactor LLMRouterCkientServices to use self.model_cls. Add payload converter for vLLM. |
| 0.5.0 |
Integration with PII masker, code refactoring |
| 0.5.1 |
Add /v1/messages endpoint (Claude Agent compatibility), streaming Cache‑Control/Pragma/Expires/Vary headers, mandatory Redis (runtime error on missing connection), non‑root Docker startup, remove ml‑utils dependency, add gnupg to requirements, update default plugin to simple_semantic_routing. |
| 0.5.2 |
Prevent network topology leak in error messages (full details remain in server‑side logs), example models-config.json no longer contains real internal IPs. Introduced LLM_ROUTER_MAX_REQUEST_BODY_SIZE to set the maximum content length. Sanitize all error messages returned to the Client. Local security. |
| 0.6.0 |
Authentication system: API key-based auth with multi-backend key stores (Memory, Redis, Vault), plaintext and secret-key lookup, enable/disable keys, seed-file persistence. Auth CLI: auth subcommands for managing API keys (create/list/enable/disable/delete) with formatted tabular output and prefix matching. Rate limiting: Per-key rate limiting via token bucket, predefined rate-limiting policies in rate_limiting-policies.json, PolicyEngine accepting dict key records. Anonymizer CLI: Migrated fast_masker to anonymizer CLI with deprecation warning; moved to masker subpackage. Infrastructure: Shared Redis client across stores and cache, dynamic column widths for CLI output, environment variable updates for auth/rate-limiting/audit logging config. |
| 0.6.1 |
Added config CLI command with discover (auto-discover local Ollama/vLLM/LM Studio providers) and merge (deep-merge multiple models-config.json files). |
| 0.6.2 |
CLI: fix _RATE_LIMIT_COMMANDS NameError in auth CLI (commit 2d9e593). Core: extract Prometheus multiprocess dir handling into MetricsHandler.prepare_multiproc_dir() with sane default; add LLM_ROUTER_AUTH_MEMORY_SEED_FILE env var for memory store seed path. CLI: restructure config commands, clean up imports across core modules, add llm_router_client.py test script with predefined model tests. |
| 0.6.3 |
Core: consolidate streaming mode flags into StreamConversion enum, simplify handler dispatch and fix swapped handler type hints in OpenAI/Ollama endpoints. Simplify error handling and response logic in endpoint_i.py. |
| 0.6.4 |
Refac: Clear REAME files. |
| 0.6.5 |
CLI discover: added KoboldCpp and TabbyAPI providers to config discover auto-discovery (now covers 6 local providers: Ollama, vLLM, LM Studio, llama.cpp, KoboldCpp, TabbyAPI). Infrastructure: fixed _scan_and_merge passing best_port as host parameter causing empty config; removed deprecated llm-router-fast-masker entry point and masker module. |
| 0.6.6 |
Auth CLI: restored working key generation by fixing _handle_key instance method call (cls._handle_key → cls()._handle_key). Config CLI: fixed _do_discover/_do_merge calls to use class methods properly (_do_discover(args) → cls._do_discover(args)). Removed all legacy backward-compatibility shim comments and module-level functions from CLI command modules. |
| 0.6.7 |
Docs: consolidate environment variable documentation into ENV_DEFINITIONS.md, simplify references across README files, add semantic_biencoder_routing env var details, remove LLM_ROUTER_HOST and update default MASKING_STRATEGY_PIPELINE. Chore: add /metrics endpoint to public endpoints list (LLM_ROUTER_AUTH_PUBLIC_ENDPOINTS) across auth and config modules; bump version to 0.6.7. |