| 0.0.1 |
Initialization, License, setup, interface for each endpoint and sample ping EP. Autoloader of builtin endpoints and for the future implementations. |
| 0.0.2 |
Add base models for api call (module llm_proxy_rest.data_models with error.py handling. Decorators to check required params and to measure the response time. |
| 0.0.3 |
Proper AutoLoading for each found endpoint. Implementation of ApiTypesDispatcher, ApiModelConfig, ModelHandler. Ollama endpoints: /, tags. Added endpoint to full proxy with params. Streaming in case when external api provides stream. |
| 0.0.4 |
All llama-service endpoints are refactored to llm-proxy-api. Refactoring base ep_run method. Proper handling system message, prompt name, model etc. |
| 0.1.0 |
Repository name changed from llm-proxy-api to llm-router. Added class HttpRequestExecutor to handle http requests from EndpointWithHttpRequestI. Handled routing between any models: openai -> ollama and ollama -> openai |
| 0.1.1 |
Prometheus metrics logging. Workers/Threads/Workers class is able to set by environments. Streaming fixes. Multi-providers for single model with default-balanced strategy. |
| 0.2.0 |
Add balancing strategies: balanced, weighted, dynamic_weighted and first_available which works for streaming and non streaming requests. Included Prometheus metrics logging via /metrics endpoint. First stage of llm_router_lib library, to simply usage of llm-router-api. |
| 0.2.1 |
Fix stream: OpenAI->Ollama, Ollama->OpenAI. Add Redis caching of availability of model providers (when using first_available strategy). Add llm_router_web module with simple flask-based frontend to manage llm-router config files. |
| 0.2.2 |
Update dockerfile and requirements. Fix routing with vLLM. |
| 0.2.3 |
New web configurator: Handling projects, configs for each user separately. First Available strategy is more powerful, a lot of improvements to efficiency. |
| 0.2.4 |
Anonymizer module, integration anonymization with any endpoint (using dynamic payload analysis and full payload anonymisation), dedicated /api/anonymize_text endpoint as memory only anonymization. Whole router may be run in FORCE_ANONYMISATION mode. |
| 0.3.0 |
Anonymization available with three strategies: fast_masker, genai, prov_masker. |
| 0.3.1 |
Refactoring lb.strategies to be more flexible modular. Introduced MaskerPipeline and GuardrailPipeline both configured via env. Removed genai-based masking endpoint. |
| 0.4.0 |
The main repository is divided into dedicated ones: plugins, services, web — separate repositories. Clean up the whole repository. Examples of integration with llamaindex, langchain, openai, litellm and haystack. |
| 0.4.1 |
Audit log is stored using GPG. Add bash script (scripts/gen_and_export_gpg.sh to prepare GPG keys and simple scripts/decrypt_auditor_logs.sh to decrypt encrypted audit logs. Moved core functionality from base to module core module. Quickstart. |
| 0.4.2 |
Fix first_available_optim Strategy. Add KeepAliveMonitor to periodically pings model endpoints to keep them warm. |
| 0.4.3 |
Add custom Prometheus metrices for logging masker/guardrail inidents. Fix OpenAI compatible v1 /models endpoint. Introduce monitors: services and keep alive models. Fixed guardrail retunr in case when streaming. |
| 0.4.4 |
Validate unique provider identifiers. Store all hosts with keep‑alive configured in a Redis. UtilsPlugin pipeline with LangChain based simple RAG plugin (extending context to GenAI with locally built databse). Add handling of v1/response endpoint |
| 0.4.5 |
Fixed sreaming to LMStudio native. Refactor streaming module. |
| 0.4.6 |
Added support for embeddings endpoints across all providers. Extended ApiModel and ApiTypesI with is_embedding flag. Added test_embeddings.py utility for verifying embedding models through the API. |
| 0.4.7 |
Integration with native Anthropic API. Add translate, generative_answer and ping methods to LLMRouterClient (with tests). Refactor LLMRouterCkientServices to use self.model_cls. Add payload converter for vLLM. |
| 0.5.0 |
Integration with PII masker, code refactoring |
| 0.5.1 |
Add /v1/messages endpoint (Claude Agent compatibility), streaming Cache‑Control/Pragma/Expires/Vary headers, mandatory Redis (runtime error on missing connection), non‑root Docker startup, remove ml‑utils dependency, add gnupg to requirements, update default plugin to simple_semantic_routing. |
| 0.5.2 |
Prevent network topology leak in error messages (full details remain in server‑side logs), example models-config.json no longer contains real internal IPs. Introduced LLM_ROUTER_MAX_REQUEST_BODY_SIZE to set the maximum content length. Sanitize all error messages returned to the Client. Local security. |
| 0.6.0 |
Authentication system: API key-based auth with multi-backend key stores (Memory, Redis, Vault), plaintext and secret-key lookup, enable/disable keys, seed-file persistence. Auth CLI: auth subcommands for managing API keys (create/list/enable/disable/delete) with formatted tabular output and prefix matching. Rate limiting: Per-key rate limiting via token bucket, predefined rate-limiting policies in rate_limiting-policies.json, PolicyEngine accepting dict key records. Anonymizer CLI: Migrated fast_masker to anonymizer CLI with deprecation warning; moved to masker subpackage. Infrastructure: Shared Redis client across stores and cache, dynamic column widths for CLI output, environment variable updates for auth/rate-limiting/audit logging config. |
| 0.6.1 |
Added config CLI command with discover (auto-discover local Ollama/vLLM/LM Studio providers) and merge (deep-merge multiple models-config.json files). |
| 0.6.2 |
CLI: fix _RATE_LIMIT_COMMANDS NameError in auth CLI (commit 2d9e593). Core: extract Prometheus multiprocess dir handling into MetricsHandler.prepare_multiproc_dir() with sane default; add LLM_ROUTER_AUTH_MEMORY_SEED_FILE env var for memory store seed path. CLI: restructure config commands, clean up imports across core modules, add llm_router_client.py test script with predefined model tests. |
| 0.6.3 |
Core: consolidate streaming mode flags into StreamConversion enum, simplify handler dispatch and fix swapped handler type hints in OpenAI/Ollama endpoints. Simplify error handling and response logic in endpoint_i.py. |
| 0.6.4 |
Refac: Clear REAME files. |
| 0.6.5 |
CLI discover: added KoboldCpp and TabbyAPI providers to config discover auto-discovery (now covers 6 local providers: Ollama, vLLM, LM Studio, llama.cpp, KoboldCpp, TabbyAPI). Infrastructure: fixed _scan_and_merge passing best_port as host parameter causing empty config; removed deprecated llm-router-fast-masker entry point and masker module. |
| 0.6.6 |
Auth CLI: restored working key generation by fixing _handle_key instance method call (cls._handle_key → cls()._handle_key). Config CLI: fixed _do_discover/_do_merge calls to use class methods properly (_do_discover(args) → cls._do_discover(args)). Removed all legacy backward-compatibility shim comments and module-level functions from CLI command modules. |
| 0.6.7 |
Docs: consolidate environment variable documentation into ENV_DEFINITIONS.md, simplify references across README files, add semantic_biencoder_routing env var details, remove LLM_ROUTER_HOST and update default MASKING_STRATEGY_PIPELINE. Chore: add /metrics endpoint to public endpoints list (LLM_ROUTER_AUTH_PUBLIC_ENDPOINTS) across auth and config modules; bump version to 0.6.7. |
| 0.6.8 |
Feature: Prometheus router metrics, Grafana dashboard, and /metrics public endpoint for improved observability and visualization. Docs: consolidate env var docs into ENV_DEFINITIONS.md, update docstrings and line wrapping across modules. Chore: expand mypy checks, add types-setuptools. |
| 0.6.9 |
Chore: update prompts to be more specific. Feature: add prepare_response_function to endpoint_i, add GenerateNewsFromTextService, add default_model parameter to Client to provide fallback model configuration |
| 0.7.0 |
Feature: add polarity_3c endpoint (/api/polarity_3c) to detect 3-class text polarity (ambivalent, positive, negative) with PL/EN system prompts. Lib: add Polarity3cModel, Polarity3cService, and LLMRouterClient.polarity_3c() method to llm_router_lib. Auth & Docs: register endpoint with builtin policy permission, update endpoint documentation and add comprehensive test suite. |
| 0.8.0 |
Feature: add endpoints: /api/simplify_text, /api/generate_questions, /api/create_full_article_from_texts, /api/generate_article_from_texts, /api/generate_label with models, services and matching LLMRouterClient methods. Auth: register the new endpoints with the builtin policy permission. Fix _ensure_alternating_roles in endpoint_i to build a valid payload by merging consecutive messages of the same role (joining their contents), folding all system messages into one leading message, and prepending/appending empty user placeholders when the dialogue starts with or ends with an assistant turn. |
| 0.9.0 |
Refactor (dependency direction): llm_router_lib no longer imports from llm_router_api. The shared DEFAULT_EP_LANGUAGE constant (and the LLM_ROUTER_ env prefix) now live in the dependency‑free llm_router_lib.core.constants; llm_router_api.base.constants_base re‑exports them for backward compatibility. BREAKING – structural rename anchored on the endpoint URL (HTTP paths are unchanged). Canonical name = URL without /api/; class/model/service/client‑method/constant names now follow it and legacy aliases are removed: TranslateTexts→Translate, TranslateTextModel→TranslateModel, TranslateTextService→TranslateService; SimplifyTexts→SimplifyText, client simplify_texts()→simplify_text(); GenerateQuestionsFromTexts→GenerateQuestions, GenerateQuestionFromTextsModel→GenerateQuestionsModel, GenerateQuestionsFromTextsService→GenerateQuestionsService (alias GenerateQuestionFromTextsService removed), client generate_questions_from_texts()→generate_questions(), constants GENERATE_Q_*→GENERATE_QUESTIONS_*; GenerateNewsFromTextHandler→GenerateArticleFromText, GenerateNewsFromTextService→GenerateArticleFromTextService, client generate_news_from_text()→generate_article_from_text(), constants GENERATE_ART_*→GENERATE_ARTICLE_FROM_TEXT_*; GenerateArticleFromTexts constants GENERATE_ARTICLES_*→GENERATE_ARTICLE_FROM_TEXTS_*; FullArticleFromTexts→CreateFullArticleFromTexts, CreateArticleFromNewsListModel→CreateFullArticleFromTextsModel, constants FULL_ARTICLE_*→CREATE_FULL_ARTICLE_FROM_TEXTS_*; AnswerBasedOnTheContext→GenerativeAnswer, AnswerBasedOnTheContextModel→GenerativeAnswerModel, constants CONTEXT_ANSWER_*→GENERATIVE_ANSWER_*; GenerativeConversationModel→ConversationWithModelRequest, ExtendedGenerativeConversationModel→ExtendedConversationWithModelRequest, ConversationService→ConversationWithModelService, ExtendedConversationService→ExtendedConversationWithModelService. Class inheritance hierarchy and the auth _ENDPOINT_PERMISSION_MAP (URL‑keyed) are unchanged. Refactor (structure only, no behavior change): shared base TextListUtilityEndpoint for the texts utility endpoints (common prepare_payload — one user message per source text, build_map_prompt hook, per‑text _prepare_response helper); message normalisation moved verbatim into pure functions in llm_router_api/endpoints/message_normalizer.py (thin delegating methods kept on the endpoint class); outbound HTTP dispatch + retry orchestration moved verbatim into llm_router_api/endpoints/http_dispatch.py (HttpDispatch; EndpointWithHttpRequestI.RetryResponse kept as a backward‑compatible alias of http_dispatch.RetryPolicy); run_ep slimmed into a readable orchestrator with extracted sequence helpers (endpoint_i.py reduced from 2123 to ~1810 lines). Fix (behavior): per‑text endpoints now actually return one result per source text — the role normalizer (introduced in 0.8.0 to build valid chat payloads) was applied unconditionally in run_ep and merged the consecutive user messages built by the texts endpoints, so polarity_3c, translate, simplify_text and generate_questions silently returned one merged result; run_ep now skips role normalisation for call_for_each_user_msg=True endpoints (the HTTP executor already dispatches each text as its own [system, user] call), while conversation endpoints (conversation_with_model, extended_conversation_with_model) keep it; client code consuming a single merged entry for those four endpoints must iterate the response list. Tests: regression grid extended to all remaining builtin endpoints; new test_http_dispatch.py (dispatch/retry, late‑binding, RetryResponse alias) and end‑to‑end test_per_text_dispatch.py (per‑text ⇒ one model call per text ⇒ results paired per text; chat endpoint ⇒ consecutive user messages still merged). Feature: Add log rotation. Feature: add method models() to client.py, add public /health endpoint |
| 0.9.1 |
BREAKING – unified LLMRouterClient calling contract (pre‑1.0): all 12 endpoint methods (conversation_with_model, extended_conversation_with_model, polarity_3c, translate, simplify_text, generative_answer, generate_article_from_text, generate_article_from_texts, create_full_article_from_texts, generate_questions, generate_label, …) now share one keyword‑only signature: payload (a Pydantic request model instance) or named domain arguments + model + optional generation options; argument order is fixed (payload → domain fields → model → temperature → max_new_tokens). Raw dict payloads are removed – passing a dict as payload raises TypeError (build the matching model explicitly, e.g. MyModel(**payload)). Hard‑coded defaults in signatures (0.2/0.75/256/512/1024/64, number_of_questions=1) are removed – generation defaults come exclusively from the Pydantic models (GenerativeOptions); optional parameters default to None so None‑valued kwargs no longer break Pydantic validation (fixes the texts=None bug). conversation_with_model / extended_conversation_with_model gain the same kwargs contract (user_last_statement, historical_messages, system_prompt, model, temperature, max_new_tokens); calling them with neither payload nor user_last_statement now raises NoArgsAndNoPayloadError (previously TypeError). Partial kwargs (required field missing, or no resolvable model name) raise NoArgsAndNoPayloadError instead of a raw Pydantic ValidationError. Type aliases _ConvPayload / _ExtConvPayload removed; Complete numpydoc docstrings added to every endpoint method (previously missing on generative_answer and generate_article_from_text). Tests: new llm_router_lib/tests/test_client_unified_api.py (89 unit tests, python -m unittest); api tests for dict payloads rewritten to assert TypeError, plus new kwargs‑based conversation client tests. Style: unify all type annotations across llm_router_lib, llm_router_api and llm_router_cli to the typing module (Optional[...] instead of X \| None, List/Dict/Tuple instead of the built-in list/dict/tuple generics, Union[...] for pure unions) so the codebase is consistent; runtime behaviour is unchanged. Docs: llm_router_lib/README.md and RESPONSE_MODELS.md updated to the keyword‑only calling contract (raw‑dict payloads rejected, previously missing methods documented). |
| 0.9.2 |
Refactor and enhance codebase with type hints, tests, and fixes |
| 0.9.3 |
Tests: Added unit tests for GPGAuditorLogStorage, AuthMiddleware, and PermissionEngine to cover encryption errors, IP whitelists, token budgets, and custom policies. Refactors: Removed auth key reveal functionality and updated related CLI commands and documentation to align with the enhanced key security model. Introduced LLM_ROUTER_TRUSTED_PROXIES and LLM_ROUTER_AUTH_FAILURE_LIMIT for better handling of proxy trust and authentication failures. Enhanced request authorization logic with a default-deny model, IP whitelists, token budgets, and more refined endpoint-level policies. Improved Lua scripts by ensuring explicit type coercion. Added SHA-256 reverse indexing for O(1) key lookups and improved key handling across memory, Redis, and Vault backends. Fixes: Enforced fatal errors on GPG encryption failures to ensure audit records' integrity. Modified auth_429_response to return an improved JSON error response containing a retry_after field for better clarity. |
| 0.9.4 |
Feature: add Redis protocol (RESP2/RESP3) support across all Redis-backed modules: new env vars LLM_ROUTER_REDIS_PROTOCOL (default 3) and LLM_ROUTER_AUTH_REDIS_PROTOCOL (default 3); protocol= is now passed to every redis.Redis client (auth key store, shared client factory, rate limiter, load-balancing strategies, CLI auth command, scripts/check_redis.py); new CLI flag --auth-redis-protocol (2/3, default 2) on all llm-router auth subcommands with --store redis. Docs: new protocol env vars and CLI flag documented in ENV_DEFINITIONS.md, AUTHENTICATION.md, RATE_LIMITING.md and llm_router_cli/README.md. Refactor: move llm_router_api docs into llm_router_api/docs/. Changelog: bump version to 0.9.4 (.version, Helm chart appVersion). |
| 0.9.5 |
Fix (correctness): automatic retry actually works now — non‑OK provider responses (429/500/502/503/504) are retried on a different provider with exponential backoff + jitter (capped, configurable policy in llm_router_api/endpoints/http_dispatch.py); when the retry budget is exhausted the provider's last status code is returned to the client instead of a masked 500; transport errors (connection refused/timeout) are retried the same way; replay safety documented (only reproducible JSON bodies are retried). Fix (correctness): client input errors (missing/invalid params, ValueError/pydantic ValidationError) now return 400 with {"error": {"message": …}} — the Flask registrar is the single owner of the exception→status mapping; real server‑side failures still return 500. Fix (correctness): provider health‑check now treats only 2xx/3xx as available (401/403 → auth_error, 404 → not_found, other 4xx/5xx diagnosed separately); ping path per api_type (vLLM /health, Ollama /api/version, OpenAI‑compatible /v1/models) with fallback probing and auth token forwarding; missing api_type no longer crashes the monitor; new hysteresis — a provider is marked down only after LLM_ROUTER_PROVIDER_MONITOR_MAX_CONSECUTIVE_FAILURES (default 2) consecutive failed pings, one success restores it (a busy host that is slow to accept a new connection but serves keep‑alive traffic no longer drops out of the pool); ping timeout configurable via LLM_ROUTER_PROVIDER_MONITOR_PING_TIMEOUT_SECONDS (default 5.0 s). Resilience: balanced/weighted load‑balancing counters moved to Redis (shared across gunicorn workers, atomic pick‑and‑increment, deterministic weighted sequence) with in‑memory fallback when Redis is unavailable (llm_router_api/core/lb/lb_counters.py). Lifecycle: FlaskEngine exposes an explicit, idempotent start()/stop() + context manager + atexit hook and a SIGTERM/SIGINT shutdown hook in rest_api — gunicorn workers no longer leak the services‑monitor thread (monitor threads are daemon=True). Data quality: /models no longer returns fabricated constants ("type":"vllm", "state":"not-loaded", "quantization":"4bit") — provider‑specific fields come from real per‑provider data or are null. Cleanup: removed dead LLM_ROUTER_AUTH_RATE_LIMIT_ENABLED switch and the broken, unregistered beta/AdaptiveStrategy; unified the Vault key‑store backend on hvac (was non‑installable hvault); setup.py no longer ships test sub‑packages in the wheel and pins the git dependencies to commits; single source of truth for the version (.version) enforced in CI against git tag and Helm appVersion; Helm deployment checksum annotation fixed to reference the real files/models-config.json; requirements*.txt pinned, [build-system] added to pyproject.toml, requirements.lock generated; print() and bare except: pass in production code replaced with logging or intentional comments; llm-router auth CLI rewritten on a single argparse pass (no argv re‑scanning) with working --store vault wiring. Tests: new suites for dispatch retry, 400‑input handling, provider‑monitor (status mapping, fallbacks, hysteresis), engine lifecycle, LB global counters, /models fields, hvac import and the auth CLI — 559 passed, 4 skipped. Changelog: bump version to 0.9.5 (.version, Helm chart appVersion). |
| 0.9.6 |
Fix LMStudio /api/v0/model returns proper response format. |
| 1.0.0-rc1 |
Chore Bump version to 1.0.0-rc1; Change python version from 3.13.9-slim-trixie to 3.14-slim-trixie, Fix: remove unused orig_params and reconnect_number from endpoint_i dispatcher. Docs: add ENDPOINT_DEV.md endpoint‑development guide (English version of endpoints/README-pl.md) and register it in site/docs.html; cross‑link it from root and llm_router_api READMEs. |
| 1.0.0 |
Docs: add "Endpoint Development Guide" to llm_router_api/docs with comprehensive explanations for creating, configuring, and extending REST endpoints, remove old Polish version. Chore: remove unused Docker entrypoint script and Supervisord configuration, update Dockerfile to simplify CMD setup. |
| 1.0.1 |
Chore: update _ENDPOINT_PERMISSION_MAP to adjust permissions for /api/ping and /api/version endpoints |
| 1.0.2 |
Feat: add the util subcommands to llm-router — translate, genai-classifier and genai-data-augmentation. These utilities are ported over from the standalone llm-router-utils repository and now run directly through the LLMRouter client. Feat: add opt-in --verbose (DEBUG) logging across the util, auth and config commands, with secrets always masked (***) in the logs. Refactor: extract the shared setup_logging/shorten helpers into llm_router_cli/log_utils.py and reuse the --verbose flag via BaseCommand.add_verbose (no per-command duplication). |
| 1.0.3 |
Features: Introduced persistent custom authentication policies with JSON support. Enabled file-based caching and improved policy management via CLI. Fixes: Resolved missing package error during __version__ lookup. Updated README with clarified flag descriptions and safer examples. Refactoring: Modularized provider discovery and configuration generation. Enhanced flexibility and error handling for CLI policy management. Improved record processing and added warnings for missing fields during tasks. Default prefix for the API Auth keys changed to sk-llmr-live. Testing: Increased test coverage for custom policy handling in CLI. Documentation: Added details on using custom auth policies and environment configurations like LLM_ROUTER_AUTH_CUSTOM_POLICIES_FILE. |
| 1.0.4 |
Tests: Add comprehensive unit tests and coverage configuration. Added unit tests for: (1) LLMRouterClient, CLI commands and helpers; (2) llm_router_api.core.server helper functions; (3) llm_router_api.rest_api CLI arguments and server dispatch logic; (4) Prometheus metrics, auth metrics, and stream converters; (5) Strategies, caching mechanisms, and monitoring functionality; (6) Endpoint handling and API conversions; (7) Constants, error handling, and environment-specific behavior; (8) Multiple endpoints, including Health, Ping, and version utilities. Add Manifest.in and author_email in setup.py. Prepared to PyPi package. |
| 1.0.5 |
Build for PyPi |
| 1.0.6 |
CLI: add the server command to llm-router for managing the REST API lifecycle: start (background daemon of python -m llm_router_api.rest_api; PID and log files under ~/.llm-router/; flags --foreground/--server/--host/--port/--models-config/--debug/--lb-strategy/--default-lang/--auth/--redis-*/--auth-redis-*; env defaults mirroring run-rest-api-gunicorn.sh, with CLI > shell > defaults precedence), status (running check, stale PID cleanup, and launch parameters recorded in the <pidfile>.run file: command, server engine, host, port, models config, start time, explicit env overrides); stop removes the .run record; new llm-router completion bash \|zsh command prints a tab‑completion script (commands, sub‑commands and long options) generated from the live CLI tree (install: eval "$(llm-router completion bash)" or source <(llm-router completion zsh)); log (colorized tail -f; --lines/--no-follow/--color), reload (graceful SIGHUP) and stop (SIGTERM; --force goes straight to SIGKILL). CLI: completion bash\|zsh gains --install (append the generated completion script to ~/.bashrc / ~/.zshrc, created if missing, idempotent marker block, optional --file PATH override); server start --foreground now writes the PID file and run record (so server status / server stop work) and removes them on exit; the .run record now snapshots all LLM_ROUTER_* environment variables the server was started with (shared ENV_PREFIX from llm_router_lib.core.constants), shown by server status. server status now renders a colored card (✓ All good / ✗ Not running header + Details and Environment sections, --color auto\|always\|never, --show-env to reveal the Environment section, hidden by default) and masks the values of sensitive LLM_ROUTER_* (passwords, secrets, tokens, API keys) so credentials are never echoed. server status Details now shows a single Log row for the application log (from LLM_ROUTER_LOG_FILENAME, default llm-router.log in the CWD; in daemon mode the daemon's stdout/stderr are captured into the same file). completion bash\|zsh now covers every sub-command nesting level (e.g. auth key generate, anonymizer run, config discover) and all their long options, server run records now store the absolute application log path (app_log_file): when LLM_ROUTER_LOG_FILENAME is a bare file name the path is anchored to the launch CWD, and server status shows it in the Log row. server log now follows the application log by default (record's app_log_file, else the shell's LLM_ROUTER_LOG_FILENAME, working in both daemon and --foreground mode); --log-file overrides it (e.g. to tail the daemon's console log ~/.llm-router/server.log). server start's daemon log (--log-file) now follows the user's LLM_ROUTER_LOG_FILENAME (a bare name anchored to the launch CWD) and falls back to ~/.llm-router/server.log only when the variable is unset. server status Details now shows the Models config as an absolute path: relative paths (e.g. resources/configs/models-config.json) are anchored to the launch CWD and stored in the run record (models_config). |
| 1.0.7 |
CLI: server is now multi-instance — every sub-command takes -i/--instance NAME (or $LLM_ROUTER_INSTANCE) and each named instance keeps a private state tree under ~/.llm-router/instances/NAME/ (pid, run record, logs, config.env, metrics, start lock), so several routers — one per project, environment or provider pool — run side by side on one host; the built-in default keeps the historical ~/.llm-router layout, so existing scripts and run-rest-api-gunicorn.sh are unaffected. New server list [--json] and server rm-instance NAME; stop/status accept --all. A named instance reads config.env with CLI flag > config.env > shell environment > defaults precedence, and start --save-config persists the flags of that command line (0600, matching keys rewritten in place, comments and ordering preserved) so a plain start -i NAME relaunches it identically. start adds a bind port pre-flight (--no-port-check to skip) and a server.start.lock against concurrent starts, and named instances get a private PROMETHEUS_MULTIPROC_DIR so their Prometheus counters are not wiped by each other. The completion script (bash / zsh) now proposes short options next to the long ones (--instance / -i) and keeps offering them mid-command. Adds llm_router_cli/tests/test_server_instances.py and the Instances section in llm_router_cli/README.md. Fix: named instances no longer share one application log — launch scripts export a bare LLM_ROUTER_LOG_FILENAME (llm-router.log), which the server resolves against its CWD, so every instance started from the same directory appended to one file whose rotating handlers truncated each other; a relative value is now anchored to ~/.llm-router/instances/NAME/ (./.. stripped, absolute and ~ paths honored), start warns when a running instance already logs to the same file, and server list shows the application log in the LOG column (--json reports app_log_file and log_file). Fix: a broken models config now fails on the console instead of inside the daemon log - start validates LLM_ROUTER_MODELS_CONFIG before spawning (--no-config-check to skip) and reports a daemon that dies immediately, with the tail of its log. |
| 1.0.8 |
CLI / Core: new --verbose flag on llm-router server start and LLM_ROUTER_VERBOSE env var to log raw, UNMASKED request parameters for debugging — endpoints (via SecureEndpointI) dump the original request payload to the log; on startup the REST API prints a prominent styled console warning and delays startup by 3 seconds so an accidental production start can still be aborted (never enable in production: PII goes to the log). API: ApiModelConfig gains safe_active_models_config — a sanitized copy of the active models configuration with sensitive fields (api_token) masked to an empty string; _utils_pipeline.apply now accepts an optional model_config parameter and endpoint_i passes the safe (token-masked) config into the utils pipeline instead of the raw config. CLI: llm-router server reload now restarts the instance instead of only signalling it — it stops the running server (SIGTERM with the 15 s grace period, SIGKILL with the new --force) and starts it again, replaying the previous launch flags from the run record (--host/--port/--server, --models-config as the absolute path that was actually loaded, the Redis/auth flags, --verbose, the daemon log file and an explicit --pid-file), while the instance config.env, the shell environment and the built-in defaults are applied again on top of them; --graceful keeps the previous behavior (SIGHUP to the Gunicorn master, no restart), reload with no running server fails with a server start hint, a stop timeout aborts before anything is killed, and a missing or unparsable recorded models-config aborts the reload before the stop so a bad path never takes a running instance down. Docs: LLM_ROUTER_VERBOSE documented in ENV_DEFINITIONS.md, ENDPOINT_DEV.md, CLI README and run scripts (gunicorn/docker), the README gaining a server reload section (restart semantics, --force/--graceful and the run-record flag replay). Tests: coverage for VERBOSE_MODE behavior, the verbose startup warning/delay, CLI --verbose handling, and the new reload lifecycle (a real stop-then-start of a detached process, --force/--graceful, the stop-refusal and models-config pre-flight aborts, and the run-record to start flag replay). CLI: server stop and server reload now wait for the whole process tree instead of only the PID-file process: Gunicorn's worker forks inherit the listening socket, so a fork that outlived the master kept the port bound and the reload start half failed on EADDRINUSE. Every fork is sampled before the master is signalled (after that it is reparented to init and no longer traceable), receives the SIGTERM itself, gets the same 15 s grace, and is SIGKILLed afterwards; the stop reports success and removes the PID file and run record only once master and workers are gone, otherwise it fails naming the surviving PIDs and leaves the state files for a retry. Tests: coverage for a lingering worker fork (the SIGTERM fan-out, the SIGKILL escalation after the grace period, the stop that keeps its state files when a fork survives, and a reload that reaches start only after the tree is empty). CLI: reload also replays -i/--instance — without it the restart resolved default, i.e. another config.env, PID file, application log and metrics tree next to the instance that was reloaded — together with the environment recorded for the stopped server, applied below config.env and the shell environment, so a launch that configured itself through LLM_ROUTER_* variables (the run-rest-api-gunicorn.sh wrapper) comes back on its own port/Redis/auth/models config instead of the built-in defaults. A stop now also counts a process that exited without being reaped (a zombie, or a task the kernel is still deleting) as gone, which previously surfaced as a spurious did not exit within 15s on a healthy server, gives the forks their own shorter drain once the master is down, and a foreground start no longer deletes the PID file of the server that already replaced it. Tests: named-instance restart argv, the recorded-environment layer against config.env and the shell, reload reaching start with the instance and its recorded port, and stopping an unreaped zombie. CLI: the server.start.lock now guards only the window between the "is it alive?" check and writing the PID file — it is released as soon as the PID is published, so a --foreground server no longer holds it for its whole lifetime and reload of such an instance no longer aborted with a start is already in progress for instance '…' (the stopped server's own start still owned the lock, and a lock under 60 s old is not treated as abandoned); the daemon path keeps the lock until the daemon publishes its PID, and release_start_lock removes a lock only when the PID recorded in it is still the releasing process, so a dying start cannot unlock a newer one. Tests: a foreground start leaves no lock while the server runs and still publishes its PID, and releasing a lock owned by another start is a no-op. Build: version bumped to 1.0.8. |
| 1.1.0 |
Features: Update llm-router completion cli, add new options. New first_available_optim_nworkers load-balancing strategy — per-provider worker slots (nworkers provider config field, default 1) on top of first_available_optim. Busy-worker slots live in a per-(model, provider) Redis sorted set <prefix>model:<model>:in_use:<provider> of expiring leases and are claimed by an atomic Lua script; a new "least loaded with a free slot" selection step (ranked by busy workers, last host and config order) runs before the first-available fallback, which waits for a slot release and raises TimeoutError when every provider is saturated. nworkers=1 reproduces the classic binary lock one-to-one; missing/invalid values fall back to 1 with a logged warning. Registered in POSSIBLE_BALANCE_STRATEGIES, the strategy facade and llm-router server start --lb-strategy; LB/config/env docs updated; new unit-test coverage (test_first_available_optim_nworkers.py). |
| 1.1.1 |
Lib: New AsyncLLMRouterClient (issue #121) — an asyncio/httpx alternative to LLMRouterClient with full parity: the same 14 methods, the same keyword‑only payload contract (shared utils/payload.py builder), the same retry policy (exponential back‑off; streaming never retried) and the same error translation (AuthenticationError / RateLimitError / LLMRouterError). Streaming: stream_conversation_with_model and stream_extended_conversation_with_model yield normalised StreamEvents (handles OpenAI SSE, Ollama NDJSON, data: [DONE] and {"error": ...} chunks) plus a collect_stream_text aggregation helper; stream_timeout (default None = no read timeout so long generations are not cut off). Transport: new AsyncHttpRequester (utils/http_async.py), StreamEvent response model and utils/stream.py line parser; HTTP status→exception mapping shared with the sync requester. Deps: httpx added to requirements.txt, pytest-asyncio to requirements-dev.txt. Tests: 196 new unit tests (test_http_async_requester.py, test_async_client.py, test_stream_events.py, all over httpx.MockTransport — no network) plus a manual async harness python -m llm_router_lib.tests.async_llm_router_client. |
| 1.1.2 |
Fix (first_available_optim_nworkers): the load-aware step now runs before the host-reuse steps, so a saturated hot host spreads to the least loaded provider instead of the next one in configuration order — previously steps "reuse any known host" and "pick an unused host" partitioned the whole provider list and the ranking could never decide anything (its tests passed only because they mocked those steps away). The "config order" tie-break now uses the caller's provider list instead of the position in a Redis SMEMBERS set. The load-aware ranking also counts busy workers in absolute terms instead of busy / nworkers, and “reuse the last host” is no longer a step of its own: it answered as soon as the previously used provider had any free slot, which filled one provider to saturation before the next one was considered at all, while the relative measure preferred the largest provider on every tie. Providers now fill in layers — one worker each in configuration order — so p1: nworkers=2, p2: nworkers=3, p3: nworkers=1 serve p1, p2, p3, p1, p2, p2; the warm host survives as a tie-breaker inside one load level. Fix: a worker slot that is never released stops being renewed after LLM_ROUTER_LB_SLOT_MAX_AGE_SECONDS (default 1800) and expires on its own — lease expiry alone only covered a dead process, so a request that forgot its token (an abandoned stream) shrank the provider's capacity until the router restarted. Fix: the “this host serves this model“ pin (host:<host>) now expires (LLM_ROUTER_LB_HOST_PIN_TTL_SECONDS, default 3600, refreshed on every selection) instead of reserving a host for its last model forever. Fix: nworkers is read from the live models-config.json instead of the health monitor's registration snapshot, so changing it takes effect on the next request, and a provider reported as active but absent from the configuration is no longer selected. Fix: start-up cleanup no longer scans a :occupancy suffix that no key has ever carried. Env: LLM_ROUTER_LB_SLOT_LEASE_SECONDS, LLM_ROUTER_LB_SLOT_MAX_AGE_SECONDS, LLM_ROUTER_LB_HOST_PIN_TTL_SECONDS, LLM_ROUTER_AUTH_VAULT_TOKEN, LLM_ROUTER_AUTH_CUSTOM_POLICIES_FILE and LLM_ROUTER_RATE_LIMITING_CONFIG now have defaults in run-rest-api-gunicorn.sh and in the CLI's built-in defaults, and the run script honours LLM_ROUTER_INSTANCE for the instance name. Tests: KeepAliveMonitor.on_tick_callback, the lease age cap, the host pin expiry and the live-configuration slot limit are covered. Fix: strategy_prefix was stored but never applied, so this strategy shared model:<model>, :last_host and :hosts with first_available / first_available_optim; model keys are namespaced under fa_optim_nworkers_ (host occupancy stays shared on purpose). Fix: start-up cleanup no longer deletes the :in_use state of every model — with clear_buffers defaulting to True, each new worker process freed the slots of the requests other workers were still serving, over-admitting past nworkers. Changed: a worker slot is an expiring lease (per-request member in a sorted set, renewed from the KeepAliveMonitor thread, throttled to a third of the lease lifetime) instead of a bare counter, so a router killed mid-request stops occupying a provider after LLM_ROUTER_LB_SLOT_LEASE_SECONDS (default 120) rather than permanently; the lease token travels through ApiModel. Acquiring now copies the provider dictionary instead of writing to the shared model configuration, which nworkers > 1 had made racy. Redis failures in the new code paths degrade selection instead of raising, and are logged. Deps: lupa added to requirements-dev.txt so fakeredis executes the real Lua scripts. Tests: harness no longer stubs _get_redis_key, _provider_field, _host_key or _get_active_providers, and reports active providers in non-configuration order; 53 tests for the strategy, 3 for the lease round trip through ApiModel. Docs: LB_STRATEGIES.md, MODELS_CONFIG.md and ENV_DEFINITIONS.md corrected. Core: Introduce LLamaCPPApiType and update dispatcher registry to use it for llama.cpp. |
| 1.1.3 |
Feature: new model-level fallback_model in models-config.json — when no provider can serve a model (no providers configured, no healthy provider, or every provider busy until the selection timeout) ModelHandler reroutes the request before load balancing to the declared fallback model, and the very same LB strategy picks a provider of that model. Chains are supported (a -> b -> c); unknown targets, non-string values and cycles are rejected at startup. The model that actually serves the request is what the client sees in model, what releases the provider lock in Redis, and what metrics report; each switch logs a WARNING, an exhausted chain logs the whole chain, and the new llm_router_model_fallback_total{model_name, fallback_model} counter counts rerouted requests. Strategies gain a tri-state, non-blocking has_available_provider(...) probe (Redis health based) so an unhealthy primary is skipped instead of waiting out the 60 s timeout; models without fallback_model behave exactly as before. |
| 1.1.4 |
Feature (provider failover): a request is now rotated over the providers of its model before the fallback_model is used. Every 4xx/5xx answer - and every transport error - replays the request on another provider, for streaming and non-streaming alike: the providers a request failed on are carried in its options (__attempted_providers), ModelHandler drops them from the candidates (_candidates_of) and the fallback_model chain is entered only when the model has no untried provider left, so the last provider error is what the client finally receives. Streams no longer open lazily: HttpDispatch.stream_or_rerun() waits for the first chunk, so a provider that rejects the stream request (error status before any content) is swapped before the client can see a truncated answer, and only a fully exhausted chain emits the single error chunk (SSE, NDJSON for the OpenAI→Ollama conversion). Such pre-content failures raise the new ProviderStreamError (core/errors.py), deliberately not a requests.RequestException, so the per-format generators still turn mid-stream failures into their error chunk; ModelHandler.has_provider_candidates(...) tells the dispatcher whether another attempt has anything left to try. Escape hatch: RETRY_ON_ANY_ERROR_STATUS = False on an endpoint's RetryResponse restores the historical allow-list (429, 500, 502, 503, 504). Unreachable providers are rotated as well: a stream whose provider never answered (connection refused, connect timeout, DNS failure) now raises the same failover signal instead of an error chunk, so another provider of the model is tried - all six stream converters open the request through the shared stream_handler._open_stream(...). A failure mid-stream, after bytes reached the client, still ends the response with the provider's error chunk: a partially delivered stream cannot be replayed. ProviderStreamError gained error_code and reason (HTTP 503 vs connection_error / timeout), and the new core/errors.connection_error_code(...) labels transport failures for llm_router_provider_errors_total (shared with the non-streaming retries). A 200 OK whose body dies before the first chunk (or times out while reading it) rotates as well: stream_handler._guarded_body(...) translates only the very first read of the provider body, so the mid-stream behaviour - the final error chunk - stays untouched. |
| 1.1.5 |
Provider Selection: Fixed has_provider_candidates to consider provider health while deciding if an attempt should proceed. Added a NoProviderAvailable exception to improve provider handling and dispatch logic, ensuring HTTP 503 is returned for exhausted fallback chains. Error Handling: Resolved an issue where a TimeoutError during provider selection incorrectly returned a 500 status, now correctly responding with 504. Improved handling of rejected streams by ensuring HTTP status codes reflect the actual error, replacing silent 200 statuses when no data was yielded. Metrics Tracking: Addressed a bug where generation_time was incorrectly measured across concurrent requests, now isolated using ContextVar. Fixed inconsistent metrics caching behavior to ensure reliable router metrics. Other Fixes: Addressed edge cases such as run_ep(None) handling to avoid unnecessary errors. Enhanced fallback behavior and restricted _get_active_providers to utilize the caller's shortlist. |
| 1.1.6 |
Build GH CI workflow updates |