llm-router/docs

Changelog#

Version Changelog
0.0.1 Initialization, License, setup, interface for each endpoint and sample ping EP. Autoloader of builtin endpoints and for the future implementations.
0.0.2 Add base models for api call (module llm_proxy_rest.data_models with error.py handling. Decorators to check required params and to measure the response time.
0.0.3 Proper AutoLoading for each found endpoint. Implementation of ApiTypesDispatcher, ApiModelConfig, ModelHandler. Ollama endpoints: /, tags. Added endpoint to full proxy with params. Streaming in case when external api provides stream.
0.0.4 All llama-service endpoints are refactored to llm-proxy-api. Refactoring base ep_run method. Proper handling system message, prompt name, model etc.
0.1.0 Repository name changed from llm-proxy-api to llm-router. Added class HttpRequestExecutor to handle http requests from EndpointWithHttpRequestI. Handled routing between any models: openai -> ollama and ollama -> openai
0.1.1 Prometheus metrics logging. Workers/Threads/Workers class is able to set by environments. Streaming fixes. Multi-providers for single model with default-balanced strategy.
0.2.0 Add balancing strategies: balanced, weighted, dynamic_weighted and first_available which works for streaming and non streaming requests. Included Prometheus metrics logging via /metrics endpoint. First stage of llm_router_lib library, to simply usage of llm-router-api.
0.2.1 Fix stream: OpenAI->Ollama, Ollama->OpenAI. Add Redis caching of availability of model providers (when using first_available strategy). Add llm_router_web module with simple flask-based frontend to manage llm-router config files.
0.2.2 Update dockerfile and requirements. Fix routing with vLLM.
0.2.3 New web configurator: Handling projects, configs for each user separately. First Available strategy is more powerful, a lot of improvements to efficiency.
0.2.4 Anonymizer module, integration anonymization with any endpoint (using dynamic payload analysis and full payload anonymisation), dedicated /api/anonymize_text endpoint as memory only anonymization. Whole router may be run in FORCE_ANONYMISATION mode.
0.3.0 Anonymization available with three strategies: fast_masker, genai, prov_masker.
0.3.1 Refactoring lb.strategies to be more flexible modular. Introduced MaskerPipeline and GuardrailPipeline both configured via env. Removed genai-based masking endpoint.
0.4.0 The main repository is divided into dedicated ones: plugins, services, web — separate repositories. Clean up the whole repository. Examples of integration with llamaindex, langchain, openai, litellm and haystack.
0.4.1 Audit log is stored using GPG. Add bash script (scripts/gen_and_export_gpg.sh to prepare GPG keys and simple scripts/decrypt_auditor_logs.sh to decrypt encrypted audit logs. Moved core functionality from base to module core module. Quickstart.
0.4.2 Fix first_available_optim Strategy. Add KeepAliveMonitor to periodically pings model endpoints to keep them warm.
0.4.3 Add custom Prometheus metrices for logging masker/guardrail inidents. Fix OpenAI compatible v1 /models endpoint. Introduce monitors: services and keep alive models. Fixed guardrail retunr in case when streaming.
0.4.4 Validate unique provider identifiers. Store all hosts with keep‑alive configured in a Redis. UtilsPlugin pipeline with LangChain based simple RAG plugin (extending context to GenAI with locally built databse). Add handling of v1/response endpoint
0.4.5 Fixed sreaming to LMStudio native. Refactor streaming module.
0.4.6 Added support for embeddings endpoints across all providers. Extended ApiModel and ApiTypesI with is_embedding flag. Added test_embeddings.py utility for verifying embedding models through the API.
0.4.7 Integration with native Anthropic API. Add translate, generative_answer and ping methods to LLMRouterClient (with tests). Refactor LLMRouterCkientServices to use self.model_cls. Add payload converter for vLLM.
0.5.0 Integration with PII masker, code refactoring
0.5.1 Add /v1/messages endpoint (Claude Agent compatibility), streaming Cache‑Control/Pragma/Expires/Vary headers, mandatory Redis (runtime error on missing connection), non‑root Docker startup, remove ml‑utils dependency, add gnupg to requirements, update default plugin to simple_semantic_routing.
0.5.2 Prevent network topology leak in error messages (full details remain in server‑side logs), example models-config.json no longer contains real internal IPs. Introduced LLM_ROUTER_MAX_REQUEST_BODY_SIZE to set the maximum content length. Sanitize all error messages returned to the Client. Local security.
0.6.0 Authentication system: API key-based auth with multi-backend key stores (Memory, Redis, Vault), plaintext and secret-key lookup, enable/disable keys, seed-file persistence. Auth CLI: auth subcommands for managing API keys (create/list/enable/disable/delete) with formatted tabular output and prefix matching. Rate limiting: Per-key rate limiting via token bucket, predefined rate-limiting policies in rate_limiting-policies.json, PolicyEngine accepting dict key records. Anonymizer CLI: Migrated fast_masker to anonymizer CLI with deprecation warning; moved to masker subpackage. Infrastructure: Shared Redis client across stores and cache, dynamic column widths for CLI output, environment variable updates for auth/rate-limiting/audit logging config.
0.6.1 Added config CLI command with discover (auto-discover local Ollama/vLLM/LM Studio providers) and merge (deep-merge multiple models-config.json files).
0.6.2 CLI: fix _RATE_LIMIT_COMMANDS NameError in auth CLI (commit 2d9e593). Core: extract Prometheus multiprocess dir handling into MetricsHandler.prepare_multiproc_dir() with sane default; add LLM_ROUTER_AUTH_MEMORY_SEED_FILE env var for memory store seed path. CLI: restructure config commands, clean up imports across core modules, add llm_router_client.py test script with predefined model tests.
0.6.3 Core: consolidate streaming mode flags into StreamConversion enum, simplify handler dispatch and fix swapped handler type hints in OpenAI/Ollama endpoints. Simplify error handling and response logic in endpoint_i.py.
0.6.4 Refac: Clear REAME files.
0.6.5 CLI discover: added KoboldCpp and TabbyAPI providers to config discover auto-discovery (now covers 6 local providers: Ollama, vLLM, LM Studio, llama.cpp, KoboldCpp, TabbyAPI). Infrastructure: fixed _scan_and_merge passing best_port as host parameter causing empty config; removed deprecated llm-router-fast-masker entry point and masker module.
0.6.6 Auth CLI: restored working key generation by fixing _handle_key instance method call (cls._handle_key → cls()._handle_key). Config CLI: fixed _do_discover/_do_merge calls to use class methods properly (_do_discover(args) → cls._do_discover(args)). Removed all legacy backward-compatibility shim comments and module-level functions from CLI command modules.
0.6.7 Docs: consolidate environment variable documentation into ENV_DEFINITIONS.md, simplify references across README files, add semantic_biencoder_routing env var details, remove LLM_ROUTER_HOST and update default MASKING_STRATEGY_PIPELINE. Chore: add /metrics endpoint to public endpoints list (LLM_ROUTER_AUTH_PUBLIC_ENDPOINTS) across auth and config modules; bump version to 0.6.7.
0.6.8 Feature: Prometheus router metrics, Grafana dashboard, and /metrics public endpoint for improved observability and visualization. Docs: consolidate env var docs into ENV_DEFINITIONS.md, update docstrings and line wrapping across modules. Chore: expand mypy checks, add types-setuptools.
0.6.9 Chore: update prompts to be more specific. Feature: add prepare_response_function to endpoint_i, add GenerateNewsFromTextService, add default_model parameter to Client to provide fallback model configuration
0.7.0 Feature: add polarity_3c endpoint (/api/polarity_3c) to detect 3-class text polarity (ambivalent, positive, negative) with PL/EN system prompts. Lib: add Polarity3cModel, Polarity3cService, and LLMRouterClient.polarity_3c() method to llm_router_lib. Auth & Docs: register endpoint with builtin policy permission, update endpoint documentation and add comprehensive test suite.
0.8.0 Feature: add endpoints: /api/simplify_text, /api/generate_questions, /api/create_full_article_from_texts, /api/generate_article_from_texts, /api/generate_label with models, services and matching LLMRouterClient methods. Auth: register the new endpoints with the builtin policy permission. Fix _ensure_alternating_roles in endpoint_i to build a valid payload by merging consecutive messages of the same role (joining their contents), folding all system messages into one leading message, and prepending/appending empty user placeholders when the dialogue starts with or ends with an assistant turn.
0.9.0 Refactor (dependency direction): llm_router_lib no longer imports from llm_router_api. The shared DEFAULT_EP_LANGUAGE constant (and the LLM_ROUTER_ env prefix) now live in the dependency‑free llm_router_lib.core.constants; llm_router_api.base.constants_base re‑exports them for backward compatibility. BREAKING – structural rename anchored on the endpoint URL (HTTP paths are unchanged). Canonical name = URL without /api/; class/model/service/client‑method/constant names now follow it and legacy aliases are removed: TranslateTexts→Translate, TranslateTextModel→TranslateModel, TranslateTextService→TranslateService; SimplifyTexts→SimplifyText, client simplify_texts()→simplify_text(); GenerateQuestionsFromTexts→GenerateQuestions, GenerateQuestionFromTextsModel→GenerateQuestionsModel, GenerateQuestionsFromTextsService→GenerateQuestionsService (alias GenerateQuestionFromTextsService removed), client generate_questions_from_texts()→generate_questions(), constants GENERATE_Q_*→GENERATE_QUESTIONS_*; GenerateNewsFromTextHandler→GenerateArticleFromText, GenerateNewsFromTextService→GenerateArticleFromTextService, client generate_news_from_text()→generate_article_from_text(), constants GENERATE_ART_*→GENERATE_ARTICLE_FROM_TEXT_*; GenerateArticleFromTexts constants GENERATE_ARTICLES_*→GENERATE_ARTICLE_FROM_TEXTS_*; FullArticleFromTexts→CreateFullArticleFromTexts, CreateArticleFromNewsListModel→CreateFullArticleFromTextsModel, constants FULL_ARTICLE_*→CREATE_FULL_ARTICLE_FROM_TEXTS_*; AnswerBasedOnTheContext→GenerativeAnswer, AnswerBasedOnTheContextModel→GenerativeAnswerModel, constants CONTEXT_ANSWER_*→GENERATIVE_ANSWER_*; GenerativeConversationModel→ConversationWithModelRequest, ExtendedGenerativeConversationModel→ExtendedConversationWithModelRequest, ConversationService→ConversationWithModelService, ExtendedConversationService→ExtendedConversationWithModelService. Class inheritance hierarchy and the auth _ENDPOINT_PERMISSION_MAP (URL‑keyed) are unchanged. Refactor (structure only, no behavior change): shared base TextListUtilityEndpoint for the texts utility endpoints (common prepare_payload — one user message per source text, build_map_prompt hook, per‑text _prepare_response helper); message normalisation moved verbatim into pure functions in llm_router_api/endpoints/message_normalizer.py (thin delegating methods kept on the endpoint class); outbound HTTP dispatch + retry orchestration moved verbatim into llm_router_api/endpoints/http_dispatch.py (HttpDispatch; EndpointWithHttpRequestI.RetryResponse kept as a backward‑compatible alias of http_dispatch.RetryPolicy); run_ep slimmed into a readable orchestrator with extracted sequence helpers (endpoint_i.py reduced from 2123 to ~1810 lines). Fix (behavior): per‑text endpoints now actually return one result per source text — the role normalizer (introduced in 0.8.0 to build valid chat payloads) was applied unconditionally in run_ep and merged the consecutive user messages built by the texts endpoints, so polarity_3c, translate, simplify_text and generate_questions silently returned one merged result; run_ep now skips role normalisation for call_for_each_user_msg=True endpoints (the HTTP executor already dispatches each text as its own [system, user] call), while conversation endpoints (conversation_with_model, extended_conversation_with_model) keep it; client code consuming a single merged entry for those four endpoints must iterate the response list. Tests: regression grid extended to all remaining builtin endpoints; new test_http_dispatch.py (dispatch/retry, late‑binding, RetryResponse alias) and end‑to‑end test_per_text_dispatch.py (per‑text ⇒ one model call per text ⇒ results paired per text; chat endpoint ⇒ consecutive user messages still merged). Feature: Add log rotation. Feature: add method models() to client.py, add public /health endpoint
0.9.1 BREAKING – unified LLMRouterClient calling contract (pre‑1.0): all 12 endpoint methods (conversation_with_model, extended_conversation_with_model, polarity_3c, translate, simplify_text, generative_answer, generate_article_from_text, generate_article_from_texts, create_full_article_from_texts, generate_questions, generate_label, …) now share one keyword‑only signature: payload (a Pydantic request model instance) or named domain arguments + model + optional generation options; argument order is fixed (payload → domain fields → model → temperature → max_new_tokens). Raw dict payloads are removed – passing a dict as payload raises TypeError (build the matching model explicitly, e.g. MyModel(**payload)). Hard‑coded defaults in signatures (0.2/0.75/256/512/1024/64, number_of_questions=1) are removed – generation defaults come exclusively from the Pydantic models (GenerativeOptions); optional parameters default to None so None‑valued kwargs no longer break Pydantic validation (fixes the texts=None bug). conversation_with_model / extended_conversation_with_model gain the same kwargs contract (user_last_statement, historical_messages, system_prompt, model, temperature, max_new_tokens); calling them with neither payload nor user_last_statement now raises NoArgsAndNoPayloadError (previously TypeError). Partial kwargs (required field missing, or no resolvable model name) raise NoArgsAndNoPayloadError instead of a raw Pydantic ValidationError. Type aliases _ConvPayload / _ExtConvPayload removed; Complete numpydoc docstrings added to every endpoint method (previously missing on generative_answer and generate_article_from_text). Tests: new llm_router_lib/tests/test_client_unified_api.py (89 unit tests, python -m unittest); api tests for dict payloads rewritten to assert TypeError, plus new kwargs‑based conversation client tests. Style: unify all type annotations across llm_router_lib, llm_router_api and llm_router_cli to the typing module (Optional[...] instead of X | None, List/Dict/Tuple instead of the built-in list/dict/tuple generics, Union[...] for pure unions) so the codebase is consistent; runtime behaviour is unchanged. Docs: llm_router_lib/README.md and RESPONSE_MODELS.md updated to the keyword‑only calling contract (raw‑dict payloads rejected, previously missing methods documented).
llm-router · docs are generated from the repository by tools/build_docs.py 0.9.2 @ ea7b8e7