llm-router/docs

llm_router_lib#

Overview#

llm_router_lib bundles Pydantic data‑model definitions and a thin, opinionated client wrapper for the LLM‑Router service.

  • The data models live in llm_router_lib/data_models and describe every request payload the router accepts.
  • The client (LLMRouterClient) offers a high‑level, Pythonic API that hides HTTP details, retries, and error handling.
  • Low‑level service classes (ConversationWithModelService, ExtendedConversationWithModelService, TranslateService, GenerativeAnswerService, health services) perform the actual HTTP calls and can be used directly when finer‑grained control is required.
  • HttpRequester (in utils/http.py) is a small wrapper around requests that adds logging, configurable retries, and unified error translation.
  • A dedicated exception hierarchy (exceptions.py) maps HTTP errors to meaningful Python exceptions.

In short, llm_router_lib provides both the contract (the “schema”) and a convenient client to consume the router service.

Installation#

The library targets Python 3.10.6 and uses a virtualenv. Install it in editable mode for development:

# Clone the repository (if you haven't already)

git clone https://github.com/radlab-dev-group/llm-router.git
cd llm-router/llm_router_lib

# Create and activate a virtual environment

python3 -m venv .venv
source .venv/bin/activate

# Install the package and its dependencies

pip install -e .

All runtime dependencies (requests, pydantic, plus the packages listed in requirements.txt) are declared in the project’s requirements.txt.

Quick start#

from llm_router_lib import LLMRouterClient

# Initialise the client – point it at the router’s host (do **not** include the `/api` prefix)

client = LLMRouterClient(
    api="http://localhost:8080",  # router host URL
    token="YOUR_ROUTER_TOKEN",  # optional, if router requires auth
)

# Build a payload using the provided data model (validation is automatic)

payload = {
    "model_name": "google/gemma-3-12b-it",
    "user_last_statement": "Hello, how are you?",
    "temperature": 0.7,
    "max_new_tokens": 128,
}

# Call the standard conversation endpoint

response = client.conversation_with_model(payload)

# response is a typed `ConversationResponse` model:
print(response.response)  # → the assistant's reply text
print(response.generation_time)  # → seconds taken by the server

You can also pass a pydantic model instance directly:

from llm_router_lib.data_models.builtin_chat import ConversationWithModelRequest

model = ConversationWithModelRequest(
    model_name="google/gemma-3-12b-it",
    user_last_statement="Hello, how are you?",
    temperature=0.7,
    max_new_tokens=128,
)

response = client.conversation_with_model(model)

Data models#

All request payloads are defined in llm_router_lib/data_models.
A common base class supplies shared options:

class BaseModelOptions(BaseModel):
    """Options shared across many endpoint models."""
    mask_payload: bool = False
    masker_pipeline: Optional[List[str]] = None

Conversation models#

Model Required fields Optional / extra fields
ConversationWithModelRequest model_name, user_last_statement temperature, max_new_tokens, historical_messages, …
ExtendedConversationWithModelRequest All of the above + system_prompt –

Utility models (selected examples)#

Model Required fields Optional fields (generation parameters)
GenerateQuestionsModel texts + model_name number_of_questions, generation opts
GenerateArticleFromTextModel text + model_name generation opts
TranslateModel texts + model_name generation opts
GenerativeAnswerModel question_str, texts + model_name doc_name_in_answer, question_prompt, system_prompt, generation opts
OpenAIChatModel (OpenAI‑compatible) model, messages stream, keep_alive, language, options
SimplifyTextModel texts, model_name generation opts
CreateFullArticleFromTextsModel user_query, texts, model_name article_type, generation opts
GenerateLabelModel texts + model_name generation opts

(All utility models inherit from BaseModelOptions and therefore share the mask_payload and masker_pipeline flags.)

Services (low‑level wrappers)#

If you need direct access to the HTTP layer, the library exposes a set of service classes in llm_router_lib/services:

Service class Endpoint (relative to api) Payload model (if any)
ConversationWithModelService /api/conversation_with_model ConversationWithModelRequest
ExtendedConversationWithModelService /api/extended_conversation_with_model ExtendedConversationWithModelRequest
TranslateService /api/translate TranslateModel
SimplifyTextService /api/simplify_text SimplifyTextModel
GenerativeAnswerService /api/generative_answer GenerativeAnswerModel
GenerateQuestionsService /api/generate_questions GenerateQuestionsModel
GenerateLabelService /api/generate_label GenerateLabelModel
PingService /api/ping none
VersionService /api/version none

These services inherit from BaseConversationServiceInterface, which provides call_post and call_get helpers that perform JSON parsing and raise the library‑specific exceptions on failure.

Example: using a service directly#

from llm_router_lib.services.conversation import ConversationWithModelService
from llm_router_lib.utils.http import HttpRequester
import logging

http = HttpRequester(base_url="http://localhost:8080", token="...", timeout=10)
logger = logging.getLogger("demo")

service = ConversationWithModelService(http, logger)
payload = {
    "model_name": "google/gemma-3-12b-it",
    "user_last_statement": "Hi!",
}
response = service.call_post(payload)
print(response)

Thin client wrapper (LLMRouterClient)#

LLMRouterClient aggregates the low‑level services and exposes a concise, high‑level API:

Method Description
conversation_with_model(payload) Calls /api/conversation_with_model. Accepts a dict or a ConversationWithModelRequest.
extended_conversation_with_model(payload) Calls /api/extended_conversation_with_model. Accepts a dict or an ExtendedConversationWithModelRequest.
translate(payload=None, texts=None, model=None) Calls /api/translate. Three usage patterns:
1️⃣ Pass a ready‑made dict.
2️⃣ Pass a TranslateModel instance.
3️⃣ Provide texts + model and let the client build the model.
simplify_text(payload=None, texts=None, model=None) Calls /api/simplify_text. Three usage patterns:
1️⃣ Pass a ready‑made dict.
2️⃣ Pass a SimplifyTextModel instance.
3️⃣ Provide texts + model and let the client build the model.
generate_questions(payload=None, texts=None, number_of_questions=1, model=None) Calls /api/generate_questions. Generates questions from input texts. Alias: generate_questions.
generate_label(payload=None, texts=None, model=None) Calls /api/generate_label. Generates a single category name (label) from a list of texts. Three usage patterns:
1️⃣ Pass a ready‑made dict.
2️⃣ Pass a GenerateLabelModel instance.
3️⃣ Provide texts + model and let the client build the model.
generative_answer(payload=None, model=None, texts=None, question_str=None) Calls /api/generative_answer. Works with a dict, a GenerativeAnswerModel instance, or explicit arguments.
ping() Calls /api/ping – health‑check endpoint.
version() Calls /api/version – retrieves router version information.
translate(...) and generative_answer(...) also raise NoArgsAndNoPayloadError if called without required arguments.

All methods return a typed response model (a Pydantic BaseModel) validated from the JSON body — see RESPONSE_MODELS.md for the full reference and the method→model mapping. Call .model_dump() on the result to get the plain dict back. Errors from the underlying HTTP layer are translated into the following exceptions (defined in exceptions.py):

  • LLMRouterError – base class for all library‑specific errors.
  • AuthenticationError – HTTP 401/403 (invalid or missing token).
  • RateLimitError – HTTP 429 (too many requests).
  • ValidationError – HTTP 400 (malformed payload). | NoArgsAndNoPayloadError – client‑side validation when required arguments are missing.

:::tip Field naming across models

Models for the built‑in generative endpoints use model_name as the field key (e.g. ConversationWithModelRequest). The OpenAI‑compatible endpoint model uses model instead, to match the official OpenAI API schema:

Model class Key for model identifier
ConversationWithModelRequest model_name
OpenAIChatModel model

:::

:::tip Context manager support

All public types (LLMRouterClient, HttpRequester) implement __enter__ / __exit__, so they can be used with the with statement to guarantee resource cleanup:

from llm_router_lib import LLMRouterClient

with LLMRouterClient(api="http://localhost:8080", token="...") as client:
    result = client.conversation_with_model(payload)  # session closed automatically

:::

Utilities#

  • utils/http.py – HttpRequester
    Handles URL construction, bearer‑token injection, configurable retries (via urllib3.Retry), and unified error mapping. It returns the raw requests.Response after validation.

  • exceptions.py – centralised exception definitions (see above).

Development & testing#

The repository includes a small test harness under llm_router_lib/tests. Example usage:

python -m llm_router_lib.tests.llm_router_client

This script spins up a LLMRouterClient instance and runs a suite of end‑to‑end tests covering conversation, extended conversation, translation, generative answering, and health checks.

llm-router · docs are generated from the repository by tools/build_docs.py 0.9.0 @ 832c832