Skip to content

reflexr.litellm

The litellm extra. See The LLM gateway.

LiteLLM as reflexr's model gateway: the [litellm] extra (ADR-0022).

An adapter (ADR-0025). The model is pydantic-ai's port: litellm_model is a model that calls a LiteLLM proxy through pydantic-ai's LiteLLMProvider, and LiteLLMGateway is a capability that attaches the tenant, the causal chain, the trace, a tenant's key and the rule's guardrails to every request::

agent = Agent(
    litellm_model("claude-sonnet", api_base="http://litellm:4000"),
    deps_type=Reaction[AppDeps],
    capabilities=[
        EventContext(emit=[IncidentOpened]),
        LiteLLMGateway(tenant_key=keys.for_tenant, guardrails=policy.for_rule),
    ],
)

Routing, fallbacks, budgets, rate limits and guardrails stay in the proxy's configuration.

The model

litellm_model

litellm_model(
    model: str,
    *,
    api_base: str | None = None,
    api_key: str | None = None,
    http_client: AsyncClient | None = None,
    settings: ModelSettings | None = None,
) -> OpenAIChatModel

Return a pydantic-ai model that calls a LiteLLM proxy's model group.

Routing, fallbacks, budgets and guardrails are the proxy's configuration; the model only names the group.

Shared verbatim with artifactr's src/artifactr/litellm/gateway.py; change both.

Parameters:

Name Type Description Default
model str

The proxy's model group, such as claude-sonnet.

required
api_base str | None

The proxy's URL, such as http://litellm:4000. Defaults to the provider's environment variables.

None
api_key str | None

The proxy key used when a request carries no tenant's key.

None
http_client AsyncClient | None

The HTTP client, for tests or a shared connection pool.

None
settings ModelSettings | None

The model's default settings, as every pydantic-ai model takes them, such as a temperature, or extra_body={"mock_response": "..."} for a reply the proxy mocks. An agent's or a run's model_settings override them key by key, so an extra_body there replaces this one; LiteLLMGateway adds its metadata to whichever extra_body a request ends up with.

None

The gateway

LiteLLMGateway dataclass

LiteLLMGateway(
    tenant_key: TenantKey | None = None,
    guardrails: GuardrailPolicy | None = None,
    tags: Sequence[str] = (),
)

Bases: AbstractCapability[Reaction[Any]]

Attach tenancy, the chain, the trace, a tenant's key and guardrails to each request.

Give it to an agent beside EventContext, with a litellm_model. Before each model request it adds, through the request's extra_body and extra_headers:

  • LiteLLM metadata: the tenant, workspace, rule, scope and run, the causal chain as the session, the person whose event fired the rule as the user, the trace id (so LiteLLM's own traces join the run's), and tags
  • the W3C trace context, as traceparent
  • the guardrails the rule's policy names
  • the tenant's virtual key, as the request's Authorization, from tenant_key

A request a guardrail blocks fails the attempt with a GuardrailBlocked: the run is dead-lettered as guardrail_blocked, never retried.

Parameters:

Name Type Description Default
tenant_key TenantKey | None

Returns a tenant's LiteLLM key; the model's own key is used without one.

None
guardrails GuardrailPolicy | None

Returns the guardrails a rule's requests use in a workspace.

None
tags Sequence[str]

More tags for every request, such as the application's name.

()

before_model_request async

before_model_request(
    ctx: RunContext[Reaction[Any]],
    request_context: ModelRequestContext,
) -> ModelRequestContext

Add the gateway's metadata, guardrails and key to the request.

request_options async

request_options(
    reaction: Reaction[Any],
) -> tuple[dict[str, Any], dict[str, str]]

Return the extra body and headers of a model request in a run attempt.

on_model_request_error async

on_model_request_error(
    ctx: RunContext[Reaction[Any]],
    *,
    request_context: ModelRequestContext,
    error: Exception,
) -> ModelResponse

Turn a guardrail's block into a typed run failure; let other errors through.

wrap_run_event_stream async

wrap_run_event_stream(
    ctx: RunContext[Reaction[Any]],
    *,
    stream: AsyncIterable[AgentStreamEvent],
) -> AsyncIterable[AgentStreamEvent]

Turn a guardrail's block into a typed run failure in a streamed run too.

A streamed request is sent as its events are first read, so its HTTP error surfaces here rather than in on_model_request_error.

TenantKey module-attribute

TenantKey = Callable[[TenantId], Awaitable[str | None]]

Returns the LiteLLM virtual key of a tenant's team, or None for the model's own key.

A port the application implements (ADR-0025): keys come from its secret store, and reflexr never logs them or puts them on spans.

GuardrailPolicy module-attribute

Returns the names of the LiteLLM guardrails a rule's requests use in a workspace.

Guardrail blocks

GuardrailBlocked

GuardrailBlocked(guardrail: str | None)

Bases: RunFailure

A model request that one of the gateway's guardrails blocked.

The run is dead-lettered with the reason guardrail_blocked and this message, which names the guardrail but never repeats what was blocked. It is permanent: retrying the same input would be blocked again.

Parameters:

Name Type Description Default
guardrail str | None

The guardrail's name, when the gateway said.

required

guardrail_block

guardrail_block(
    error: ModelHTTPError,
) -> GuardrailBlocked | None

Return the guardrail block a model request's HTTP error reports, if it is one.

LiteLLM answers a request a guardrail blocks with HTTP 400 and an error that names the guardrail; other errors are not guardrail blocks.

GUARDRAIL_BLOCKED module-attribute

GUARDRAIL_BLOCKED = 'guardrail_blocked'

The failure reason of a run attempt whose model request a guardrail blocked.