Skip to content

artifactr.litellm

The litellm extra. See The LLM gateway.

LiteLLM as artifactr's model gateway: the [litellm] extra (ADR-0031).

An adapter (ADR-0034). The model is pydantic-ai's port: litellm_model is a model that calls a LiteLLM proxy through pydantic-ai's LiteLLMProvider, and LiteLLMGateway is a capability that attaches the tenant, the session, the trace, a tenant's key and the workspace's guardrails to every request::

agent = Agent(
    litellm_model("claude-sonnet", api_base="http://litellm:4000"),
    deps_type=Session[None],
    capabilities=[
        ArtifactWorkspace(types=[Doc]),
        LiteLLMGateway(tenant_key=keys.for_tenant, guardrails=policy.for_workspace),
    ],
)

Routing, fallbacks, budgets, rate limits and guardrails stay in the proxy's configuration.

The model and the gateway

litellm_model

litellm_model(
    model: str,
    *,
    api_base: str | None = None,
    api_key: str | None = None,
    http_client: AsyncClient | None = None,
    settings: ModelSettings | None = None,
) -> OpenAIChatModel

Return a pydantic-ai model that calls a LiteLLM proxy's model group.

Routing, fallbacks, budgets and guardrails are the proxy's configuration; the model only names the group.

Shared verbatim with reflexr's src/reflexr/litellm/gateway.py; change both.

Parameters:

Name Type Description Default
model str

The proxy's model group, such as claude-sonnet.

required
api_base str | None

The proxy's URL, such as http://litellm:4000. Defaults to the provider's environment variables.

None
api_key str | None

The proxy key used when a request carries no tenant's key.

None
http_client AsyncClient | None

The HTTP client, for tests or a shared connection pool.

None
settings ModelSettings | None

The model's default settings, as every pydantic-ai model takes them, such as a temperature, or extra_body={"mock_response": "..."} for a reply the proxy mocks. An agent's or a run's model_settings override them key by key, so an extra_body there replaces this one; LiteLLMGateway adds its metadata to whichever extra_body a request ends up with.

None

LiteLLMGateway dataclass

LiteLLMGateway(
    tenant_key: TenantKey | None = None,
    guardrails: GuardrailPolicy | None = None,
    tags: Sequence[str] = (),
)

Bases: AbstractCapability[Session[Any]]

Attach tenancy, the session, the trace, a tenant's key and guardrails to each request.

Give it to an agent beside ArtifactWorkspace, with a litellm_model. Before each model request it adds, through the request's extra_body and extra_headers:

  • LiteLLM metadata: the tenant, workspace, thread and run, the thread as the session, the person as the user, the trace id (so LiteLLM's own traces join the turn's), and tags
  • the W3C trace context, as traceparent
  • the guardrails the workspace's policy names
  • the tenant's virtual key, as the request's Authorization, from tenant_key

A request a guardrail blocks fails the run with a GuardrailBlocked, recorded as guardrail_blocked, and is not retried.

Parameters:

Name Type Description Default
tenant_key TenantKey | None

Returns a tenant's LiteLLM key; the model's own key is used without one.

None
guardrails GuardrailPolicy | None

Returns the guardrails a workspace's requests use.

None
tags Sequence[str]

More tags for every request, such as the application's name.

()

before_model_request async

before_model_request(
    ctx: RunContext[Session[Any]],
    request_context: ModelRequestContext,
) -> ModelRequestContext

Add the gateway's metadata, guardrails and key to the request.

request_options async

request_options(
    session: Session[Any],
) -> tuple[dict[str, Any], dict[str, str]]

Return the extra body and headers of a model request in a session's run.

on_model_request_error async

on_model_request_error(
    ctx: RunContext[Session[Any]],
    *,
    request_context: ModelRequestContext,
    error: Exception,
) -> ModelResponse

Turn a guardrail's block into a typed run failure; let other errors through.

wrap_run_event_stream async

wrap_run_event_stream(
    ctx: RunContext[Session[Any]],
    *,
    stream: AsyncIterable[AgentStreamEvent],
) -> AsyncIterable[AgentStreamEvent]

Turn a guardrail's block into a typed run failure in a streamed run too.

A streamed request is sent as its events are first read, so its HTTP error surfaces here rather than in on_model_request_error.

TenantKey module-attribute

TenantKey = Callable[[TenantId], Awaitable[str | None]]

Returns the LiteLLM virtual key of a tenant's team, or None for the model's own key.

A port the application implements (ADR-0034): keys come from its secret store, and artifactr never logs them or puts them on spans.

GuardrailPolicy module-attribute

GuardrailPolicy = Callable[
    [TenantId, WorkspaceId], Awaitable[Sequence[str]]
]

Returns the names of the LiteLLM guardrails a workspace's requests use.

Guardrail blocks

GuardrailBlocked

GuardrailBlocked(guardrail: str | None)

Bases: RunFailure

A model request that one of the gateway's guardrails blocked.

The run fails with the reason guardrail_blocked and this message, which names the guardrail but never repeats what was blocked.

Parameters:

Name Type Description Default
guardrail str | None

The guardrail's name, when the gateway said.

required

guardrail_block

guardrail_block(
    error: ModelHTTPError,
) -> GuardrailBlocked | None

Return the guardrail block a model request's HTTP error reports, if it is one.

LiteLLM answers a request a guardrail blocks with HTTP 400 and an error that names the guardrail; other errors are not guardrail blocks.

GUARDRAIL_BLOCKED module-attribute

GUARDRAIL_BLOCKED = 'guardrail_blocked'

The run_ended reason of a run whose model request a guardrail blocked.