reflexr.litellm¶
The litellm extra. See The LLM gateway.
LiteLLM as reflexr's model gateway: the [litellm] extra (ADR-0022).
An adapter (ADR-0025). The model is pydantic-ai's port: litellm_model is a model that
calls a LiteLLM proxy through pydantic-ai's LiteLLMProvider, and LiteLLMGateway
is a capability that attaches the tenant, the causal chain, the trace, a tenant's key and the
rule's guardrails to every request::
agent = Agent(
litellm_model("claude-sonnet", api_base="http://litellm:4000"),
deps_type=Reaction[AppDeps],
capabilities=[
EventContext(emit=[IncidentOpened]),
LiteLLMGateway(tenant_key=keys.for_tenant, guardrails=policy.for_rule),
],
)
Routing, fallbacks, budgets, rate limits and guardrails stay in the proxy's configuration.
The model¶
litellm_model
¶
litellm_model(
model: str,
*,
api_base: str | None = None,
api_key: str | None = None,
http_client: AsyncClient | None = None,
settings: ModelSettings | None = None,
) -> OpenAIChatModel
Return a pydantic-ai model that calls a LiteLLM proxy's model group.
Routing, fallbacks, budgets and guardrails are the proxy's configuration; the model only names the group.
Shared verbatim with artifactr's src/artifactr/litellm/gateway.py; change both.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
model
|
str
|
The proxy's model group, such as |
required |
api_base
|
str | None
|
The proxy's URL, such as |
None
|
api_key
|
str | None
|
The proxy key used when a request carries no tenant's key. |
None
|
http_client
|
AsyncClient | None
|
The HTTP client, for tests or a shared connection pool. |
None
|
settings
|
ModelSettings | None
|
The model's default settings, as every pydantic-ai model takes them, such as
a |
None
|
The gateway¶
LiteLLMGateway
dataclass
¶
LiteLLMGateway(
tenant_key: TenantKey | None = None,
guardrails: GuardrailPolicy | None = None,
tags: Sequence[str] = (),
)
Bases: AbstractCapability[Reaction[Any]]
Attach tenancy, the chain, the trace, a tenant's key and guardrails to each request.
Give it to an agent beside EventContext, with a litellm_model. Before each
model request it adds, through the request's extra_body and extra_headers:
- LiteLLM metadata: the tenant, workspace, rule, scope and run, the causal chain as the session, the person whose event fired the rule as the user, the trace id (so LiteLLM's own traces join the run's), and tags
- the W3C trace context, as
traceparent - the guardrails the rule's policy names
- the tenant's virtual key, as the request's
Authorization, fromtenant_key
A request a guardrail blocks fails the attempt with a GuardrailBlocked: the run is
dead-lettered as guardrail_blocked, never retried.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
tenant_key
|
TenantKey | None
|
Returns a tenant's LiteLLM key; the model's own key is used without one. |
None
|
guardrails
|
GuardrailPolicy | None
|
Returns the guardrails a rule's requests use in a workspace. |
None
|
tags
|
Sequence[str]
|
More tags for every request, such as the application's name. |
()
|
before_model_request
async
¶
before_model_request(
ctx: RunContext[Reaction[Any]],
request_context: ModelRequestContext,
) -> ModelRequestContext
Add the gateway's metadata, guardrails and key to the request.
request_options
async
¶
Return the extra body and headers of a model request in a run attempt.
on_model_request_error
async
¶
on_model_request_error(
ctx: RunContext[Reaction[Any]],
*,
request_context: ModelRequestContext,
error: Exception,
) -> ModelResponse
Turn a guardrail's block into a typed run failure; let other errors through.
wrap_run_event_stream
async
¶
wrap_run_event_stream(
ctx: RunContext[Reaction[Any]],
*,
stream: AsyncIterable[AgentStreamEvent],
) -> AsyncIterable[AgentStreamEvent]
Turn a guardrail's block into a typed run failure in a streamed run too.
A streamed request is sent as its events are first read, so its HTTP error surfaces
here rather than in on_model_request_error.
TenantKey
module-attribute
¶
Returns the LiteLLM virtual key of a tenant's team, or None for the model's own key.
A port the application implements (ADR-0025): keys come from its secret store, and reflexr never logs them or puts them on spans.
GuardrailPolicy
module-attribute
¶
Returns the names of the LiteLLM guardrails a rule's requests use in a workspace.
Guardrail blocks¶
GuardrailBlocked
¶
GuardrailBlocked(guardrail: str | None)
Bases: RunFailure
A model request that one of the gateway's guardrails blocked.
The run is dead-lettered with the reason guardrail_blocked and this message, which names
the guardrail but never repeats what was blocked. It is permanent: retrying the same input
would be blocked again.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
guardrail
|
str | None
|
The guardrail's name, when the gateway said. |
required |
guardrail_block
¶
guardrail_block(
error: ModelHTTPError,
) -> GuardrailBlocked | None
Return the guardrail block a model request's HTTP error reports, if it is one.
LiteLLM answers a request a guardrail blocks with HTTP 400 and an error that names the guardrail; other errors are not guardrail blocks.
GUARDRAIL_BLOCKED
module-attribute
¶
The failure reason of a run attempt whose model request a guardrail blocked.