The LLM gateway¶
Applications route their agents through a LiteLLM proxy, which owns model routing (groups, fallbacks, load balancing), budgets, rate limits and guardrails (ADR-0022). stackr runs the proxy. The litellm extra connects reflexr to it through pydantic-ai's LiteLLM provider, with no dependency on the litellm package. This page covers the model, what each request carries, tenants' keys, and guardrails.
A model over the proxy¶
Install reflexr[litellm], then give an agent a model from litellm_model and the LiteLLMGateway capability beside EventContext:
from pydantic_ai import Agent
from reflexr.agent import AgentAction, EventContext
from reflexr.litellm import LiteLLMGateway, litellm_model
from reflexr.workspace import Reaction
triage_agent = Agent(
litellm_model("claude-sonnet", api_base="http://litellm:4000", api_key=proxy_key),
deps_type=Reaction[None],
capabilities=[
EventContext(emit=[IncidentOpened]),
LiteLLMGateway(tenant_key=tenant_key, guardrails=guardrails, tags=["oncall"]),
],
)
triage = AgentAction(triage_agent, name="triage")
litellm_model names one of the proxy's model groups; which provider and model serve it, and what happens when one fails, is the proxy's configuration. api_key is the key used for requests that carry no tenant's key. Without api_base and api_key, the provider reads its environment variables.
Its settings are the model's default settings, as every pydantic-ai model takes them: a temperature, a max_tokens, or an extra_body the proxy reads. A smoke test can have the proxy answer without calling a provider, with LiteLLM's mock_response:
from pydantic_ai.settings import ModelSettings
model = litellm_model(
"claude-sonnet",
api_base="http://litellm:4000",
settings=ModelSettings(temperature=0, extra_body={"mock_response": "Done."}),
)
The agent's and each run's model_settings override these key by key, so an extra_body there replaces the model's. The gateway adds its metadata to whichever extra_body a request ends up with.
What each request carries¶
LiteLLMGateway adds to every model request of a run attempt:
| Where | What | Used by LiteLLM for |
|---|---|---|
metadata.tenant_id, workspace_id, rule, scope, run_id |
reflexr's ids | Spend and logs by tenant, workspace and rule |
metadata.tags |
reflexr, tenant:<id>, workspace:<id>, rule:<name>, and yours |
Tag-based spend tracking and routing |
metadata.session_id |
The run's causal chain | Grouping every request one incident caused into one session |
metadata.trace_user_id |
The person whose event made the rule fire, when a person published it | Users in its Langfuse logging |
metadata.existing_trace_id, and the traceparent header |
The run attempt's trace, when telemetry is configured | Joining the attempt's trace in Langfuse and OpenTelemetry |
guardrails |
The rule's guardrails in this workspace | Which guardrails check the request |
Authorization |
The tenant's key | Budgets, rate limits and allowed models per tenant |
The trace fields are there only when a tracer provider records spans (Observability); without one there is no trace to join.
Tenants and keys¶
Each tenant is a LiteLLM team with its own virtual keys, budgets and rate limits. The gateway asks your application for a tenant's key on each request, through tenant_key, a port your application implements. Keep keys in your secret store:
async def tenant_key(tenant_id: str) -> str | None:
return await secrets.get(f"litellm/{tenant_id}") # None: use the model's own key
reflexr never records keys: not in the log, not on spans. Keep HTTP header capture off in your OpenTelemetry instrumentation, which is its default.
Guardrails¶
A rule's policy names the guardrails its requests use, as the proxy configures them. The policy is called with the tenant, the workspace and the rule, so an application can guard some workflows more than others:
async def guardrails(tenant_id: str, workspace_id: str, rule: str) -> list[str]:
if rule == "ops:error-spike" and await is_regulated(tenant_id):
return ["presidio-pii", "prompt-injection"]
return []
When a guardrail blocks a request, the proxy answers HTTP 400. The gateway turns that into GuardrailBlocked, a permanent RunFailure (ADR-0036), so the run is dead-lettered at once with the reason guardrail_blocked, whatever the rule's retry policy allows: retrying the same input would be blocked again. The message names the guardrail but never repeats what was blocked:
{"type": "reflexr:run_dead_lettered", "run_id": "fir_823258f37c0bddf1", "rule": "ops:error-spike",
"attempts": 1, "error": "The model request was blocked by the presidio-pii guardrail.",
"reason": "guardrail_blocked"}
The reflexr.runs metric counts it with reflexr.run.reason="guardrail_blocked", so a dashboard shows guardrail blocks beside other failures. A person can still retry_run it after changing the policy or the input (Workspaces and the log).
Your own actions and tools can fail a run the same way. Raise RunFailure with a stable reason code, and permanent=True when retrying cannot help (The reactor).
Testing¶
The gateway works with any pydantic-ai model, so the patterns in Testing your application apply. To test against the proxy's wire format without a proxy, give litellm_model an httpx2.AsyncClient over an httpx2.MockTransport that answers chat completions, and read the requests it received, as reflexr's own tests do in tests/litellm/test_gateway.py:
import httpx2
client = httpx2.AsyncClient(transport=httpx2.MockTransport(proxy.handle))
model = litellm_model("claude-sonnet", api_base="http://litellm.test", http_client=client)
A handler that answers HTTP 400 with a body naming a guardrail reproduces a block.