Skip to content

The LLM gateway

Applications route their agents through a LiteLLM proxy, which owns model routing (groups, fallbacks, load balancing), budgets, rate limits and guardrails (ADR-0031). stackr runs the proxy. The litellm extra connects artifactr to it through pydantic-ai's LiteLLM provider, with no dependency on the litellm package (ADR-0043).

A model over the proxy

from artifactr.litellm import LiteLLMGateway, litellm_model

agent = Agent(
    litellm_model("claude-sonnet", api_base="http://litellm:4000", api_key=proxy_key),
    deps_type=Session[AppDeps],
    capabilities=[
        ArtifactWorkspace(types=[Doc, Plan]),
        LiteLLMGateway(tenant_key=tenant_key, guardrails=guardrails, tags=["docplan"]),
    ],
)

litellm_model names one of the proxy's model groups; which provider and model serve it, and what happens when one fails, is the proxy's configuration.

Its settings are the model's default settings, as every pydantic-ai model takes them: a temperature, a max_tokens, or an extra_body the proxy reads. A smoke test can have the proxy answer without calling a provider, with LiteLLM's mock_response:

model = litellm_model(
    "claude-sonnet",
    api_base="http://litellm:4000",
    settings=ModelSettings(temperature=0, extra_body={"mock_response": "Done."}),
)

The agent's and each run's model_settings override these key by key, so an extra_body there replaces the model's. The gateway adds its metadata to whichever extra_body a request ends up with.

What each request carries

LiteLLMGateway adds to every model request of a run:

Where What Used by LiteLLM for
metadata.tenant_id, workspace_id, thread_id, run_id artifactr's ids Spend and logs by tenant and workspace
metadata.tags artifactr, tenant:<id>, workspace:<id>, and yours Tag-based spend tracking and routing
metadata.session_id, trace_user_id The thread, and the person who asked Sessions and users in its Langfuse logging
metadata.existing_trace_id, and the traceparent header The turn's trace Joining the turn's trace in Langfuse and OpenTelemetry
guardrails The workspace's guardrails Which guardrails check the request
Authorization The tenant's key Budgets, rate limits and allowed models per tenant

Tenants and keys

Each tenant is a LiteLLM team with its own virtual keys, budgets and rate limits. The gateway asks your application for a tenant's key on each request; keep keys in your secret store:

async def tenant_key(tenant_id: str) -> str | None:
    return await secrets.get(f"litellm/{tenant_id}")  # None: use the model's own key

artifactr never records keys: not in the log, not on spans. Keep HTTP header capture off in your OpenTelemetry instrumentation, which is its default.

Guardrails

A workspace's policy names the guardrails its requests use, as the proxy configures them:

async def guardrails(tenant_id: str, workspace_id: str) -> list[str]:
    return ["presidio-pii", "prompt-injection"] if await is_regulated(tenant_id) else []

When a guardrail blocks a request, the proxy answers HTTP 400. The gateway turns that into a GuardrailBlocked error, a RunFailure (ADR-0042). The run ends failed with the reason guardrail_blocked, and a message that names the guardrail but never repeats what was blocked:

{"type": "run_ended", "status": "failed", "reason": "guardrail_blocked",
 "error": "The model request was blocked by the presidio-pii guardrail."}

The request is not retried, and the artifactr.runs metric counts it by reason, so the Agent and LLM dashboard shows guardrail blocks beside other failures.

Testing

The gateway works with any pydantic-ai model, so the patterns in Testing your application apply. To test against the proxy's wire format without a proxy, give litellm_model an httpx2.AsyncClient over an httpx2.MockTransport that answers chat completions, as artifactr's own tests do.