Skip to content

LiteLLM configuration

The gateway's configuration is deploy/litellm/config.yaml, mounted read-only into the litellm service (The LLM gateway). A value os.environ/NAME is read from the service's environment, which compose.yaml sets from .env; make validate checks that every one of them is set there, and that no key is written into the file.

Models

Applications name a model group or an alias from the first column, never a provider's model.

Model name Routes to Key API base
claude-sonnet anthropic/claude-sonnet-4-6 ANTHROPIC_API_KEY the provider's
claude-opus anthropic/claude-opus-4-7 ANTHROPIC_API_KEY the provider's
claude-haiku anthropic/claude-haiku-4-5-20251001 ANTHROPIC_API_KEY the provider's
gpt-4o openai/gpt-4o OPENAI_API_KEY the provider's
gpt-4o-mini openai/gpt-4o-mini OPENAI_API_KEY the provider's
gemini-2.5-pro gemini/gemini-2.5-pro GEMINI_API_KEY the provider's
gemini-2.5-flash gemini/gemini-2.5-flash GEMINI_API_KEY the provider's
openrouter/* openrouter/* OPENROUTER_API_KEY the provider's
lmstudio lm_studio/local-model a placeholder LM_STUDIO_API_BASE
omlx/* openai/* a placeholder OMLX_API_BASE
default anthropic/claude-sonnet-4-6 ANTHROPIC_API_KEY the provider's
fast anthropic/claude-haiku-4-5-20251001 ANTHROPIC_API_KEY the provider's

A model whose key is empty in .env fails, and its group falls back. LM Studio and oMLX need no key; the file gives them placeholders.

Fallbacks

Model group Falls back to, in order
default gpt-4o then gemini-2.5-pro
fast gpt-4o-mini then gemini-2.5-flash

Guardrails

Requests choose guardrails with "guardrails": [...] in their body.

Guardrail Integration Mode On by default Checks
pii-mask litellm_content_filter pre_call no email (mask), us_phone (mask), us_ssn (mask), visa (mask), mastercard (mask), amex (mask), aws_access_key (mask), aws_secret_key (mask), github_token (mask)
prompt-injection litellm_content_filter pre_call no prompt_injection_jailbreak (block at medium), prompt_injection_system_prompt (block at medium), prompt_injection_data_exfiltration (block at medium)

Settings

Setting Value
router_settings.routing_strategy simple-shuffle
router_settings.num_retries 2
router_settings.allowed_fails 3
router_settings.cooldown_time 30
router_settings.redis_url os.environ/REDIS_URL
litellm_settings.drop_params true
litellm_settings.request_timeout 600
litellm_settings.cache true
litellm_settings.cache_params.type redis
litellm_settings.cache_params.redis_url os.environ/REDIS_URL
litellm_settings.cache_params.ttl 600
callback_settings.otel.attributes.exclude_list [metadata.user_api_key_hash, metadata.user_api_key_alias, hidden_params]
general_settings.master_key os.environ/LITELLM_MASTER_KEY
general_settings.database_url os.environ/DATABASE_URL

REDIS_URL is database 1 of the shared Valkey, and DATABASE_URL the litellm database on the database adapter, both set in compose.yaml.

The file

The whole file: deploy/litellm/config.yaml
# The LiteLLM proxy: the LLM gateway port (ADR-0005, ADR-0010). Applications
# call its OpenAI-compatible API with a team's virtual key and a model name
# from this file, never a provider's; the providers are adapters behind it.
# Provider keys come from .env; a missing one fails only the requests that
# need it.

model_list:
  # --- aliases: one provider model each ---------------------------------------

  # Anthropic (Console API key)
  - model_name: claude-sonnet
    litellm_params:
      model: anthropic/claude-sonnet-4-6
      api_key: os.environ/ANTHROPIC_API_KEY
  - model_name: claude-opus
    litellm_params:
      model: anthropic/claude-opus-4-7
      api_key: os.environ/ANTHROPIC_API_KEY
  - model_name: claude-haiku
    litellm_params:
      model: anthropic/claude-haiku-4-5-20251001
      api_key: os.environ/ANTHROPIC_API_KEY

  # OpenAI
  - model_name: gpt-4o
    litellm_params:
      model: openai/gpt-4o
      api_key: os.environ/OPENAI_API_KEY
  - model_name: gpt-4o-mini
    litellm_params:
      model: openai/gpt-4o-mini
      api_key: os.environ/OPENAI_API_KEY

  # Google Gemini (AI Studio key)
  - model_name: gemini-2.5-pro
    litellm_params:
      model: gemini/gemini-2.5-pro
      api_key: os.environ/GEMINI_API_KEY
  - model_name: gemini-2.5-flash
    litellm_params:
      model: gemini/gemini-2.5-flash
      api_key: os.environ/GEMINI_API_KEY

  # OpenRouter: any of its models, as openrouter/<vendor>/<model>
  - model_name: openrouter/*
    litellm_params:
      model: openrouter/*
      api_key: os.environ/OPENROUTER_API_KEY

  # Local model servers on the host, reached from the container through
  # host.docker.internal: LM Studio, and oMLX (any of its models, as
  # omlx/<model>).
  - model_name: lmstudio
    litellm_params:
      model: lm_studio/local-model
      api_base: os.environ/LM_STUDIO_API_BASE
      api_key: lm-studio
  - model_name: omlx/*
    litellm_params:
      model: openai/*
      api_base: os.environ/OMLX_API_BASE
      api_key: omlx-local

  # --- model groups: what applications ask for --------------------------------
  # A group can hold several deployments, which the router balances; on
  # errors it falls back to the groups in router_settings.fallbacks.

  - model_name: default
    litellm_params:
      model: anthropic/claude-sonnet-4-6
      api_key: os.environ/ANTHROPIC_API_KEY
  - model_name: fast
    litellm_params:
      model: anthropic/claude-haiku-4-5-20251001
      api_key: os.environ/ANTHROPIC_API_KEY

router_settings:
  routing_strategy: simple-shuffle
  num_retries: 2
  allowed_fails: 3
  cooldown_time: 30
  # Each group falls back across providers, in order.
  fallbacks:
    - default: [gpt-4o, gemini-2.5-pro]
    - fast: [gpt-4o-mini, gemini-2.5-flash]
  # Routing state (cooldowns, usage) shared through the Redis protocol,
  # database 1.
  redis_url: os.environ/REDIS_URL

# Guardrails, defined once and selected per request with
# `"guardrails": ["pii-mask"]` in the body: the libraries' [litellm] extras
# choose them from each workspace's or rule's policy. Both run in the proxy,
# without another service.
guardrails:
  - guardrail_name: pii-mask
    litellm_params:
      guardrail: litellm_content_filter
      mode: pre_call
      default_on: false
      patterns:
        - {pattern_type: prebuilt, pattern_name: email, action: MASK}
        - {pattern_type: prebuilt, pattern_name: us_phone, action: MASK}
        - {pattern_type: prebuilt, pattern_name: us_ssn, action: MASK}
        - {pattern_type: prebuilt, pattern_name: visa, action: MASK}
        - {pattern_type: prebuilt, pattern_name: mastercard, action: MASK}
        - {pattern_type: prebuilt, pattern_name: amex, action: MASK}
        - {pattern_type: prebuilt, pattern_name: aws_access_key, action: MASK}
        - {pattern_type: prebuilt, pattern_name: aws_secret_key, action: MASK}
        - {pattern_type: prebuilt, pattern_name: github_token, action: MASK}
  - guardrail_name: prompt-injection
    litellm_params:
      guardrail: litellm_content_filter
      mode: pre_call
      default_on: false
      categories:
        - {category: prompt_injection_jailbreak, enabled: true, action: BLOCK, severity_threshold: medium}
        - {category: prompt_injection_system_prompt, enabled: true, action: BLOCK, severity_threshold: medium}
        - {category: prompt_injection_data_exfiltration, enabled: true, action: BLOCK, severity_threshold: medium}

litellm_settings:
  drop_params: true
  request_timeout: 600
  # Responses cached in the Redis protocol's database 1, for 10 minutes.
  cache: true
  cache_params:
    type: redis
    redis_url: os.environ/REDIS_URL
    ttl: 600

# The gateway's metrics (token usage, cost, durations) keep the model, the
# provider and the tenant's team, but not per-key or per-deployment ids, which
# would make a series per key.
callback_settings:
  otel:
    attributes:
      exclude_list: [metadata.user_api_key_hash, metadata.user_api_key_alias, hidden_params]

general_settings:
  master_key: os.environ/LITELLM_MASTER_KEY
  database_url: os.environ/DATABASE_URL