migrate from litellm

Running LiteLLM?
Point your config here.

Translate a LiteLLM config.yaml with one command, then point your OpenAI-format clients at SBproxy. The request shape and the endpoints stay the same, so your application code does not change. What you add is a real proxy in front of the traffic, guardrails that run in-stream, and a spend ledger you can verify.

tagline
Point your existing clients at a new URL and keep their code.
terminal
one command
# translate, validate, run
$ sbproxy config import-litellm config.yaml --out sb.yml
# warning: callback 'my_module.Handler' is a Python hook -> rewrite in CEL
# import-litellm: 1 key needs manual attention
$ sbproxy validate sb.yml
# sb.yml: ok
$ sbproxy sb.yml
▶ config.yaml → sb.ymlfull guide →
01 / one command

Start with
the translator.

import-litellm reads your config.yaml and writes an equivalent sb.yml. It prints a warnings report for every key that needs manual attention and never fails on an unmapped key, so you see exactly what carries over. os.environ/VAR references become ${VAR}.

S01

Translate

sbproxy config import-litellm config.yaml --out sb.yml emits the config and a warnings report.

S02

Validate

sbproxy validate sb.yml compiles the result and fails fast on anything that will not run.

S03

Point clients

Run sbproxy sb.yml and change one base URL. OpenAI-format clients need no other change.

02 / before and after

A LiteLLM config
and its translation.

Two providers behind public model names, a routing strategy, and a cache. Each model_list entry becomes a provider; two entries sharing a name become a load-balanced model group; router_settings and litellm_settings map onto routing and semantic_cache.

before  ·  litellm config.yaml
model_list:
  - model_name: gpt-4
    litellm_params:
      model: openai/gpt-4o
      api_key: os.environ/OPENAI_API_KEY
      rpm: 100
  - model_name: claude
    litellm_params:
      model: anthropic/claude-haiku-4-5
      api_key: os.environ/ANTHROPIC_API_KEY
router_settings:
  routing_strategy: latency-based-routing
litellm_settings:
  cache: true
after  ·  sb.yml
origins:
  "ai.local":
    action:
      type: ai_proxy
      routing: lowest_latency
      providers:
        - name: openai
          provider_type: openai
          api_key: ${OPENAI_API_KEY}
          models: [gpt-4]
          model_map: { gpt-4: gpt-4o }
        - name: anthropic
          provider_type: anthropic
          api_key: ${ANTHROPIC_API_KEY}
          models: [claude]
          model_map: { claude: claude-haiku-4-5 }
      model_rate_limits:
        gpt-4: { requests_per_minute: 100 }
      semantic_cache: { enabled: true }
03 / what maps

The keys you
already wrote.

The translator maps the known surface and warns on the rest. The common shape of a LiteLLM proxy carries over directly.

model_list

model_list -> providers + model groups

Each entry becomes a provider; its provider-prefixed model string splits into a provider_type and an upstream model via model_map. Two entries that share a model_name become one load-balanced model group behind that public name.

router_settings

routing_strategy -> routing

simple-shuffle maps to round_robin, latency-based-routing to lowest_latency, usage-based-routing to least_token_usage, least-busy to least_connections, cost-based-routing to cost_optimized. Retries, timeouts, and fallbacks land on routing and resilience.

rpm / tpm

per-deployment limits -> model_rate_limits

litellm_params.rpm and tpm become per-model requests_per_minute and tokens_per_minute, enforced fail-fast at the gateway before the upstream call.

litellm_settings

cache -> semantic_cache

litellm_settings.cache turns on the built-in embedding semantic cache, with a similarity threshold you control. os.environ/VAR references anywhere become ${VAR} interpolation.

guardrails

guardrails -> built-in or external adapters

A simple PII or moderation guardrail maps to the built-in detector; a pointer at a specific Presidio, Lakera, or Bedrock endpoint maps to an external guardrail adapter. The translator picks the closer target and documents the choice.

04 / beyond parity

What you get
once you switch.

A callback-based gateway adds these as config and after-the-fact hooks. SBproxy runs them inline, in the same binary that proxies the request. Each is off by default.

  • A verifiable usage ledger: hash-chained, optionally Ed25519-signed spend receipts you can re-derive and verify.
  • One sandboxed CEL policy over guardrails, budgets, routing, and principal, instead of Python file-path hooks.
  • A guardrail mesh: collect every verdict, block on a quorum or redact and continue, with a verdict cache.
  • Outcome-aware routing that scores providers by realized cost-per-success and demotes the ones that refuse or error.
  • Predictive budgets that warn, then downgrade to a cheaper model, before the hard cap.
  • LLM-aware resilience: per-error retry, context-window compression, hedged requests, content-policy fallback.
  • A live key plane: mint, rotate, and revoke virtual keys through the admin API with no reload, hashed at rest with HMAC-SHA256 and a server pepper. A revoke takes effect on the next request, not after a database round trip.
  • Mesh clustering: run a fleet and the keys, budgets, and rate counters stay coherent through a gossip mesh and CRDT counters. LiteLLM leans on an external Redis to share that state across instances; here the cluster coordinates itself.
what needs a rewrite

The parts no translator can move for you.

The importer warns on each and points at the docs.

  • 01
    Python hookscustom_auth, custom_sso, custom_key_generate, and callback classes given as module paths. The analog is sandboxed CEL, Lua, JavaScript, or WebAssembly; rewrite the logic in one of those.
  • 02
    Open-ended litellm_paramsArbitrary completion kwargs are passed through and warned on; review each warned key against the provider config.
  • 03
    External guardrail endpointsPresidio, Lakera, Aporia, and Bedrock guardrails map to a built-in detector or an external adapter; confirm the chosen target per entry.
  • 04
    master_key and key storegeneral_settings.master_key becomes explicit proxy authentication. LiteLLM's database-backed virtual-key and spend store maps to the built-in dynamic key plane: mint, rotate, and revoke at runtime, backed by the embedded store, Redis, or a secrets manager. Re-issue keys here rather than copying database rows.
who it's for
You, if
  • You run the LiteLLM proxy and want a real reverse proxy in front of the same traffic.
  • You want guardrails, budgets, and routing enforced in one binary, not bolted on as callbacks.
  • You need a spend record you can prove for chargeback or compliance.
  • You are deploying in your own VPC and want a single self-hostable process.
Carries over
OpenAI-format clientsmodel_listrouting_strategyfallbacksrpm / tpm limitscachingos.environ refs/v1/chat/completions/v1/embeddings
06 / talk to us

Want to talk
before you install?

The gateway is on GitHub under Apache 2.0 and you can install it right now, so nothing here is gated behind a sales call. Use this form for deployment questions, migration scoping, or anything you would rather not file as a public issue.

01
Share your stack. A few details on providers, scale, and the gateway you might be replacing.
02
Talk to engineering. A real conversation about whether SBproxy fits the deployment shape you have in mind. No demo theatre.
03
Get a straight answer. Including when the answer is that SBproxy is the wrong fit and you should run something else.