Translate
sbproxy config import-litellm config.yaml --out sb.yml emits the config and a warnings report.
Translate a LiteLLM config.yaml with one command, then point your OpenAI-format clients at SBproxy. The request shape and the endpoints stay the same, so your application code does not change. What you add is a real proxy in front of the traffic, guardrails that run in-stream, and a spend ledger you can verify.
# translate, validate, run
$ sbproxy config import-litellm config.yaml --out sb.yml
# warning: callback 'my_module.Handler' is a Python hook -> rewrite in CEL
# import-litellm: 1 key needs manual attention
$ sbproxy validate sb.yml
# sb.yml: ok
$ sbproxy sb.ymlimport-litellm reads your config.yaml and writes an equivalent sb.yml. It prints a warnings report for every key that needs manual attention and never fails on an unmapped key, so you see exactly what carries over. os.environ/VAR references become ${VAR}.
sbproxy config import-litellm config.yaml --out sb.yml emits the config and a warnings report.
sbproxy validate sb.yml compiles the result and fails fast on anything that will not run.
Run sbproxy sb.yml and change one base URL. OpenAI-format clients need no other change.
Two providers behind public model names, a routing strategy, and a cache. Each model_list entry becomes a provider; two entries sharing a name become a load-balanced model group; router_settings and litellm_settings map onto routing and semantic_cache.
model_list:
- model_name: gpt-4
litellm_params:
model: openai/gpt-4o
api_key: os.environ/OPENAI_API_KEY
rpm: 100
- model_name: claude
litellm_params:
model: anthropic/claude-haiku-4-5
api_key: os.environ/ANTHROPIC_API_KEY
router_settings:
routing_strategy: latency-based-routing
litellm_settings:
cache: trueorigins:
"ai.local":
action:
type: ai_proxy
routing: lowest_latency
providers:
- name: openai
provider_type: openai
api_key: ${OPENAI_API_KEY}
models: [gpt-4]
model_map: { gpt-4: gpt-4o }
- name: anthropic
provider_type: anthropic
api_key: ${ANTHROPIC_API_KEY}
models: [claude]
model_map: { claude: claude-haiku-4-5 }
model_rate_limits:
gpt-4: { requests_per_minute: 100 }
semantic_cache: { enabled: true }The translator maps the known surface and warns on the rest. The common shape of a LiteLLM proxy carries over directly.
Each entry becomes a provider; its provider-prefixed model string splits into a provider_type and an upstream model via model_map. Two entries that share a model_name become one load-balanced model group behind that public name.
simple-shuffle maps to round_robin, latency-based-routing to lowest_latency, usage-based-routing to least_token_usage, least-busy to least_connections, cost-based-routing to cost_optimized. Retries, timeouts, and fallbacks land on routing and resilience.
litellm_params.rpm and tpm become per-model requests_per_minute and tokens_per_minute, enforced fail-fast at the gateway before the upstream call.
litellm_settings.cache turns on the built-in embedding semantic cache, with a similarity threshold you control. os.environ/VAR references anywhere become ${VAR} interpolation.
A simple PII or moderation guardrail maps to the built-in detector; a pointer at a specific Presidio, Lakera, or Bedrock endpoint maps to an external guardrail adapter. The translator picks the closer target and documents the choice.
A callback-based gateway adds these as config and after-the-fact hooks. SBproxy runs them inline, in the same binary that proxies the request. Each is off by default.
The importer warns on each and points at the docs.
The gateway is on GitHub under Apache 2.0 and you can install it right now, so nothing here is gated behind a sales call. Use this form for deployment questions, migration scoping, or anything you would rather not file as a public issue.