open source · apache 2.0 · v1.14.0→

Take control of
your AI traffic.

Your APIs, your MCP tools, and every model your teams call, behind one binary you run yourself. In your VPC or air-gapped, with no vendor control plane in the path.

$brew install soapbucket/tap/sbproxy
tls + wafmcp federation200+ modelsguardrails + budgetsyour gpus
sbproxy · request loglive
$ sbproxy --config sb.yml
✓ listener :443 · tls issued via acme · 3 nodes joined
→ POST /v1/chat/completions key=team-web
auth pass virtual key · scope chat
guardrail pass pii clean · injection clean
route anthropic · claude-sonnet-4-5 · 212ms
→ POST /v1/chat/completions key=team-batch
budget block workspace cap $150/day · 402
→ GET /articles/2026 agent=gptbot
crawl price unpaid crawler → 402 · rsl manifest served
↻ failover openrouter → anthropic p50 1.2s · degraded
works with the stack you already run
OpenAI SDKAnthropic SDKVercel AI SDKLangChainvLLMllama.cppcurl

Every request walks the same line.

REST call, tool invocation, or completion, SBproxy terminates the TLS, checks the caller, applies the policy, then routes it. One policy graph and one audit log for all three kinds of traffic, so there is no second tool where the rules can drift.

See how routing works →
Terminate TLSdone · 0.2ms
certs issued via acme, rotated in place
Authenticatedone · key ok
virtual key team-web · scope chat
Policy + guardrailschecking
pii · injection · budget · rate limit
Route + failoverqueued
anthropic → openrouter → local-vllm

One runtime for all of it.

Most teams end up with a proxy for the APIs, a second thing for MCP, and a third for the model calls. Three configs, three sets of credentials, three places the policy can be wrong. These are the same problem wearing different protocols.

api

The APIs you already run

Put an existing HTTP service behind policy without adding a hop, because SBproxy is the proxy. TLS it issues itself from ACME, a CRS-derived baseline WAF, fifteen auth providers from API keys through JWT to full OIDC, load balancing with health checks, transforms, and caching.

tls + acmewafoidcgrpcopenapi
mcp

The MCP tools agents reach for

Federate internal MCP servers behind one endpoint with OAuth2, DPoP, and PKCE. Scope which tools each agent can see, meter every call into a session ledger, and re-export an OpenAPI spec you already have as MCP tools without writing a server.

federationoauth2 + dpopper-agent rbacrest → mcp
ai

The models your teams call

Two hundred models across 70 providers on one OpenAI-compatible endpoint, on your keys at your prices. Or point a serve: block at your own GPU and SBproxy fits an engine to the card and governs it with the same keys, budgets, and guardrails.

70 providers18 strategiesguardrailsbudgetsyour gpus

Change one line. Reach 200+ models.

Keep the SDK you already use. Point its base_url at SBproxy and call any model by name, across 70 providers or your own GPUs. The request shape never changes, and your provider keys stay on your side.

OpenAI SDKAnthropic SDKVercel AI SDKLangChaincurl
client.py
client = OpenAI(
    # your SBproxy, in your VPC
    base_url="https://sbproxy.acme.internal/v1",
    api_key=os.environ["SB_KEY"],
)

# any of 200+ models, by name
client.chat.completions.create(
    model="claude-sonnet-4-5",
    messages=[...],
)
# provider down? SBproxy fails over.

EVERY POLICY.
ONE FILE.

Model routing, failover, budgets, caching, guardrails, MCP exposure, and crawler policy declare in the same YAML. Versioned in git, reviewed like code, hot-reloaded on save.

sb.yml1 file · your entire traffic layer
# the ai you call
action:
  type: ai_proxy
  providers:
    - name: anthropic
      api_key: ${ANTHROPIC_API_KEY}
    - name: local-vllm
      base_url: http://gpu-01:8000/v1
  routing:
    strategy: fallback_chain
  guardrails:
    input:
      - type: pii
      - type: injection
  budget:
    limits:
      - scope: workspace
        max_cost_usd: 150
        period: daily
# the ai that calls you
action:
  type: mcp
  federated_servers:
    - type: openapi
      origin: https://api.acme.com
      spec_path: openapi.yml
  oauth:
    authorization_servers:
      - https://auth.acme.com
policies:
  - type: ai_crawl_control
    price: 0.001
    currency: USD
  - type: rate_limit
    requests_per_minute: 600

The admin console ships in the gateway.

Turn on the admin block in sb.yml and keys, spend, provider health, and live traffic are already wired. The console is a URL on the gateway itself, so there is nothing extra to deploy or buy.

sbproxy.acme.internal/adminlive · 3 nodes
overview
keys
credentials
config
logs
metrics
spend
playground
requests / min3,182
spend today$41.20 / $150 cap
blocked by policy37
providerstatusp50routing
anthropichealthy212msprimary
local-vllmhealthy89msllama-3.3 · qwen3
openaihealthy340msfallback
openrouterdegraded1.2sfailing over
▲ prompt-injection blocked · key team-web402 unpaid crawler · gptbot↻ failover openrouter → anthropic
50,713 rpsfull proxy chain · 0.6ms p99
185K rpsWAF rejection · 0.17ms p99
70 providers200+ models · one catalog
Apache 2.0self-hosted · air-gap ready

Route your first request in five minutes.

01

Install the binary

One static Rust binary. Homebrew, Docker, or a straight download.

$ brew install soapbucket/tap/sbproxy
02

Write sb.yml

Providers, routing, budgets, and guardrails in one file you review like code.

$ sbproxy --config sb.yml
03

Point your client at it

Change the base URL. Every model in the catalog answers on one endpoint.

base_url="https://sbproxy.internal/v1"