Semantic model routing
Run vLLM Semantic Router as a managed Docker sidecar and route supported proxy traffic between OpenAI-compatible model backends.
Semantic routing lets the DefenseClaw gateway choose a configured model backend after prompt inspection and before the upstream request. DefenseClaw starts a version-pinned vLLM Semantic Router container, asks it to classify each supported request, resolves the returned model alias locally, and then forwards through the normal guardrail proxy path.
Proxy connectors only
Automatic routing applies to OpenClaw and ZeptoClaw requests that
reach the proxy as OpenAI-compatible */chat/completions traffic. Hook-based
connectors (Claude Code, Codex, Cursor, Devin, GitHub Copilot CLI, OpenHands,
Antigravity, Hermes, OpenCode, Amp, OmniGent and Kiro) keep their existing
provider path and are not rerouted.
Compatibility at a glance
| Traffic path | Semantic routing | Guardrails |
|---|---|---|
| OpenClaw or ZeptoClaw → OpenAI Chat Completions proxy | Yes | Pre-call and post-call |
| Hook-based connector → hook API | No | Existing hook enforcement |
OpenAI Responses API (/v1/responses) | No | Existing proxy behavior |
Anthropic Messages API (/v1/messages) | No | Existing proxy behavior |
| Native Gemini or Bedrock endpoints | No | Existing proxy behavior |
Routing fails open to the connector's original upstream: if the router is unavailable, times out, or returns an unknown model alias, DefenseClaw logs the routing failure and continues with the original provider selection. Guardrail enforcement does not fail open with it.
Prerequisites
- OpenClaw or ZeptoClaw is configured in proxy mode.
- For a managed router, Docker Desktop or Docker Engine is installed and running Linux containers. A remote router does not require local Docker.
- Every routed backend exposes an OpenAI-compatible Chat Completions endpoint.
- Each
base_urlis an origin such ashttps://api.openai.com, without/v1/chat/completionsappended. - Provider credentials are stored in process environment variables or
~/.defenseclaw/.env; never put key values inconfig.yaml.
Host support
| Mode | Host requirement | Runtime architecture |
|---|---|---|
| Managed sidecar | macOS or Linux with Docker running Linux containers | Pinned image includes linux/amd64 and linux/arm64 |
| Remote router | macOS or Linux host that can reach the configured HTTPS endpoint | Router is operator-owned |
| Native Windows | Not supported | Docker Desktop, WSL, a VM, or a remote classifier does not add the missing native proxy-connector lifecycle |
DefenseClaw is hook-only on native Windows. The only connectors eligible for semantic routing—OpenClaw and ZeptoClaw—both require the guardrail proxy and are unsupported there. Semantic routing is a macOS and Linux feature, even when the classifier runs remotely.
The managed sidecar image runs on amd64 and arm64 hosts, including Apple silicon Macs with Docker Desktop.
Configure routing
Activation command, not a model wizard
defenseclaw setup routing --enable validates and activates an existing routing
catalog; it does not invent provider or model choices. Add at least one model and
one fallback decision to config.yaml first, then run the command. This keeps
model names, endpoints, and credential references explicit and reviewable.
Add the routing block to ~/.defenseclaw/config.yaml. Model name values are the
aliases returned by the semantic router; model, base_url, and api_key_env are
resolved inside DefenseClaw and are not sent to the classifier as credentials.
routing:
enabled: true
version: "0.3.0"
port: 8080
algorithm: static
models:
- name: fast
provider: openai
model: gpt-4o-mini
base_url: https://api.openai.com
api_key_env: OPENAI_API_KEY
- name: deep
provider: openai
model: gpt-4.1
base_url: https://api.openai.com
api_key_env: OPENAI_API_KEY
signals:
keywords:
- name: code-request
keywords: ["debug", "implement", "refactor"]
operator: OR
decisions:
- name: code
priority: 100
conditions:
- type: keyword
name: code-request
model_refs: [deep]
- name: default
priority: 1
model_refs: [fast]Then enable the managed sidecar:
defenseclaw setup routing --enable
defenseclaw-gateway restartThe setup command checks Docker, saves the routing state, and tells the gateway to
start the pinned router image on its next launch. The generated router configuration
lives under ~/.defenseclaw/semantic-router/; it is runtime state, not a file you
need to maintain.
Use a remote router
To use a router you run yourself, set routing.remote.endpoint. The gateway then
sends classifier requests there and does not start a local container:
routing:
enabled: true
remote:
endpoint: https://router.example.com
timeout_ms: 1500 # 0 uses the default; the maximum is 5000
models:
# ... same models, signals and decisions as aboveNon-loopback endpoints must use HTTPS.
Store routed-model keys with DefenseClaw
Every non-empty routing.models[].api_key_env is part of DefenseClaw's normal key
registry. When routing is enabled, keys list marks it as required; keyless local
models such as Ollama do not create a credential requirement.
defenseclaw keys list --missing-only
defenseclaw keys set OPENAI_API_KEYThe gateway resolves the selected backend's key only after classification. Secret values and real provider URLs are not written into the generated router catalog or sent in the classifier request. See Keys and credentials for storage precedence and automation-safe commands.
Prompt privacy boundary
The classifier receives message roles and content because it must evaluate the request. In managed mode that traffic stays on the loopback-bound Docker API. A configured remote classifier receives the same prompt content, so DefenseClaw requires HTTPS for non-local endpoints and does not follow redirects. API keys, request headers, tools, user IDs, session IDs, and gateway policy metadata are not sent to the v0.3 classifier.
Verify the running state
defenseclaw setup routing --status
docker ps --filter label=com.defenseclaw.component=semantic-routerSend a normal request through OpenClaw or ZeptoClaw. A routed response includes:
X-Semantic-Router: routed
X-Semantic-Router-Reason: decision=… model=… confidence=…Each guarded request also emits one request-correlated v8 routing decision
(defenseclaw.semantic_routing.decision plus latency) on the same trace as the
model operation. Labels are a closed set: applied or fallback, a failure
code (none, timeout, upstream_status, decode_failure, unknown_alias,
credential_failure, configuration_failure), and whether an override was
applied. Prompt text, credentials, endpoints, and the header reason string are
not metric labels.
No separate public chat endpoint is exposed on the router port. Port 8080 is the
loopback classifier API used by the gateway; clients should continue using their
normal DefenseClaw connector endpoint. The container uses a read-only config mount,
a read-only root filesystem, no Linux capabilities, and a digest-pinned multi-platform
v0.3.0 image.
Disable routing
defenseclaw setup routing --disable
defenseclaw-gateway restartDisabling routing leaves connector and guardrail configuration intact. Requests use their original provider again after the restart.
Troubleshooting
| Symptom | Check |
|---|---|
| Router is enabled but no container appears | Run docker info, then restart the gateway. |
| Requests never show the routing response header | Confirm the connector is OpenClaw or ZeptoClaw and the path ends in /chat/completions. |
Gateway logs unknown model alias | Ensure the router's returned alias exactly matches a routing.models[].name. |
| Routed provider returns authentication errors | Check the named api_key_env with defenseclaw keys list; do not copy the secret into YAML. |
| Router startup fails after editing YAML | Validate the configured port, unique model names, decision model_refs, and Docker logs for the labeled container. |
Next steps
- Unified LLM key for the provider keys routed models use.
- OpenClaw and ZeptoClaw, the connectors that support routing.
Disabling guardrail
defenseclaw setup guardrail --disable is the global rollback. Connector hooks are removed or restored from the hash-checked backup, and agents run without DefenseClaw until you turn the guardrail back on.
Unified LLM key
Wire up DEFENSECLAW_LLM_KEY — the single environment variable that powers the LLM judge, the MCP / skill / plugin scanners, and any custom LLM call DefenseClaw makes through Bifrost.