Set up the guardrail
The central command. Configures how DefenseClaw guards your agent (connector, observe or action mode, scanner, judge, human approval and rule pack), then restarts the gateway.
defenseclaw setup guardrail is the central operator command. It picks the
connector, the mode, the scanner backend, the optional LLM judge and the
human-in-the-loop (HITL) posture. Connector setup aliases such as
defenseclaw setup claude-code configure one connector with a smaller flag
set; other setup commands manage webhooks, observability, scanners and
providers on their own.
Run it interactively the first time: the wizard explains each choice and picks safe defaults. Once you are happy with the configuration, re-run it with explicit flags and --non-interactive for unattended setups and CI.
Watch the interactive flow
The animation replays a real defenseclaw setup guardrail session on a host
where no connector was configured yet: connector pick, observe or action, hook
fail mode, scanner engine, the judge step and the summary. In observe mode the
wizard skips the human-approval and judge questions, because neither applies
until a connector enforces. Hover or press Pause to study a frame; press
Restart to replay.
When a connector is already configured, the wizard skips the connector picker
and says that it is editing the global guardrail policy for the configured
connectors. Adding or switching a connector is the job of
defenseclaw setup <connector>; see Multi-connector.
Same setup, no prompts
Every choice the wizard makes has a flag, apart from the two noted below. The CI-friendly equivalent of the session above is one command:
defenseclaw setup guardrail \
--non-interactive \
--connector claudecode \
--mode observe \
--scanner-mode local \
--detection-strategy regex_only \
--restart--non-interactive (aliases --accept-defaults and --yes) skips every
prompt; missing flags fall back to the defaults the wizard would have offered.
The judge stays off until you pass --judge-model. The saved detection
strategy defaults to regex_judge, but with the judge off only the regex lane
runs; --detection-strategy regex_only makes that explicit.
Don't hand-roll the flags
The Command generator builds a non-interactive defenseclaw setup guardrail invocation for any connector with all the options below (mode, scanner, judge, HITL and advanced settings) and shows validation warnings inline.
Prompt-flag mapping
Each row is one wizard prompt, with the default that Enter selects and the flag to use in scripts.
| Prompt (default) | Flag, or how to script it |
|---|---|
| Which agent framework are you using? (first setup only; suggests an agent it finds) | --connector / --agent: claudecode, codex, cursor, devin, copilot, hermes, openhands, antigravity, opencode, amp, omnigent, kiro, openclaw, zeptoclaw. |
| Enable guardrail? (Y) | Implied by running setup. --disable turns the guardrail off. |
| Select mode (observe on first setup, then the current mode) | --mode observe|action. observe records; action enforces. |
| Select hook fail mode (current setting; asked on first setup or when the mode changes) | No flag on setup guardrail. Use defenseclaw setup <connector> --fail-mode open|closed, defenseclaw init --fail-mode, or defenseclaw guardrail fail-mode open|closed afterwards. |
| Human approval for risky actions? (action mode only; current setting) | --human-approval / --no-human-approval. |
| Approval minimum severity (if HITL is on; HIGH) | --hilt-min-severity high|medium|low|critical, matched case-insensitively. |
| Select engine (local) | --scanner-mode local|remote|both. The wizard offers only local and remote; pass both to run the two together. |
| API endpoint / API key env var name / Timeout (remote scanner only) | --cisco-endpoint, --cisco-api-key-env, --cisco-timeout-ms. Defaults come from the existing cisco_ai_defense config. |
| Add LLM judge on top of rule scanning? (action mode only; current setting) | --judge-model turns the judge on. --detection-strategy regex_only turns it off; there is no --no-judge flag on this command. |
| Use this LLM for the judge? (Y, when an LLM is already configured) | --judge-provider, --judge-model, --judge-api-key-env; --inherit-llm or --inherit-from guardrail|scanners.skill|scanners.mcp|scanners.plugin copies another LLM block. To start from the saved judge settings, use defenseclaw setup llm --role judge --inherit-from guardrail.judge. |
| How should DefenseClaw use the LLM? (proxy connectors only) | --llm-role judge_only|judge_and_agent. Hook connectors always get judge_only; see Hook-based vs proxy-based connectors. |
| Configure fallback models? (N) | No flag. Set guardrail.judge.fallbacks (a list of model ids) in ~/.defenseclaw/config.yaml. |
| Configure advanced options? (N) | Opens the next two prompts. |
| Guardrail proxy port (4000) | --port. Used by the proxy connectors; hook connectors bind no proxy. |
| Use a custom block message? (action mode only) | --block-message. |
On a multi-connector install, --connector <name> scopes this setup run to that
connector. When --block-message is passed without --connector, setup treats
it as broad operator intent: it updates the shared block message and reconciles
the active connector overrides, so no connector keeps stale block text.
Two wizard choices have no setup guardrail flag. The hook fail mode is set with defenseclaw setup <connector> --fail-mode or defenseclaw guardrail fail-mode. Judge fallback models live in guardrail.judge.fallbacks in config.yaml.
Redaction has its own workflow
Telemetry redaction is configured per destination with defenseclaw setup redaction, not by this command. See Redaction.
Two modes you have to choose between
observe
Log findings to the audit DB and sinks. Block nothing. Run this for at least a week before promoting.
action
Apply the selected policy thresholds. In the default balanced profile, CRITICAL blocks, HIGH alerts or confirms with HITL, and MEDIUM alerts.
Connector resolution
When you omit --connector, DefenseClaw resolves it in this order:
- The
--connectorflag (operator intent always wins). - The
guardrail.connectorthat an earlier setup run saved. - The
<data_dir>/picked_connectorhint, written byinstall.sh --connector ...and by connector setup. - Otherwise pass
--connector. The interactive wizard suggests an agent it finds on this machine; the non-interactive path does no detection.
Tabs by mode (non-interactive recipes)
defenseclaw setup guardrail \
--non-interactive \
--connector claudecode \
--mode observe \
--scanner-mode local \
--restartMandatory SQLite history fills with every collected prompt and tool call.
Nothing blocks. Open defenseclaw tui for the live audit panel. For a scripted
stream, configure an explicit v8 kind: jsonl destination with the buckets and
redaction profile you intend to expose.
defenseclaw setup guardrail \
--non-interactive \
--connector claudecode \
--mode action \
--scanner-mode local \
--rule-pack default \
--restartWith the default balanced profile, CRITICAL findings block immediately, HIGH and MEDIUM findings alert, and LOW findings allow. Operators see the block or alert context supported by their connector; the audit log captures every verdict.
defenseclaw setup guardrail \
--non-interactive \
--connector claudecode \
--mode action \
--rule-pack default \
--human-approval \
--hilt-min-severity high \
--restartHIGH findings are eligible for confirmation. CRITICAL still blocks unconditionally. On Claude Code, PreToolUse can surface a native ask; on Codex, confirm falls back to an alert/system message with raw_action preserved and does not create a TUI approval. See the HITL page for the full matrix.
defenseclaw keys set DEFENSECLAW_LLM_KEY
defenseclaw setup guardrail \
--non-interactive \
--connector claudecode \
--mode action \
--detection-strategy regex_judge \
--judge-model anthropic/claude-sonnet-4-20250514 \
--judge-api-key-env DEFENSECLAW_LLM_KEY \
--restartkeys set stores the key in ~/.defenseclaw/.env behind a hidden prompt.
Regex still runs first (cheap and offline) and the judge adjudicates what regex
flags. --detection-strategy judge_first flips the order, which helps when
regex is too noisy. Hook connectors are judged only in action mode; use
--judge-hook-connectors to choose which hook connectors the judge covers.
defenseclaw setup guardrail \
--non-interactive \
--connector claudecode \
--mode action \
--detection-strategy regex_judge \
--judge-provider bedrock \
--judge-model us.anthropic.claude-sonnet-4-6 \
--judge-bedrock-region us-east-1 \
--judge-bedrock-auth-mode iam_credentials \
--judge-bedrock-access-key-env AWS_ACCESS_KEY_ID \
--judge-bedrock-secret-key-env AWS_SECRET_ACCESS_KEY \
--judge-bedrock-inference-profile us. \
--restartRoutes the judge through AWS Bedrock instead of a SaaS endpoint. boto3 ships in the base install, so no extra pip install is needed. Swap --judge-provider vertex_ai and the --judge-vertex-* flags for GCP Vertex AI, or --judge-provider azure and the --judge-azure-* flags for Azure OpenAI (see Every flag). For self-signed lab endpoints, add --judge-tls-ca-cert-file /etc/ssl/lab-root.pem. See Unified LLM key → Regional providers for the full matrix.
If the same Bedrock, Vertex AI or Azure settings are already on a custom-provider overlay entry (region, auth mode, deployment aliases, TLS; see Bedrock, Vertex AI, Azure on a custom instance), the judge can inherit them with a single --judge-instance-name <name> instead of repeating every flag. Role-level judge flags still win field by field, so a shared overlay can supply the credentials while --judge-bedrock-region pins a different region per environment.
Every flag
The table follows defenseclaw setup guardrail --help.
Prop
Type
What setup writes
setup guardrail saves config.yaml. The gateway writes the connector files
when it starts: for Claude Code, the hook scripts under ~/.defenseclaw/hooks/
and the hook entries in ~/.claude/settings.json that call
claude-code-hook.sh. A hash-checked backup is stored before edits; teardown
restores the file or removes only the DefenseClaw-owned entries. See the
connector pages for the exact files each agent uses.
Verify it worked
defenseclaw doctor
defenseclaw status
defenseclaw guardrail status
defenseclaw alerts --limit 25doctor prints the full health report. status shows the environment, the
enforcement counts and the active connector roster with each connector's mode;
add --json for automation. guardrail status shows the resolved posture of
each connector. alerts lists recent decisions as a table; --connector <name>
filters by connector, --show <n> prints one alert in full and --json prints
the list as JSON. For a live view, open defenseclaw tui.
Which commands prompt and which are flag-only is listed in the interactive vs non-interactive matrix.
Common follow-ups
Quick aliases
defenseclaw setup claude-code, setup codex, setup cursor, ...
Multi-connector
One gateway enforcing several hook connectors through guardrail.connectors. Choose Add at the prompt.
Changing connectors
Add, reconfigure, replace or remove connector wiring with the setup aliases.
Disabling
--disable rolls everything back, including agent-side hook entries.
HITL
Per-connector native ask events and non-pausing fallbacks.