Set up the guardrail

The central command. Configures how DefenseClaw guards your agent (connector, observe or action mode, scanner, judge, human approval and rule pack), then restarts the gateway.

defenseclaw setup guardrail is the central operator command. It picks the connector, the mode, the scanner backend, the optional LLM judge and the human-in-the-loop (HITL) posture. Connector setup aliases such as defenseclaw setup claude-code configure one connector with a smaller flag set; other setup commands manage webhooks, observability, scanners and providers on their own.

Run it interactively the first time: the wizard explains each choice and picks safe defaults. Once you are happy with the configuration, re-run it with explicit flags and --non-interactive for unattended setups and CI.

Watch the interactive flow

The animation replays a real defenseclaw setup guardrail session on a host where no connector was configured yet: connector pick, observe or action, hook fail mode, scanner engine, the judge step and the summary. In observe mode the wizard skips the human-approval and judge questions, because neither applies until a connector enforces. Hover or press Pause to study a frame; press Restart to replay.

~/code/your-agent-repo
Recorded with --no-restart, so it ends by asking you to restart the gateway; without that flag, setup restarts it for you. Defaults are in [brackets]; Enter accepts them. The orange characters are the operator's replies.

When a connector is already configured, the wizard skips the connector picker and says that it is editing the global guardrail policy for the configured connectors. Adding or switching a connector is the job of defenseclaw setup <connector>; see Multi-connector.

Same setup, no prompts

Every choice the wizard makes has a flag, apart from the two noted below. The CI-friendly equivalent of the session above is one command:

defenseclaw setup guardrail \
  --non-interactive \
  --connector claudecode \
  --mode observe \
  --scanner-mode local \
  --detection-strategy regex_only \
  --restart

--non-interactive (aliases --accept-defaults and --yes) skips every prompt; missing flags fall back to the defaults the wizard would have offered. The judge stays off until you pass --judge-model. The saved detection strategy defaults to regex_judge, but with the judge off only the regex lane runs; --detection-strategy regex_only makes that explicit.

Don't hand-roll the flags

The Command generator builds a non-interactive defenseclaw setup guardrail invocation for any connector with all the options below (mode, scanner, judge, HITL and advanced settings) and shows validation warnings inline.

Prompt-flag mapping

Each row is one wizard prompt, with the default that Enter selects and the flag to use in scripts.

Prompt (default)Flag, or how to script it
Which agent framework are you using? (first setup only; suggests an agent it finds)--connector / --agent: claudecode, codex, cursor, devin, copilot, hermes, openhands, antigravity, opencode, amp, omnigent, kiro, openclaw, zeptoclaw.
Enable guardrail? (Y)Implied by running setup. --disable turns the guardrail off.
Select mode (observe on first setup, then the current mode)--mode observe|action. observe records; action enforces.
Select hook fail mode (current setting; asked on first setup or when the mode changes)No flag on setup guardrail. Use defenseclaw setup <connector> --fail-mode open|closed, defenseclaw init --fail-mode, or defenseclaw guardrail fail-mode open|closed afterwards.
Human approval for risky actions? (action mode only; current setting)--human-approval / --no-human-approval.
Approval minimum severity (if HITL is on; HIGH)--hilt-min-severity high|medium|low|critical, matched case-insensitively.
Select engine (local)--scanner-mode local|remote|both. The wizard offers only local and remote; pass both to run the two together.
API endpoint / API key env var name / Timeout (remote scanner only)--cisco-endpoint, --cisco-api-key-env, --cisco-timeout-ms. Defaults come from the existing cisco_ai_defense config.
Add LLM judge on top of rule scanning? (action mode only; current setting)--judge-model turns the judge on. --detection-strategy regex_only turns it off; there is no --no-judge flag on this command.
Use this LLM for the judge? (Y, when an LLM is already configured)--judge-provider, --judge-model, --judge-api-key-env; --inherit-llm or --inherit-from guardrail|scanners.skill|scanners.mcp|scanners.plugin copies another LLM block. To start from the saved judge settings, use defenseclaw setup llm --role judge --inherit-from guardrail.judge.
How should DefenseClaw use the LLM? (proxy connectors only)--llm-role judge_only|judge_and_agent. Hook connectors always get judge_only; see Hook-based vs proxy-based connectors.
Configure fallback models? (N)No flag. Set guardrail.judge.fallbacks (a list of model ids) in ~/.defenseclaw/config.yaml.
Configure advanced options? (N)Opens the next two prompts.
Guardrail proxy port (4000)--port. Used by the proxy connectors; hook connectors bind no proxy.
Use a custom block message? (action mode only)--block-message.

On a multi-connector install, --connector <name> scopes this setup run to that connector. When --block-message is passed without --connector, setup treats it as broad operator intent: it updates the shared block message and reconciles the active connector overrides, so no connector keeps stale block text.

Two wizard choices have no setup guardrail flag. The hook fail mode is set with defenseclaw setup <connector> --fail-mode or defenseclaw guardrail fail-mode. Judge fallback models live in guardrail.judge.fallbacks in config.yaml.

Redaction has its own workflow

Telemetry redaction is configured per destination with defenseclaw setup redaction, not by this command. See Redaction.

Two modes you have to choose between

observe

Log findings to the audit DB and sinks. Block nothing. Run this for at least a week before promoting.

action

Apply the selected policy thresholds. In the default balanced profile, CRITICAL blocks, HIGH alerts or confirms with HITL, and MEDIUM alerts.

Connector resolution

When you omit --connector, DefenseClaw resolves it in this order:

  1. The --connector flag (operator intent always wins).
  2. The guardrail.connector that an earlier setup run saved.
  3. The <data_dir>/picked_connector hint, written by install.sh --connector ... and by connector setup.
  4. Otherwise pass --connector. The interactive wizard suggests an agent it finds on this machine; the non-interactive path does no detection.

Tabs by mode (non-interactive recipes)

defenseclaw setup guardrail \
  --non-interactive \
  --connector claudecode \
  --mode observe \
  --scanner-mode local \
  --restart

Mandatory SQLite history fills with every collected prompt and tool call. Nothing blocks. Open defenseclaw tui for the live audit panel. For a scripted stream, configure an explicit v8 kind: jsonl destination with the buckets and redaction profile you intend to expose.

defenseclaw setup guardrail \
  --non-interactive \
  --connector claudecode \
  --mode action \
  --scanner-mode local \
  --rule-pack default \
  --restart

With the default balanced profile, CRITICAL findings block immediately, HIGH and MEDIUM findings alert, and LOW findings allow. Operators see the block or alert context supported by their connector; the audit log captures every verdict.

defenseclaw setup guardrail \
  --non-interactive \
  --connector claudecode \
  --mode action \
  --rule-pack default \
  --human-approval \
  --hilt-min-severity high \
  --restart

HIGH findings are eligible for confirmation. CRITICAL still blocks unconditionally. On Claude Code, PreToolUse can surface a native ask; on Codex, confirm falls back to an alert/system message with raw_action preserved and does not create a TUI approval. See the HITL page for the full matrix.

defenseclaw keys set DEFENSECLAW_LLM_KEY

defenseclaw setup guardrail \
  --non-interactive \
  --connector claudecode \
  --mode action \
  --detection-strategy regex_judge \
  --judge-model anthropic/claude-sonnet-4-20250514 \
  --judge-api-key-env DEFENSECLAW_LLM_KEY \
  --restart

keys set stores the key in ~/.defenseclaw/.env behind a hidden prompt. Regex still runs first (cheap and offline) and the judge adjudicates what regex flags. --detection-strategy judge_first flips the order, which helps when regex is too noisy. Hook connectors are judged only in action mode; use --judge-hook-connectors to choose which hook connectors the judge covers.

defenseclaw setup guardrail \
  --non-interactive \
  --connector claudecode \
  --mode action \
  --detection-strategy regex_judge \
  --judge-provider bedrock \
  --judge-model us.anthropic.claude-sonnet-4-6 \
  --judge-bedrock-region us-east-1 \
  --judge-bedrock-auth-mode iam_credentials \
  --judge-bedrock-access-key-env AWS_ACCESS_KEY_ID \
  --judge-bedrock-secret-key-env AWS_SECRET_ACCESS_KEY \
  --judge-bedrock-inference-profile us. \
  --restart

Routes the judge through AWS Bedrock instead of a SaaS endpoint. boto3 ships in the base install, so no extra pip install is needed. Swap --judge-provider vertex_ai and the --judge-vertex-* flags for GCP Vertex AI, or --judge-provider azure and the --judge-azure-* flags for Azure OpenAI (see Every flag). For self-signed lab endpoints, add --judge-tls-ca-cert-file /etc/ssl/lab-root.pem. See Unified LLM key → Regional providers for the full matrix.

If the same Bedrock, Vertex AI or Azure settings are already on a custom-provider overlay entry (region, auth mode, deployment aliases, TLS; see Bedrock, Vertex AI, Azure on a custom instance), the judge can inherit them with a single --judge-instance-name <name> instead of repeating every flag. Role-level judge flags still win field by field, so a shared overlay can supply the credentials while --judge-bedrock-region pins a different region per environment.

Every flag

The table follows defenseclaw setup guardrail --help.

Prop

Type

What setup writes

config.yaml
picked_connector
settings.json (DefenseClaw hook entries and OTEL_* env)

setup guardrail saves config.yaml. The gateway writes the connector files when it starts: for Claude Code, the hook scripts under ~/.defenseclaw/hooks/ and the hook entries in ~/.claude/settings.json that call claude-code-hook.sh. A hash-checked backup is stored before edits; teardown restores the file or removes only the DefenseClaw-owned entries. See the connector pages for the exact files each agent uses.

Verify it worked

defenseclaw doctor
defenseclaw status
defenseclaw guardrail status
defenseclaw alerts --limit 25

doctor prints the full health report. status shows the environment, the enforcement counts and the active connector roster with each connector's mode; add --json for automation. guardrail status shows the resolved posture of each connector. alerts lists recent decisions as a table; --connector <name> filters by connector, --show <n> prints one alert in full and --json prints the list as JSON. For a live view, open defenseclaw tui.

Which commands prompt and which are flag-only is listed in the interactive vs non-interactive matrix.

Common follow-ups