Defaults
What every fresh DefenseClaw install ships with — three OPA policies (permissive / default / strict), three matching guardrail rule packs, the operator-config defaults, and how to pick the combination that fits your team's risk tolerance.
DefenseClaw ships with opinionated defaults that are immediately useful without tuning, and stay out of your way until you opt into stricter behaviour. This page documents what actually ships — grounded in the policy YAMLs in policies/, the schema in internal/config/config.go, and what defenseclaw setup guardrail actually writes.
The two layers you can swap
DefenseClaw separates admission policy from runtime guardrail rules. They ship as matching triples but are independent knobs.
Admission policy (OPA)
What happens when a skill / MCP / plugin gets installed or executed. Lives in policies/<name>.yaml, activated by defenseclaw policy activate <name>.
Guardrail rule pack
The rules, LLM-judge prompts, and suppressions that check prompts, completions, and tool calls in flight. Lives in policies/guardrail/<name>/, pointed at by guardrail.rule_pack_dir.
Both layers ship in three named profiles — default, strict, permissive — and you flip them independently:
defenseclaw policy activate default # OPA layer
defenseclaw guardrail use-pack default # Guardrail layerdefenseclaw policy activate strict does not change
guardrail.rule_pack_dir, and vice versa. If you want strict everywhere,
flip both. Policies explains why admission policy and
in-flight guardrail detection are independent layers.
OPA admission policy — what each profile ships
The OPA policy file (policies/<name>.yaml) drives admission decisions: what happens to a finding by severity, whether the allow-list lets first-party assets bypass scanning, and what threshold a Cisco AI Defense verdict has to clear before it blocks.
| Knob | permissive | default | strict |
|---|---|---|---|
admission.allow_list_bypass_scan | true | true | false |
skill_actions.critical | quarantine + disable + block | quarantine + disable + block | quarantine + disable + block |
skill_actions.high | none + enable + none | quarantine + disable + block | quarantine + disable + block |
skill_actions.medium | none + enable + none | none + enable + none | quarantine + disable + block |
scanner_overrides | empty | MCP LOW/MEDIUM and plugin MEDIUM/HIGH overrides | MCP LOW/MEDIUM and plugin MEDIUM/HIGH overrides |
guardrail.block_threshold | 4 (CRITICAL) | 4 (CRITICAL) | 2 (MEDIUM) |
guardrail.alert_threshold | 3 (HIGH) | 2 (MEDIUM) | 1 (LOW) |
guardrail.cisco_trust_level | advisory | full | full |
guardrail.hilt.enabled | false (key omitted) | false | false (key omitted) |
guardrail.hilt.min_severity | HIGH | HIGH | HIGH |
Severity ranks are the rego convention from policies/rego/guardrail.rego: 1 = LOW, 2 = MEDIUM, 3 = HIGH, 4 = CRITICAL. cisco_trust_level: advisory means even Cisco AI Defense's own verdicts are surfaced but never escalated to a block.
The columns are deliberately conservative. We'd rather you opt into stricter behaviour than have an upgrade silently start blocking your traffic.
Tool-call block and alert levels
The guardrail.block_threshold and alert_threshold rows above apply to LLM
traffic through the guardrail proxy. Tool calls that agent hooks send, from
Claude Code, Codex and the other hook connectors, take their levels from the
first of these that is set:
- The connector's own level:
defenseclaw guardrail block-at LEVEL --connector <name>writesguardrail.connectors.<name>.block_at. - The global level:
defenseclaw guardrail block-at LEVELwritesguardrail.block_at. - The rule pack:
strictblocks MEDIUM and above,defaultandpermissiveblock CRITICAL only.
defenseclaw guardrail alert-at sets the alert level the same way. Without
one, strict alerts on LOW and above, default on MEDIUM and above and
permissive on HIGH and above. The levels apply in action mode, and
defenseclaw guardrail status shows the result for each connector in its
block/alert field.
Guardrail rule pack — what each profile ships
The rule pack directory (policies/guardrail/<name>/) holds the regex YAMLs, judge prompts, sensitive-tool definitions, and suppressions the in-flight scanner consumes.
| Pack | rules/ files | judge/ prompts | suppressions.yaml | sensitive-tools.yaml |
|---|---|---|---|---|
permissive | c2, cognitive, commands, enterprise-data, local-patterns, secrets, sensitive-paths, trust-exploit | pii and tool-injection (higher judge thresholds); injection ships disabled | broadest; additionally suppresses all IP findings and selected file-inspection PII | same six sensitive tool definitions as the other packs |
default | same eight families; balanced variants where profiles differ | injection, pii, tool-injection | private/loopback IPs, platform IDs, expected system metadata | same six sensitive tool definitions as the other packs |
strict | same eight families; stricter variants where profiles differ | injection, pii, tool-injection (lower judge thresholds) | minimal structural suppressions; no tool suppressions | same six sensitive tool definitions as the other packs |
All three packs share the same severity rubric and the same signal_strength output schema — only the per-category thresholds and suppression scope differ between packs. Switching the rule pack does not enable the LLM judge — that's a separate guardrail.judge.enabled toggle in your operator config (default: false). Flipping the rule pack only changes which prompt YAMLs the judge will run if you've enabled it.
What setup guardrail actually writes
After defenseclaw init, this command:
defenseclaw setup guardrail --connector claudecode --mode action \
--rule-pack default --non-interactivewrites these settings (other generated fields omitted):
claw:
mode: claudecode
guardrail:
connector: claudecode
enabled: true
scanner_mode: local
llm_role: judge_only
mode: action
rule_pack_dir: /home/<you>/.defenseclaw/policies/guardrail/default
observability: {} # no optional destinations; local SQLite still recordsSettings it doesn't write keep their defaults: guardrail.hook_fail_mode is
closed on a new install, guardrail.judge.enabled is false, and
guardrail.hilt.enabled is false with min_severity: HIGH.
Things to notice:
- The LLM judge is off by default. It only turns on if you pass
--judge-modeltosetup guardrailor answer "yes" to the interactive judge prompt. The schema default isguardrail.judge.enabled = falseininternal/config/config.go. Keeping it off keeps cost predictable; turn it on once you have aDEFENSECLAW_LLM_KEYconfigured. - Human approval (HITL) is off by default. The shipped severity floor is
HIGH, but withenabled: falseit never asks.--human-approvalturns it on and--hilt-min-severitychanges the floor. The config key and flags are spelledhilt. - The local store still writes.
observability: {}means no optional destination, not no telemetry. Every collected log is stored, unredacted, in the local SQLite database. Add akind: jsonldestination if you need a file stream, or a remote destination withsetup splunkorsetup local-observability. - New destinations get unredacted data. A destination without its own redaction policy receives every capability and bucket under profile
none. Configuresensitive,content,strictor a custom profile before you export to a less trusted place.
Tuning by risk tolerance
You usually don't need a custom policy or rule pack — just a few knob changes.
"I'm in pilot, just observe"
defenseclaw policy activate permissive
defenseclaw setup guardrail \
--connector claudecode \
--mode observe \
--rule-pack permissive \
--restart \
--non-interactiveObserve mode never blocks; the permissive rule pack only decides what gets logged. Everything still flows to mandatory SQLite and any matching configured destinations so you can review what would have happened. Recommended first week of any rollout.
"Move fast, stop only the obvious harm"
defenseclaw policy activate default
defenseclaw setup guardrail \
--connector claudecode \
--mode action \
--rule-pack default \
--human-approval \
--hilt-min-severity high \
--restart \
--non-interactiveDefault rules: CRITICAL findings block, and HIGH findings pause for your approval inside Claude Code, which supports a native PreToolUse ask. Most engineering teams in the early or middle phase of a rollout land here.
"Regulated workload, lock it down"
defenseclaw policy activate strict
defenseclaw setup guardrail \
--connector claudecode \
--mode action \
--rule-pack strict \
--detection-strategy regex_judge \
--judge-model anthropic/claude-haiku-4-5 \
--restart \
--non-interactiveStrict policy (block ≥ MEDIUM, no allow-list bypass), strict rule pack (stricter profile variants and minimal suppressions), and the LLM judge enabled. Because the strict block threshold runs before HITL, MEDIUM-and-higher findings block rather than prompt. Combine this with a reviewed first-party allow-list.
"Block risky plugins at install"
How DefenseClaw treats skills, MCP servers and plugins with findings is set in
the admission policy, independent of the rule pack. The shipped default
policy installs a plugin with a MEDIUM finding and lets it run. To block and
quarantine those plugins instead:
defenseclaw policy edit scanner --type plugin --severity medium \
--install block --file quarantine --runtime disableThe first edit of a built-in policy saves your own copy in
~/.defenseclaw/policies/default.yaml, and the command asks the running
gateway to reload it. defenseclaw policy show default lists the result under
Scanner Overrides, and defenseclaw policy delete default restores the
built-in policy.
Defaults defenseclaw init does not ask about
init does not prompt for the following lower-level defaults. Use the
documented setup/guardrail command when it exposes the knob; otherwise edit
~/.defenseclaw/config.yaml and validate the result.
| Knob | Default | Why fixed |
|---|---|---|
observability.local SQLite history | always present | Mandatory local durability for collected logs and the compliance floor |
guardrail.hook_fail_mode | closed on new installs | Delivery, authentication, and invalid-response failures fail closed where the connector can block; upgrades from the pre-change default are migrated to open for compatibility |
guardrail.judge.timeout | 30s | Hot-path latency budget for the judge |
guardrail.judge.adjudication_timeout | 5s | Per-prompt adjudication budget |
guardrail.detection_strategy | regex_judge | Regex first, then the judge for MEDIUM and higher findings. Until you turn the judge on it runs as regex_only, which is what defenseclaw guardrail status shows on a new install |
| Bifrost retry policy | 3 attempts, exp backoff | Tested LLM-routing baseline |
defenseclaw setup llm --max-retries, defenseclaw setup guardrail --detection-strategy, and defenseclaw guardrail fail-mode provide supported
CLI paths for three of these controls. Run defenseclaw config validate after
any direct edit.
Per-connector overrides (guardrail.connectors)
When you run more than one hook connector from a single gateway, override guardrail policy per connector under guardrail.connectors.<name> in ~/.defenseclaw/config.yaml. Every field is optional and inherits the global guardrail.* value when unset, so a connector block only carries what differs:
guardrail:
mode: action # global default
hook_fail_mode: closed
connectors:
claudecode:
mode: action # enforce for Claude Code
codex:
mode: observe # softer for Codex than for Claude Code
hook_fail_mode: open # explicit softer override for CodexWith more than one connector active, claw.mode keeps naming the primary connector, and the gateway leaves the defenseclaw.claw.mode telemetry attribute off. Manage these blocks with defenseclaw setup <connector> (choosing Add) and the defenseclaw guardrail ... --connector X command group, for example defenseclaw guardrail block-at HIGH --connector codex — see Setup → Multi-connector and Reference → Configuration. The OPA admission policy is still global — there's no per-connector policy override surface yet.
Legacy top-level connector blocks are deprecated
Older installs used top-level claude_code: / codex: blocks (the AgentHookConfig fields) for per-connector overrides:
claude_code:
enabled: true
mode: action
fail_mode: open # LEGACY hint, NOT consumed by hooks; see Reference → Fail modesThese are still parsed for backward compatibility, but fail_mode here does nothing at runtime (see Reference → Fail modes). Prefer guardrail.connectors.<name> for new configuration — it's the surface the per-connector CLI writes and the gateway resolves at request time.
Inspect the active defaults
defenseclaw config show # every setting, with defaults (secrets masked)
defenseclaw config get asset_policy.enabled # one value of that view
defenseclaw policy list # all policies on disk + which is active
defenseclaw policy show default # normalized summary of one named policyFor configuration v8, config show prints every section: the values
config.yaml sets plus the defaults that apply to the keys it leaves out, with
the observability section resolved the way the gateway runs it. A fresh
install's config.yaml sets only a few sections, so this is how you read
asset_policy and the rest before you change them.
--section NAME keeps one section (for example asset_policy), and
config get KEY prints one dotted key such as
asset_policy.mcp.registry_required. A default value prints with a note on
stderr; the exit code is 1 when the key has no value and no default, and 2
for an unknown section. Use --source for the masked source document as written,
--effective for the resolved observability section alone, or
--effective --provenance for explicit compiler annotations. V8 rejects
--reveal; neither source nor effective output resolves secret values.
policy show <name> prints a normalized summary of the named file (default,
strict, permissive, or any custom policies/<name>.yaml you've added).
It does not dump the source YAML or individual guardrail rules. To find a rule
by ID, look up the directory a connector enforces with list-packs, then
search it:
defenseclaw guardrail list-packs
grep -rn "RULE_ID" ~/.defenseclaw/policies/guardrail/default/rules/(built-in default) in the list-packs output means no pack has been
selected yet, so the default pack applies.
Reset to defaults
There's no --reset flag. Two real paths exist:
Soft reset (most common) — just re-run setup with the defaults you want. setup guardrail overwrites the relevant guardrail.* keys idempotently:
defenseclaw setup guardrail --rule-pack default --no-human-approval
defenseclaw policy activate defaultReset user state — first make your own private backup if you need one,
then use the destructive reset command. It deletes resettable state; it does
not create a rollback archive. The managed runtime and installed plugin are
retained so quickstart can reinstall. The example is for macOS and Linux. On
Windows, back up %USERPROFILE%\.defenseclaw with Copy-Item -Recurse, then
run the same defenseclaw reset and defenseclaw quickstart. A managed
(enterprise) install is reset through its
lifecycle commands instead.
umask 077
cp -a "$HOME/.defenseclaw" "$HOME/.defenseclaw.before-reset"
defenseclaw reset
defenseclaw quickstartSee also
- Policies — the layered architecture (regex → judge → suppressions → OPA admission)
- Setup Guardrail — the CLI that consumes these defaults
- HITL — what
guardrail.hilt.enabledandmin_severityactually change for the operator - Reference → Fail modes — the three "fail open vs closed" knobs disambiguated
- Reference → Configuration — every key surfaced here, with type and default
Policies
How DefenseClaw decides — OPA/Rego policy, structured tool-call rules, regex and judge detection, scanner policy, and noise suppression.
Deterministic detection reference
Complete repository-backed reference for DefenseClaw ActionFacts, CEL, regex, semantic proofs, YARA, bounded chains, profiles, and protection packs.