Thresholds: block_at and alert_at
The single threshold model for DefenseClaw guardrails. Where block_at and alert_at can be set, which one wins, what the rule pack supplies when you set nothing, and how to read, change and hot-apply them.
Two settings decide how severe a finding must be before the guardrail acts:
block_atis the lowest severity that is blocked.alert_atis the lowest severity that raises an alert.
Both take CRITICAL, HIGH, MEDIUM or LOW. They are the only threshold source: one pair of settings, resolved the same way on every guardrail surface. There is no separate tool-call level, no per-direction level and no threshold hidden in a policy file. config.yaml holds the values, and the rule pack supplies a default for anything you leave unset.
What the two levels do
With block_at: HIGH and alert_at: MEDIUM:
| Finding severity | Action mode | Observe mode |
|---|---|---|
CRITICAL, HIGH | Blocked | Alert |
MEDIUM | Alert | Alert |
LOW | Allowed | Allowed |
- Anything that blocks also alerts. An
alert_atthat sits aboveblock_aton the severity scale (for exampleblock_at: LOWwithalert_at: MEDIUM) alerts from the block level instead.defenseclaw config get guardrail.alert_at --effectivelabels that case(clamped to block_at). - Observe mode never blocks. Findings at or above the alert level are logged as alerts and nothing is stopped. The levels take effect when the connector runs in action mode.
- Human approval sits below the block level. When HITL is on, a finding at or above
guardrail.hilt.min_severitybut belowblock_atbecomes aconfirmon a connector event that can ask. Findings at or aboveblock_atblock without asking.
Where the levels apply
The resolved levels decide every guardrail verdict that carries a severity:
| Surface | Covered |
|---|---|
| Hook tool calls | Yes, for every hook connector (bounded chains, CodeGuard file writes and sandbox shell commands included) |
| Hook prompts | Yes |
| Guardrail proxy prompts | Yes |
| Guardrail proxy completions | Yes |
| Guardrail proxy tool calls in model responses | Yes |
| Tool responses and message content | Yes |
These levels apply only after a finding is eligible to block. A tool-call
pattern match is detection-only, and a rule whose trusted proof is explicitly
alert-only stays an alert at every block_at value. For example,
impact.sql_destructive_mutation reports destructive SQL as an alert in the
standard packs because the operation can be authorized administration.
The inline hook notice identifies that rule as alert-only. Other alert-only
tool findings include attack.http_command_injection,
attack.http_sql_injection, and exec.postgresql_copy_program. See
CEL rule authoring for the proof boundary.
A preset (defenseclaw policy activate), defenseclaw guardrail block-at and a hand edit of config.yaml all change the same two keys.
Upgrading from config_version 8
Before config_version: 9, block_at and alert_at applied only to hook tool calls, and prompts, completions and proxy traffic took their levels from the block_threshold and alert_threshold in the policy data file (data.json). The v9 migration reads those two thresholds once, at the upgrade from 8, and writes them as guardrail.block_at and guardrail.alert_at. It leaves a level unset when it equals the rule pack's default, and it prints a note that the keys now apply everywhere. Nothing reads data.json after that. See the v8 to v9 migration guide.
Scopes and precedence
A level can be set at four scopes. The most specific scope that sets a value wins, and an unset scope falls through to the next one.
List view for small screens. Use the expand button to open the drawing.
- 1.Profile connectorprofiles.P.connectors.C
- if unsetProfile
- 2.Profileprofiles.P.block_at
- if unsetConnector
- 3.Connectorconnectors.C.block_at
- if unsetGlobal
- 4.Globalguardrail.block_at
- if unsetRule pack default
- 5.Rule pack defaultpack manifest or name
From most to least specific:
- Profile connector:
guardrail.profiles.<profile>.connectors.<connector>.block_at - Profile:
guardrail.profiles.<profile>.block_at - Connector:
guardrail.connectors.<connector>.block_at - Global:
guardrail.block_at - Rule-pack default: the
postureof the pack that scope uses (below)
Application protection overlays (application_protection.connectors.<connector>.guardrail) can override some other guardrail fields, such as the mode, HITL and the rule pack, for connectors you have not set up by hand. They cannot carry block_at or alert_at: a config that sets either under application_protection is refused. For the thresholds that tier is therefore always empty, and the order reads profile connector > profile > guardrail.connectors > application protection (empty) > global > rule-pack default.
Note that a profile's own level outranks guardrail.connectors.<connector>: a subject a profile selects gets the profile's block_at even on a connector that has its own, unless the profile sets that connector too. Profiles are how you give different people different levels; see User- and group-based policies for how assignments pick a profile.
All four scopes in one file:
guardrail:
block_at: HIGH # global: HIGH and CRITICAL block everywhere
alert_at: MEDIUM
connectors:
codex:
block_at: CRITICAL # connector: Codex blocks only CRITICAL
profiles:
contractors:
block_at: MEDIUM # profile: contractors block MEDIUM and above
alert_at: LOW
connectors:
codex:
block_at: HIGH # profile connector: except on CodexRule-pack default (posture)
A scope with no level of its own takes the default of the rule pack it enforces. The default comes from the pack's posture:
| Posture | block_at | alert_at |
|---|---|---|
default | CRITICAL | MEDIUM |
strict | MEDIUM | LOW |
permissive | CRITICAL | HIGH |
The posture of a pack is, in order:
- the
posturefield ofdefenseclaw-pack.jsonat the pack's root (default,strictorpermissive), if the file exists; - the pack's own name, for the built-in packs;
- otherwise the folder name when it is
strictorpermissive, anddefaultfor every other folder.
The three built-in packs need no manifest. A custom pack is default unless it says otherwise, even when you copied it from strict. To keep strict levels for a copy, add a manifest to the pack directory:
{"posture": "strict"}The manifest is part of the pack digest: adding or editing it changes the custom_packs pin and the effective policy digest, like a rule file, so a pinned pack whose manifest changed is refused until you pin the new digest. defenseclaw guardrail validate-pack accepts a pack with or without it. See Rules and custom packs for pinning a pack.
Because a profile or connector can select a different rule_pack, the default can differ per scope. Setting block_at explicitly at a scope removes that dependency.
Cisco AI Defense trust level
guardrail.cisco_trust_level decides how much a Cisco AI Defense verdict counts when you run the remote scanner (scanner_mode: remote or both):
| Value | Effect |
|---|---|
full (default) | The verdict counts like a local finding and can reach block_at. |
advisory | A Cisco verdict that alone would reach the block level raises an alert instead, unless a local finding also reaches the alert level. |
none | Cisco AI Defense severities are ignored when computing the verdict. |
It is a single global setting; connectors and profiles do not override it. Change it with defenseclaw policy edit guardrail --cisco-trust-level advisory or defenseclaw config set guardrail.cisco_trust_level advisory. The permissive preset sets advisory.
Read the effective value
config get --effective prints the value the gateway enforces and where it comes from. The source line goes to stderr, so a script that captures the output gets only the value.
$ defenseclaw config get guardrail.block_at --effective # block_at: HIGH
(source: config:guardrail.block_at)
HIGH
$ defenseclaw config get guardrail.connectors.codex.block_at --effective # codex sets its own
(source: config:guardrail.connectors.codex.block_at)
HIGH
$ defenseclaw config get guardrail.alert_at --effective # alert_at not set anywhere
(source: pack-default:default)
MEDIUMThe source is config:guardrail.<key> for the global scope, config:guardrail.connectors.<connector>.<key> for a connector, or pack-default:<posture> when nothing is set and the rule pack decides. --effective reads the global and connector scopes. For a profile scope, ask which profile a request resolves to and what it applies:
defenseclaw guardrail profile explain --connector codexGuardrail profile resolution
profile: contractors
match: connector
applies (codex): mode=action block_at=MEDIUM alert_at=pack hilt=off rule_pack=...alert_at=pack means the profile sets no alert level, so the rule pack's default applies. Add --user alice to resolve a specific user. Two more views:
defenseclaw guardrail status
defenseclaw-gateway policy showguardrail status prints the result per connector in its Block/alert column (MEDIUM+/LOW+: blocks MEDIUM and above, alerts LOW and above). defenseclaw-gateway policy show prints what the running binary compiles from config.yaml, under guardrail. It is the one to use on a managed host, where the defenseclaw Python CLI is not installed.
Set the levels
| To | Run |
|---|---|
| Set the global block level | defenseclaw guardrail block-at HIGH |
| Set the global alert level | defenseclaw guardrail alert-at MEDIUM |
| Set one connector's levels | defenseclaw guardrail block-at CRITICAL --connector codex |
| Clear a level so it inherits | defenseclaw guardrail block-at inherit --connector codex |
| Set a profile's level | defenseclaw config set guardrail.profiles.contractors.block_at MEDIUM |
| Set a profile connector's level | defenseclaw config set guardrail.profiles.contractors.connectors.codex.block_at HIGH |
| Remove any key | defenseclaw config unset guardrail.profiles.contractors.block_at |
| Take a preset's levels | defenseclaw policy activate strict |
| Set levels and trust level together | defenseclaw policy edit guardrail --block-threshold HIGH --alert-threshold MEDIUM |
block-at and alert-at take CRITICAL, HIGH, MEDIUM, LOW (any case) or inherit, and have no --profile option; use config set for profile scopes. Every change goes through the one config writer, which validates the result before writing it, so a bad level is refused and config.yaml stays as it was.
policy activate writes guardrail.block_at and guardrail.alert_at from the preset, but only when the preset's level differs from the default of the globally selected rule pack. Activating default on the default pack writes neither key. strict writes block_at: MEDIUM and alert_at: LOW. permissive writes alert_at: HIGH and cisco_trust_level: advisory. See Defaults for the full preset table.
Managed devices refuse local changes
On a device managed by MDM or a management plane, the config writer refuses. defenseclaw config set exits with status 3 and the message This device is managed: change config.yaml in the admin config (MDM or management plane), not on the device, and the commands built on the writer refuse too. Set the levels in the admin config instead.
Hot reload
A change to a threshold, a rule pack, guardrail.rules, a profile, the mode, HITL or the trust level is applied by the running gateway without a restart. The gateway builds a new policy generation in the background and swaps it in atomically. defenseclaw guardrail block-at HIGH ends with The running gateway applies it now., and the next decision uses the new level.
If a change is invalid, the previous generation keeps running. The rejected change shows up as policy.last_reload_error in the gateway /health output and as a failing Policy row in defenseclaw doctor, which also lists the generation and the digest still being enforced.
A restart is still needed for the process-level guardrail keys: guardrail.host, guardrail.port, guardrail.enabled, guardrail.connector, guardrail.scanner_mode, guardrail.retain_judge_bodies and the hook self-heal settings. A connector's own enabled applies in place: the gateway sets up or removes its hooks. guardrail.hook_fail_mode reloads hot: the gateway rewrites the installed hook scripts (with hook self-heal off it runs the connector setup again), and defenseclaw guardrail fail-mode rewrites them itself. The writer prints Restart the gateway to apply <key>: defenseclaw-gateway restart when you change one. The rest of such a change still applies at once; the gateway logs which keys wait for the restart. With gateway.config_reload.mode: restart the gateway restarts itself when a change needs it, instead of waiting for you.
Secure Client
When Cisco Secure Client integration is active, the gateway keeps the rule pack's posture levels for content decisions. guardrail.block_at and alert_at do not change them.
See also
- Rules and custom packs:
guardrail.rule_pack,custom_packsandguardrail.rules - Defaults: what each shipped preset and pack sets
- HITL: the confirm band below the block level
- User- and group-based policies: assign profiles by user, group, connector or agent
- Reference: Configuration: every key
Defaults
What every fresh DefenseClaw install ships with — three policy presets (permissive / default / strict), three matching guardrail rule packs, the operator-config defaults, and how to pick the combination that fits your team's risk tolerance.
Rules: guardrail.rules and custom packs
Choose a guardrail rule pack, add your own pack pinned by digest, and adjust individual rules, suppressions and sensitive tools in config.yaml without copying a pack. Covers scopes, layering order and what happens when a pack or digest is wrong.