Skill scanner
Scan connector skills on demand and whenever one is installed or changed. DefenseClaw wraps cisco-ai-skill-scanner and writes its verdicts into the same skill_actions admission policy as the watcher.
Agents are getting comfortable installing skills. A "skill" is a small bundle — instructions, optional tools, sometimes auto-applied — that an LLM picks up and follows. The same shape is also a great place to hide a prompt-injection or an exfiltration tool.
DefenseClaw integrates Cisco's open-source cisco-ai-skill-scanner so operator-installed and modified skills get a deterministic + LLM-assisted review. Verdicts feed the skill_actions admission policy by severity bucket, and the install watcher scans every new or changed skill and quarantines it when the policy says so. Exact proven vendor bundles remain visible as discovery-only inventory and do not enter scanner or enforcement paths.
Guided example · Synthetic skill bundle
Catch a malicious skill and quarantine it
The watcher scans a new skill where it landed; static and optional intent checks feed the skill admission policy, which quarantines it.
{ "skill": "workspace-helper", "severity": "critical", "findings": ["path_escape", "external_exfiltration_intent"], "file": "quarantined", "runtime": "disabled", "install": "blocked"}{ "skill": "workspace-helper", "severity": "critical", "findings": ["path_escape", "external_exfiltration_intent"], "file": "quarantined", "runtime": "disabled", "install": "blocked"}
CRITICAL skill findings
Move to quarantine, disable runtime, block install, write audit event
What DefenseClaw did — and did not do
What it did
- Scan the skill where it was installed
- Combine deterministic checks with optional LLM analysis
- Quarantine on the skill_actions verdict
What it did not do
- Hide the skill from the agent while the scan runs
- Treat LLM analysis as mandatory or sufficient alone
- Skip the audit trail for a manual allow
What you just saw
The watcher scanned a new skill where it was installed. Deterministic checks and optional LLM-assisted intent analysis produced a CRITICAL result, which skill_actions mapped to quarantine (the files move out of the skill folder), disabled runtime, blocked installation, and an audit event.
What it scans
The wrapper at cli/defenseclaw/scanner/skill.py accepts four target shapes — all positional except --path. URL fetches in particular take the URL itself as the positional target (no --remote involved). The --remote flag does exist, but for a different scenario: it forwards the scan request to a sidecar API for a skill installed on a remote host (see the "Remote sidecar" tab below).
- Local skill name (
my-skill) — resolves through configured connector skill directories. On a multi-connector install, a bare name checks matching connector copies; pass--connector <name>to force one connector's skill source. --path <dir>— scan a directory you specify, useful inside a skill repo's pre-commit hook.https://...— fetch a remote skill bundle into a temp dir, scan, then clean up. The downloaded bytes never touch your skill directory.clawhub://author/skill@version— same fetch-to-temp flow, but resolves through the configured Clawhub registry source.
For each skill it runs:
- Static checks — manifest validation, allow-list of tool names, suspicious filesystem paths.
- Optional LLM-assisted analysis — by default it runs when a model is configured for
scanners.skilland is off otherwise;--use-llm/--no-use-llmturns it on or off for one scan. The upstream scanner auto-detects the provider from the LiteLLM-shapedprovider/modelstring, so OpenAI, Anthropic, Bedrock, Gemini, Vertex AI, Azure, Groq, Mistral, DeepSeek, OpenRouter, Ollama, vLLM, and other LiteLLM-supported providers all work without special-casing. - A consolidated
ScanResultwith severity, findings, and (when applicable) a recommended action.
One-shot scan
defenseclaw skill scan --allFor automation, add --json. Batch JSON remains a single top-level array for
backward compatibility. Successful results and per-skill error rows are always
written before the command returns nonzero for a scanner failure; a telemetry
recording warning is included on the affected result but does not fail or erase
the completed scan.
Walks every configured connector's skill directory, scans each bundle, prints connector-tagged findings, and exits non-zero if a scan fails. --all is also the default when you give no target. Add --json to pipe into CI. Use --connector <name> to narrow the scan to one connector.
defenseclaw skill scan --path ./my-skillUseful inside a skill repo's pre-commit hook — fail the build if the skill regresses.
defenseclaw skill scan https://github.com/some-org/skill-bundle/archive/refs/heads/main.tar.gzPass the URL of a .tar.gz or .zip archive as the positional target. The CLI downloads it into a temporary directory, extracts it with size limits, scans it, and deletes the directory. A repository page URL is not an archive and is rejected. Note: --remote is for a different scenario (see the next-to-last tab).
defenseclaw skill scan clawhub://author/skill@1.2.0Resolves through your configured Clawhub registry source. URL and clawhub:// scans are pre-screening only: --action is rejected for them.
defenseclaw skill scan my-skill --remote--remote means "POST the scan request to a remote sidecar over HTTP" — the scanner runs on a remote host (e.g. an SSM port-forward). Useful when the skill lives on a different machine than the operator CLI.
Configure the scanner
defenseclaw setup skill-scanner chooses the analyzers, the scan policy and the
optional external checks. Run it bare for the wizard, which asks about each
analyzer in turn, or pass flags with --non-interactive:
defenseclaw setup skill-scanner # interactive
defenseclaw setup skill-scanner \
--non-interactive \
--use-behavioral \
--use-llm --llm-provider anthropic --llm-model claude-sonnet-4-5 \
--enable-meta \
--policy balanced| Flag | What it turns on or sets |
|---|---|
--use-behavioral | Behavioral (dataflow) analyzer. |
--use-llm | LLM analyzer (semantic analysis). |
--enable-meta | Meta-analyzer that filters false positives (with the LLM analyzer). |
--use-trigger | Trigger analyzer that flags vague skill descriptions. |
--use-virustotal | VirusTotal check of bundled binaries. The key is read from VIRUSTOTAL_API_KEY. |
--use-aidefense | Cisco AI Defense analyzer. |
--llm-provider anthropic|openai, --llm-model <id> | LLM for the analyzer. Other providers are set with defenseclaw setup llm --provider .... |
--llm-consensus-runs <n> | Run the LLM analysis n times and combine the results (0 turns it off). |
--policy strict|balanced|permissive|none | Scanner policy preset. |
--lenient | Tolerate malformed skills instead of failing them. |
--verify / --no-verify | Run connectivity checks after setup (default on). |
With --non-interactive, the analyzer flags only turn analyzers on and leave
the others as they are. To turn one off, run the wizard.
The LLM provider and model are written to the unified top-level llm: block,
which the MCP and plugin scanners and the guardrail judge share. Cisco AI
Defense settings stay in cisco_ai_defense.
For VirusTotal, store the key before you turn the check on (the wizard asks for it and saves it the same way):
defenseclaw keys set VIRUSTOTAL_API_KEY
defenseclaw setup skill-scanner --non-interactive --use-virustotalContinuous protection: the install watcher
Running a one-shot scan is a starting point. The real value is the install watcher. It runs inside the gateway, is on by default, and watches the skill folders of every configured connector:
List view for small screens. Use the expand button to open the drawing.
- New or changed skillin a watched folder
- watcher checksBlock or allow list?
- Block or allow list?checked before scanning
- allow listAllowed
- neitherSkill scanner
- block listQuarantined
- Allowednot scanned
- Skill scannerstatic + optional LLM
- severityskill_actions
- scan failsQuarantined
- skill_actionsmaps the top severity
- no blockStays installed
- blockQuarantined
- Quarantinedand install blocked
- Stays installedclean, or a warning
The order is check the lists → scan in place → act on the verdict. The skill is not hidden while the scan runs (scans time out after five minutes), so an agent session that is already open can see it in that window. A scan that fails counts as a block. Watcher implementation lives in internal/watcher/watcher.go.
You don't have to wire any of that yourself. To record the severity verdicts without acting on them, set gateway.watcher.skill.take_action: false in ~/.defenseclaw/config.yaml; block-list matches and failed scans are still quarantined.
Configure the action mapping
The scanner's verdict is a severity. What that severity does lives in the active OPA policy file (policies/<name>.yaml) under skill_actions. The schema is real per-severity buckets (not threshold floors):
admission:
scan_on_install: true
allow_list_bypass_scan: true # first-party skills bypass scanning when the allow_list matches
skill_actions:
critical:
file: quarantine
runtime: disable
install: block
high:
file: quarantine
runtime: disable
install: block
medium:
file: none # no filesystem action
runtime: enable # agent can still call it
install: none
low:
file: none
runtime: enable
install: none
info:
file: none
runtime: enable
install: none
first_party_allow_list:
- target_type: skill
target_name: codeguard
reason: first-party DefenseClaw skill
source_path_contains: [".openclaw/workspace/skills/codeguard", ".openclaw/skills/codeguard", ".zeptoclaw/skills/codeguard", ".claude/skills/codeguard"]Each severity bucket has three independent actions:
file—noneorquarantine. Whether to move the skill bundle out of the agent's skill directory.runtime—enableordisable. Whether the agent can invoke the skill at all.install—noneorblock. Whether to refuse a fresh install.
first_party_allow_list is a policy-file key (not ~/.defenseclaw/config.yaml) — bundles whose source path contains any of the listed substrings are exempted from skill_actions. The shipped file also lists the first-party DefenseClaw plugin. See Policies for the full schema and Defaults for what each shipped rule pack sets these to.
Vendor-bundled skills are inventory only
A few skill folders hold copies the agent vendor ships, and DefenseClaw lists
them without scanning or enforcing on them. For Codex
that is only the exact vendor cache at $CODEX_HOME/skills/.system; every
other Codex skill, including CodeGuard under .agents/skills, is scanned. For
Hermes, a skill under $HERMES_HOME/skills is
inventory only when it matches the installer's .bundled_manifest and the
bundled source exactly; modified, extra or unverifiable copies are scanned.
aibom scan counts these skills as discovery-only.
To switch profiles for the entire skill_actions table:
defenseclaw policy activate strict # quarantines medium+ — see Defaults page for the diff
defenseclaw policy activate default # quarantines high+
defenseclaw policy activate permissive # quarantines critical onlyProvide the LLM key (optional but recommended)
LLM-assisted analysis is what catches the new malicious skills you don't yet have a regex for.
defenseclaw keys set DEFENSECLAW_LLM_KEYUnder the hood DefenseClaw injects the unified key into the scanner subprocess as SKILL_SCANNER_LLM_API_KEY and SKILL_SCANNER_LLM_MODEL, plus the matching provider-native variable (OPENAI_API_KEY, ANTHROPIC_API_KEY, AWS_BEARER_TOKEN_BEDROCK, GOOGLE_API_KEY, AZURE_OPENAI_API_KEY, GROQ_API_KEY, etc.) via inject_llm_env. See Unified LLM key for the resolution order and Bifrost provider catalog.
setup skill-scanner --llm-provider offers only anthropic and openai. To use another provider, configure the unified LLM with defenseclaw setup llm --provider .... LLM mode works with any LiteLLM-supported provider. The upstream Skill Scanner SDK auto-detects the provider from the LiteLLM-shaped provider/model string in your unified config (e.g. bedrock/anthropic.claude-3-5-sonnet-20240620-v1:0, gemini/gemini-2.5-pro, vertex_ai/gemini-1.5-pro). For self-hosted endpoints (vLLM, LM Studio, LiteLLM proxy) set llm.base_url in ~/.defenseclaw/config.yaml.
Availability and recovery
cisco-ai-skill-scanner is a hard runtime dependency, and supported
release-managed runtimes include it. If it is absent from a packaged runtime,
treat the installation as damaged; do not patch its managed virtual
environment with pip. Use the
supported install or repair path. Source
contributors should synchronize the checkout's locked environment with
uv sync.
If the SDK is missing when you run a scan, DefenseClaw prints a concise dependency error and exits cleanly; it does not crash the gateway.
Manage individual skills
The watcher's verdict is the default. Operators always have manual override:
defenseclaw skill list # on-disk skills + state, across ALL active connectors
defenseclaw skill list --connector codex # one connector
defenseclaw skill scan my-skill --connector codex # scan Codex's matching copy
defenseclaw skill block my-skill --reason "untrusted publisher"
defenseclaw skill block my-skill --connector codex --reason "Codex-only override"
defenseclaw skill quarantine my-skill --reason "review pending"
defenseclaw skill restore my-skill # un-quarantine
defenseclaw skill allow my-skill --reason "vetted"
defenseclaw skill unblock my-skill # clear block, file and runtime state
defenseclaw skill disable my-skill # runtime off, file untouched
defenseclaw skill enable my-skill
defenseclaw skill info my-skill # detailed viewTo install a skill through the scanner, use defenseclaw skill install <name>. It installs a copy into each configured connector's skill directory and scans each copy. By default it only reports findings; pass --action to apply skill_actions. Pass --force to overwrite an existing skill and --connector to install for one connector. defenseclaw skill search <query> searches the ClawHub registry with a locally installed or npx-cached clawhub client. Add --allow-remote-fetch only if you accept that npx downloads and runs the clawhub package from the npm registry.
Without --connector, policy and runtime commands apply to matching configured connector copies where present or the unscoped fallback policy where that command uses one. Pass --connector <name> when the same skill exists under more than one connector and you want only that connector's copy.
skill list, skill info, skill scan and the AIBOM inventory report the same policy verdict: blocked, quarantined or allowed from an explicit action, otherwise rejected, warning or clean from the latest scan and the skill_actions mapping. The list's Status column only says whether the skill is loaded (ready), removed from disk (removed) or missing. A rejected verdict means the policy refuses the skill at install time. A copy that is already on disk stays loaded until you run skill block, skill disable or skill quarantine, or accept it with skill allow.
Every action is audited (skill-block, skill-quarantine, etc.) so the trail is intact even when the manual override contradicts the watcher.
See also
- MCP scanner — sibling scanner for MCP servers, same admission pattern via
mcp_actions - Unified LLM key — the env var the LLM-assisted analysis reads
- Defaults — the
skill_actionstable for each shipped rule pack - Policies — the layered architecture (admission → scanner → action mapping)
- Reference → CLI — full
defenseclaw skillsubcommand reference
Regex cookbook
RE2 regex patterns for DefenseClaw guardrails, quoted from the shipped strict pack — secrets, prompt injection, exfiltration endpoints and enterprise data — with what makes each one precise.
MCP scanner
Behavioural scan of every Model Context Protocol server an agent might call. DefenseClaw wraps cisco-ai-mcp-scanner to surface hidden tool intents and shadow capabilities, and maps each verdict through the mcp_actions admission policy.