Frequently Asked Questions
What is Skill Scanner?
Skill Scanner is an open-source security scanner for AI agent skill packages. It detects prompt injection, data exfiltration, command injection, obfuscated code, and other threats using a combination of pattern-based detection (YAML + YARA), behavioral dataflow analysis, and LLM-powered semantic reasoning. Published as cisco-ai-skill-scanner on PyPI.
What skill formats does it support?
Skill Scanner supports OpenAI Codex Skills, Cursor Agent Skills, and related formats following the Agent Skills specification. With --lenient mode, it can also scan non-standard formats such as Claude Code .claude/commands/*.md files and any directory containing Markdown instruction files.
Do I need API keys to use it?
Not necessarily, but you do need a model for the LLM judge (--use-llm), which every recommended setup uses. That can be a hosted provider with a key, Bedrock or Vertex AI with cloud credentials, or a model you run locally with no key and no network (see LLM Providers). The deterministic rules need nothing, but on their own they catch only about 8% of malicious skills. OSV.dev scanning (--use-osv) needs network access but no key. VirusTotal (--use-virustotal) and Cisco AI Defense (--use-aidefense) need their own keys.
Which LLM providers are supported?
Anthropic (the default, anthropic/claude-sonnet-5-5), OpenAI, AWS Bedrock, Google Vertex AI and Gemini, Azure OpenAI, OpenRouter, OrcaRouter, Ollama, the on-device Apple Foundation Model (experimental), and any OpenAI-compatible gateway or local server through SKILL_SCANNER_LLM_PROVIDER=openai-compatible. Install the matching extra ([bedrock], [vertex], [google], [azure], or [all]) for managed cloud services; Apple FM needs pip install "apple-fm-sdk>=0.2.1,<0.3". See LLM Providers.
How do I choose which analyzers to enable?
Start from Recommended Settings, which picks a setup by goal with measured numbers.
- Always run the LLM judge (
--use-llm). It is what catches most malicious skills: held-out recall rises from 8% with the rules alone to about 67%. On a local model, nothing leaves your machine. - Your own skills: the judge with
--policy low-noise. - Third-party skills: the judge with
--policy balanced(highest F1) or--policy quiet(lowest false-positive rate). - Optional extras:
--use-osvchecks for known-vulnerable Python and JavaScript dependency pins.--use-behavioraladds Python dataflow analysis.- VirusTotal or Cisco AI Defense add cloud-based binary and content scanning.
- Leave
--enable-metaoff. It is off by default because it cost 16.4 points of recall in measurement.
What is the adjudicator?
The adjudicator is an optional pass between deterministic analysis and LLM enrichment. It asks an LLM to review deterministic HIGH/CRITICAL matches and can demote high-confidence literal-regex false positives to INFO. It never promotes a finding, and errors preserve the original severity. Enable it with --adjudicate.
What is the meta-analyzer?
The meta-analyzer is a second-pass LLM that reviews the findings from other analyzers, correlates them, and can filter ones it judges false positives. It is off by default, and we recommend keeping it off. Measured on held-out skills, it cost 16.4 points of recall for 0.3 points of false-positive rate. Enable it with --enable-meta only if you have measured it on your own skills.
What is LLM consensus mode?
Consensus mode runs the LLM analyzer several times and keeps only the findings a majority of runs agree on. It multiplies the cost by the number of runs. Its effect on false positives has not been measured on the evaluation splits, so check it on your own skills before relying on it. Enable with --llm-consensus-runs 3 (or any odd number).
Does a clean scan mean the skill is safe?
No. A scan that returns no findings means no known threat patterns were detected. It does not guarantee the skill is secure, benign, or free of vulnerabilities. Coverage is inherently incomplete, and novel attacks may evade all detection engines. Human review remains essential.
How does it handle binary files?
Binary files are classified into tiers based on their content type. Executable and opaque binaries lower the analyzability score. When VirusTotal is enabled (--use-virustotal), binary file hashes are checked against the VirusTotal database. Archives (ZIP, TAR) are extracted and inspected with protections against zip bombs, path traversal, and symlinks.
Can I add custom detection rules?
Yes. Skill Scanner supports three rule types: YAML signature rules (regex patterns), YARA rules (binary and text matching), and Python checks (programmatic analysis). Use --custom-rules /path/to/rules to load custom rules alongside built-in packs. See Writing Custom Rules.
How do scan policies work?
Policies control every aspect of scanner behavior through YAML configuration: which rules fire, severity levels, file limits, analyzability thresholds, and more. Five built-in presets are available: balanced (the default), low-noise and quiet, which report at LOW the rules whose real-world flags the LLM judge most often clears, strict (maximum sensitivity, for audits), and permissive. Only balanced, low-noise and quiet are measured as gates; see Recommended Settings. Generate and customize policies with skill-scanner generate-policy and skill-scanner configure-policy. See Scan Policies.
How do I integrate with CI/CD?
Use the reusable GitHub Actions workflow for zero-config integration:
jobs:
scan:
uses: cisco-ai-defense/skill-scanner/.github/workflows/scan-skills.yml@2.2.1
with:
scanner_version: "2.2.1"
skill_path: .cursor/skills
policy: low-noise
use_llm: true
llm_model: anthropic/claude-sonnet-5-5
secrets:
llm_api_key: ${{ secrets.SKILL_SCANNER_LLM_API_KEY }}
permissions:
security-events: write
contents: read
actions: read
Results appear as inline annotations in PRs via GitHub Code Scanning. Use --fail-on-severity high in any CI system to fail builds on findings. See GitHub Actions.
What output formats are available?
Six formats: summary (terminal), json (automation), markdown (reports), table (compact terminal), sarif (GitHub Code Scanning), and html (interactive triage). Repeat --format to produce several reports in one scan, and use per-format paths such as --output-json and --output-sarif.
Can I use it as a Python library?
Yes. Import SkillScanner, scan_skill, or scan_directory directly in Python. The SDK provides typed models for findings, results, and reports. See Python SDK.
How do I get the fewest false positives?
Use the LLM judge with the quiet preset and block at HIGH: 33.1% of held-out malicious skills blocked at a 5.5% false-positive rate. Review MEDIUM+ as well to reach 50.3% recall at 7.2%. Don't use quiet without the judge. For your own skills, scoped suppressions and severity_overrides remove the remaining noise without switching rules off; see Results and Tuning.
What setting gives the highest F1?
The balanced preset with the LLM judge, reviewing everything at MEDIUM or above: 66.7% recall at 15.4% false-positive rate, F1 about 75.5% on the held-out split. See Recommended Settings for every preset.
My skill was not analysed by the LLM. Is it clean?
No. A skill the judge could not read gets an INFO LLM_ANALYSIS_FAILED finding, and INFO never fails a scan. Treat it as unreviewed and route it to a person. Results and Tuning shows how to catch these in CI.
What is cross-skill analysis?
Cross-skill analysis detects coordinated attacks across multiple skills, including data relay patterns, shared external URLs, complementary triggers, and shared suspicious patterns. Enable with --check-overlap on scan-all.
What platforms are supported?
Skill Scanner supports CPython 3.11–3.14 on glibc Linux x86-64/ARM64, macOS 14+ x86-64/ARM64, and Windows x86-64. Other platforms install from the source distribution, which needs Go 1.27.1+.
How does it relate to DefenseClaw?
DefenseClaw is the governance layer that orchestrates Skill Scanner (and other Cisco AI Defense scanners) with enforcement, policy, and audit capabilities. Skill Scanner is the standalone scanning engine that DefenseClaw wraps.
How does it relate to the IDE AI Security Scanner?
The IDE AI Security Scanner VS Code extension integrates Skill Scanner for in-editor skill scanning. You can also use Skill Scanner standalone from the command line without VS Code.
Where can I report bugs or request features?
Open an issue on the GitHub repository. For security vulnerabilities, see SECURITY.md.
How do I contribute?
See CONTRIBUTING.md for guidelines. Clone the repo, install with uv sync --all-extras, and run pytest to verify your setup.