Cloudflare Security Audit Skill: How a 450-Line Skill Becomes a 128-Repo Fleet Scanner — aniketkarneai.com | aniketkarneai.com
Sunday, September 27, 2026 Field notes on autonomous systems ● Amsterdam, NL
daily

Cloudflare Security Audit Skill: How a 450-Line Skill Becomes a 128-Repo Fleet Scanner

Cloudflare just published the open-source skill that seeds their internal Vulnerability Discovery Harness — six phases, three verdict types, sandboxed execution, and a findings.json schema designed so a separate agent literally cannot file its own bugs. This is the post about why the schema ordering matters more than the prompts.

Cloudflare Security Audit Skill: How a 450-Line Skill Becomes a 128-Repo Fleet Scanner

The thing I keep coming back to, after reading cloudflare/security-audit-skill cover to cover and then re-reading their “Build your own vulnerability harness” post for a third time, is the schema. Not the prompts. Not the seven-stage pipeline. Not even the SQLite-backed resumability. The thing that makes me trust their confirmed verdict is the way report-schema.json is ordered — verdict first, then fingerprint, then title, then description, then root_cause, then intended_behavior, then trace[]. The Hunter agent has to declare the threat model before it can file anything. If you skip the threat model, the schema rejects your finding. That single bit of ordering is doing more work than any of the prompt engineering.

The skill is MIT-licensed, six-phase, agent-neutral (it calls them parent / research / general / subagent_type, but the README explicitly says to use your platform’s equivalents). It runs against a single repository at a time. It writes seven files inside an output directory outside the target: run-metadata.json, architecture.md, coverage-ledger.json, findings.json, REPORT.md, FINDINGS-DETAIL.md, NEEDS-VALIDATION.md. Two zero-dependency Node.js validators (validate-coverage-ledger.cjs and validate-findings.cjs) gate the writes mechanically — schema adherence, not correctness, but a Hunter can’t quietly file confirmed without its cited file and line numbers resolving to a real, non-stale, single-link regular file under a byte budget. The validation block in the parent has eleven enumerated steps and reads like a thesis on why you should never trust a cp from a sandboxed directory.

What’s actually interesting is that this single-repo skill is the same one Cloudflare used as the seed for their internal fleet harness. The README links to the blog post. The blog post walks the progression: start with the ~450-line skill, lift each phase into its own agent, put a SQLite database behind it, put an orchestrator in front, run for six weeks, and you end up with 128 distinct repos under continuous scanning, 25,472 wishlist writes since launch, and 13,841 findings currently sitting in their Vulnerability Validation System. The skill is the single-layer prompt that enabled the next architecture. Without it, the harness would have been designed around different primitives.

The six phases, and why they aren’t seven

The README’s “What it does” lists six phases. Reconnaissance → Coverage-led hunting → Candidate validation → Structured output → Independent record verification → Target-neutral reporting. Each phase has a named file and a named validator. The seventh stage is implicit: the parent must update run-metadata.json whenever a fact changes (run_id, repo, target, source_ref, profile, scope_paths, budget, execution_policy, run_status). The schema makes it impossible to mutate candidate state from anywhere except the coverage ledger and findings.json — which means the only writers of those two are the parent. Hunters and verifiers can only touch their own scratch/ and read the parent’s shared files.

Phase 1 (recon) writes architecture.md and coverage-ledger.json. Three parallel research agents map the target’s trust boundaries, input surfaces, prior evidence, and deterministic coverage. The coverage-led hunting phase then assigns planned ledger units to isolated hunters — one hunter per attack class, no agent-count limit silently swallowing units. If the budget can’t reserve critics and validators for a unit, the unit is explicitly deferred with the reason budget_cannot_reserve_critics_and_validation. The unit is never invisible; it’s visible in the ledger with a state.

Phase 3 (validation) is where the schema earns its keep. The Hunter hands the candidate to a fresh verifier that tries to disprove it. The verifier is in a separate context, has no memory of the Hunter’s reasoning, and cannot log findings of its own. The verdict branches in report-schema.json are three: confirmed, needs_validation, rejected. The confirmed branch requires a trace[] with kind ∈ {entrypoint, propagation, sink}, a root_cause, an intended_behavior, a list of evidence[] file+line+description triples, and a list of conditions[] with kind ∈ {authentication_level, authorization_role, user_interaction, system_configuration, network_routing, environmental_dependency, data_state, timing_dependency, third_party_dependency}. If any of those fields is empty, the validator rejects the finding. The Hunter has to know the trust boundary and the precondition; it can’t just say “this looks weird.”

Phase 5 (independent record verification) is the second-pass verifier. Even after a candidate is confirmed, a fresh agent re-checks it against the source. If the source has changed since the original verification, the unit goes back to planned and a new ledger unit is created. Material replacements receive another independent verifier. The blog post makes the philosophy explicit: “If your harness doesn’t actively fight this, all you’ve built is a faster way to produce junk.”

The three-verdict contract and why needs_validation is a feature

Most security tooling I read collapses to a binary: “vulnerable / not vulnerable.” Cloudflare’s schema adds a third bucket: needs_validation. The semantic is “this looks like a real boundary failure but the decisive fact is outside source or sandboxed fixture.” The skill’s design principle page says it explicitly: “Only confirm established boundary failures. Keep a source-grounded blocked lead as needs_validation with its exact unresolved fact.” The needs_validation record ships a fingerprint, a description, and an exact unresolved fact. No severity. No priority. No patch. The reason is honest: the engineer who picks it up needs to know what’s missing, not what the model thinks the priority is.

This shows up operationally in two places. First, the wishlist: the blog post notes that 25,472 wishlist writes have been logged across 128 repos, and the example write was* “I need a FreeBSD VM to confirm this PoC end-to-end.” That’s a needs_validation finding that the system can’t close without human infrastructure. Second, the prior-run carry-forward logic in SKILL.md: a still-external needs_validation record from a prior run may be carried into the current run only after the current source trace is checked and linked by fingerprint to a current planned unit whose verifier re-check supplies its owner and evidence. The record keeps its unresolved blocker. Prior needs_validation records never suppress a current unit; they remain visible in the ledger. That’s the part most LLM pipelines miss — they treat “I don’t know” as a soft signal and drop it, when in a security audit “I don’t know” is exactly the record that needs to outlive the run.

The sandbox model and why scratch/ is non-negotiable

The Universal execution safety section in SKILL.md is the longest block of prose in the file, and it’s worth reading twice. The agent, outside the target-controlled process, may make a disposable source copy in an assigned scratch/ directory when a build must write beside source. But the agent itself, and every target-controlled process it spawns, may write only to scratch/. Retained artifacts/ is parent-owned and is never exposed to the sandbox. The promotion of a scratch file to retained artifacts is an eleven-step deterministic procedure that reads like a postmortem on race conditions: validate the declared relative path, walk each parent component from the retained scratch-root descriptor with no-follow directory-relative operations, never reopen by path, verify with fstat that it’s a regular file with link count exactly one, copy the verified size, repeat fstat, reject a changed identity, type, link count, or size. The destination walk is the same pattern. The validator disallows recursive globs, archive extraction, symlink-following, FIFOs, sockets, devices, directories, hard-linked files, files larger than bound, and any path that can’t be enforced. If any check is unavailable, the scratch entry is discarded; if it’s decisive evidence, the finding is retained as needs_validation with the exact promotion blocker.

The blog post explains the origin: Hunters move past code reading into active execution, and the quality jump came from giving them a sandbox built on unshare to crash binaries. There’s a one-paragraph aside that’s worth quoting because it’s the kind of footnote that saves a day of debugging: “If the harness itself runs inside Docker, that sandbox needs seccomp=unconfined and apparmor=unconfined or it will silently fail to start. It’s a one-line fix that saves you a day of head-scratching if you aren’t an expert in nested containerization, like us.” I have not seen that note in any other security-AI write-up. It’s exactly the kind of field knowledge that makes the gap between “the paper” and “the production system” shrinkable.

The actual numbers from the fleet run

The fleet is one unified harness, no per-language tuning, 128 distinct repos across Rust, Go, C, Lua, TypeScript, Python, plus various configuration management systems. The Vulnerability Discovery Harness has eight stages (the original seven plus the Sibling Forking and Wishlist mechanisms, which I’ll get to in a second). Stages four through eight run as a continuous producer-consumer loop: Gapfill enqueues new hunt tasks for empty coverage cells, Feedback rewrites queued prompts based on validation failures and shallow runs, Trace walks the dependency graph and spawns consumer-repo hunt tasks, and Report is just a script — no model. The cost-to-coverage lever is Gapfill: each additional pass costs roughly half as much as the initial hunt. The total compute is “budget per repo, not per run,” with a strict task cap per repository and a worker pool of 50 to 200 concurrent workers.

The VVS — Vulnerability Validation System — is the second half of the pipeline. It currently holds 13,841 findings across 145 repos. The Dedup stage is interesting because it refuses to scale LLM-to-LLM at O(N²). Instead, deterministic code builds inverted indexes over touched files, functions, trust boundaries, and rare tokens, generates a short candidate list, and then hands it to a Dedup agent that reasons over the short list. Stable cross-run keys mean a re-found bug reopens an existing one rather than spawning a duplicate. The Judgment stage pulls from MCP servers, Jira, the wiki, git, config, and other sources to score the bug against production reachability; it also rechecks against the latest main. The Fixing stage takes the proposed patch and unit tests, applies them, runs the test filtered to the affected test, and demands a clean fail→pass flip on the target test. If the post-patch test fails, the commit is blocked and flagged for human intervention. The Fixer never merges code on its own. A human reviews the dry-run branch.

The fleet’s most-used tool is the wishlist. The blog post notes that they plumbed Semgrep all the way through the harness and the Hunters invoked it zero times in a month of runs. They would rather read and run the code. The wishlist, by contrast, was the single most-used tool — agents wrote to it 25,472 times across 128 repos. “It’s worth paying attention to what the agents actually reach for, rather than what you think they’ll want.”

The sibling-forking and wishlist patterns, briefly

Two mechanisms sit alongside the eight pipeline stages. Sibling Forking is a tool call that lets a Hunter spawn a sibling agent with a precise structural seed when it trips over an interesting code path outside its current scope. The sibling handles the new path; the Hunter doesn’t lose focus. Fleet-wide, this accounts for roughly 9% of tasks, but the rate is highly model-dependent — from near-zero to about a fifth of tasks, depending on which model is hunting. If you swap the model, the sibling-fork rate changes. That’s the kind of detail that breaks the assumption “the harness is the bit that lasts, the model is interchangeable.” It’s true at the architecture level. It’s not true at the per-task behavior level.

The Wishlist is the agent’s escape hatch when it needs a tool it doesn’t have — often a Validator confirming a PoC or a Hunter wanting a specific build environment, a VM, or a prod config file. The agent writes a structured wishlist entry with enough context for the system to re-run that exact task once a human provides the dependency. Some wishes are self-healing: if the container needs to be rebuilt, a generic coding harness monitors the logs and triggers the rebuild. The wishlist is also the main way the agents talk back to humans. Reading the wishlist is the operationally honest feedback channel.

What this skill is and what it isn’t

This skill is the precursor to a fleet-grade harness. It’s not the harness. The blog post is explicit about that: “Start with a skill in your development environment, get your prompts working well, and only build the next architectural stage when not having it is the specific thing slowing you down.” The advice block at the end of the recon section is even more pointed: a real but minimal harness consists of just Recon, Hunt, and Validate stages kept in a database, alongside a separate Validator that can’t file its own findings. Skip cross-repo tracing until you have more than one repository that matters. Skip a dedicated Dedup agent until you are actively drowning in noise. The schema and the validators and the three-verdict contract are what you keep when you scale up; everything else is overhead until you need it.

What I think the skill gets right, and most LLM-pipeline write-ups miss: the needs_validation verdict, the schema ordering that forces the threat model before the file path, the eleven-step promotion procedure with byte budgets, the SQLite persistence keyed by (run_id, repo, stage), the requirement that every confirmed finding ships a PoC written as a test that runs against the untouched codebase, and the wishlist as the agents’ primary feedback channel. What I’d want to test before depending on it: how the skill behaves when the target repo’s build system needs network access to fetch dependencies — Do not install dependencies or let builds fetch them is in SKILL.md, but the friction of building an offline mirror for every language in the target’s polyglot tree is real and the skill is silent about it. And how the prior-run carry-forward interacts with a target where the source has been force-pushed since the last run — A prior source ref alone is not evidence that a path is unchanged is in the skill, but the machinery for detecting a force-push and invalidating all carried prior-runs in one pass is implicit, not explicit.

The skill is MIT, 13 commits since June, 16.4k stars, last rework on Sep 14. The blog post went up the same week. Cloudflare is publishing the seed, not the fleet, and the seed is enough to study.

Aniket Karne
DevOps & AI Engineer · Amsterdam
Back to all posts
Reader correspondence

Comments

Powered by GitHub Discussions via Giscus. Sign in with GitHub to leave a comment.