False-positive rates

What the checks flag on the most popular listings of every source they read · each figure from its dated measurement record · none of them is a catch rate

What this is. Every roster RepoGates runs was measured, before it shipped, on the most popular listings of its own platform — things nobody has reported as malicious. What a roster flags on that population is its false-positive side: the warnings and blocks a developer would meet on ordinary, well-known code. This page puts those runs side by side, with the date each was measured, the misreads we found when we read the flagged rows by hand, and what we changed because of them.

None of these figures is a catch rate. A catch rate needs a labelled set of malicious listings that are still there to be read. On most of these platforms there is no such set: the published VS Code campaigns were removed before we measured, no malicious Docker Desktop extension has been publicly reported, gitlab.com has no campaign of the FakeGit shape on the record, and the named ClawHub listings are gone. A PASS below means the checks on that tier found nothing to flag — never that the code was read, and never that the listing is harmless.

How to read the numbers

A verdict is PASS, REVIEW or BLOCK (what a verdict means). A REVIEW is not a false alarm by default: most REVIEWs on a popular population are true statements — a last release over a year old, an executable committed to the tree, an agent config file — that a developer may reasonably decide not to care about. So each run below gives two things: the verdict split as measured, and, where the record read the flagged rows by hand, how many were misreads — the check matched something it was not built to match. Where a record did not grade a row, this page does not grade it for it.

Three caveats hold for every row. The populations are the most popular listings, which are old and established; a roster flags a young, obscure listing far more often, by design. Several rosters were calibrated on the first run and measured again on the same population, so the second figure is not a held-out number — both are printed. And every run was made through the same code that answers the verdict API, from a laptop, on the date shown; a listing changes, and a verdict is only as old as its scan.

Every source, one table

SourcePopulationMeasuredPASS / REVIEW / BLOCKMisreads foundRecord
VS Code Marketplace100 most-installed extensions18 Sep 202699 / 1 / 00 BLOCK; the one REVIEW is a true findingdocs/32
Docker Desktopall 50 listed extensions — the whole Marketplace18 Sep 202650 / 0 / 0none to finddocs/33 §6
GitLab100 most-starred gitlab.com projects19 Sep 202614 / 84 / 2not graded; see belowdocs/34 §5
GitHub (the control)100 most-starred repositories19 Sep 20260 / 96 / 43 of the 4 BLOCKs came from a rule changed since; see belowdocs/34 §5
skills.sh189 skill keys20 Sep, restated 27 Sep 2026153 / 8 / 280 of 36 REVIEWs on 20 Sep; the 28 BLOCKs are a product decisiondocs/36 §9, §11
Anthropic's official plugin marketplaceall 310 plugins20 Sep, restated 27 Sep 2026283 / 27 / 08 of 310, every one a REVIEWdocs/36 §9, §11
Anthropic's community plugin marketplacefirst 100 plugins20 Sep, restated 27 Sep 202694 / 5 / 12 of 100, both a REVIEWdocs/36 §9, §11
ClawHubthe platform's top 100 by downloads (it answers 99)25 Sep 202683 / 16 / 01 of 99, a REVIEWdocs/36 §10
Agent check-upall 310 official plugins, installed one at a time21 Sep 2026run action: allow 287 / warn 23 / block 00 critical; 1 of 26 medium findingsdocs/38 §14
Agent check-upfirst 100 community plugins21 Sep 2026run action: allow 90 / warn 10 / block 00 critical; 1 of 10 medium findingsdocs/38 §14

The records are RepoGates' own measurement files, kept with the code; every figure on this page is quoted from the one named, and the platform's intelligence page prints the same figures with their context. VS Code Marketplace and Docker Desktop verdicts are answered through the MCP server and the Claude Code plugin, not in the browser. The Hugging Face roster's run on the 100 most-downloaded models (12 September 2026) is published on its own page and has no numbered record of this series, so it is not restated here.

VS Code Marketplace — 99 / 1 / 0

The 100 most-installed extensions, 18 September 2026, through the same code as /v1/vsx/score: 99 PASS, 1 REVIEW, 0 BLOCK, no partial scans, every listing read, median 1.1 s. The REVIEW is abusaidm.html-snippets — an established listing whose publisher has no verified domain and declares no source repository, a medium finding under V5 that is true as stated.

What we got wrong first. The first run of the same day read 87 / 13 / 0. Thirteen REVIEWs came from three shapes that are ordinary on this Marketplace: executables inside a dependency under node_modules/ (four extensions), closed-source extensions from domain-verified publishers such as Copilot and Claude Code (four), and agent config files shipped inside the package (five), which an agent never opens as a workspace. Each was regraded to a note on an established listing and kept as a finding on a new one — the 2025 campaign hid its JavaScript under node_modules/, so the signal was narrowed, not dropped. Details on the VS Code page.

Docker Desktop — 50 / 0 / 0, the whole Marketplace

Every one of the 50 extensions in Docker's Marketplace index, 18 September 2026, anonymously against the registry, no image pulled: 50 PASS, 0 REVIEW, 0 BLOCK, 0 errors, 0 partial scans. That is not the roster being silent. 21 extensions declare host binaries, 29 a backend inside Docker Desktop's VM (12 of them mounting the Docker socket), 39 come from a publisher without a verified badge and 33 were not pushed in a year — and every one of those is recorded as a note, because each is listed in the Marketplace and at least 853 days old. The same shapes on an image that is not in the Marketplace are findings: the control library/nginx, a Hub image that is no extension, reads REVIEW 64. Docker paused new Marketplace submissions on 16 June 2026, so this population is closed; the run is the whole of it. Details on the Docker Desktop page.

GitLab — 14 / 84 / 2, against GitHub's 0 / 96 / 4

The 100 most-starred gitlab.com projects, 19 September 2026, anonymously, through the same 22 checks RepoGates runs on GitHub: 14 PASS, 84 REVIEW, 2 BLOCK, 12 partial scans, 0 errors. The same engine the same day on GitHub's 100 most-starred repositories, the control: 0 PASS, 96 REVIEW, 4 BLOCK. The REVIEW rate belongs to the roster, not to either host: on the largest, oldest, busiest projects the 22 flag execution surface that is ordinary there — submodules, wrapper JARs, committed binaries, stale releases, agent config files. On GitHub the star-velocity note alone fires on 99 of the 100.

The record does not grade these rows true or false, and neither does this page. The two GitLab BLOCKs are gitlab-org/cli and gitlab-org/gitaly, each carrying a committed .git directory in its test fixtures — the nested-repository class, CRITICAL because a recursive clone can run its hooks. On GitHub, three of the four BLOCKs were the credential-redirect pattern inside an AGENTS.md and one was an owner account under the critical age.

What changed since, and what is not re-measured. On 21 September 2026 the credential-redirect check (C21) was changed to fire on an assignment of a base-URL variable to another host, not on the variable's name alone — the name-only match was the largest single source of false BLOCKs in the skill and check-up runs below. Three of the control's four BLOCKs came from that rule. The GitHub control has not been run again since, so this page prints the 19 September split and no after-figure.

What we got wrong first. The first GitLab run of the day scored gitlab-org/gitlab — 136,997 tree entries — a clean PASS 100: the file list comes in pages, the first page was full, and a truncated list was recorded as a note rather than as a partial scan. Twelve of the 100 had more than one page. A full page is now a partial scan, cached briefly and never a clean PASS; four of the 14 PASS above are partial, and say so. There is no GitLab catch rate: the record holds no gitlab.com campaign to measure one against. Details on the GitLab page.

Skills and plugins — 0 false BLOCK on 20 September, 29 BLOCKs by decision since

Three catalogues, measured 20 September 2026 on the 15-check roster after its first calibration: skills.sh (189 keys) 153 / 36 / 0, Anthropic's official plugin marketplace (all 310) 282 / 28 / 0, the community marketplace's first 100 94 / 6 / 0. No false BLOCK on any of the three — that was the gate, and it was met under the rule of that day. 159 of the rows were partial scans: this tier reads at most twelve files per key at 64 KB each, and eight skills.

Every REVIEW was read by hand. On skills.sh none of the 36 was a misread: 28 were five young-owner repositories and their listings riding on the source repository's own verdict, two were archives beside a SKILL.md, six were one author's name mismatches. On the official marketplace 8 of 310 were misreads — an installer from the vendor's own short domain, a sentence describing what a tool blocks, defence text telling the agent to ignore instructions it finds, a placeholder URL read as a destination. On the community hundred, 2. Every one a REVIEW, never a BLOCK; each is recorded with the rule that would close it, and none is fixed yet.

What changed on 27 September. The owner decided that a skill is its source repository: when the repository's own 22-check verdict is BLOCK, so is the skill's (S1). Re-measured that day, the three lists read skills.sh 153 / 8 / 28, official 283 / 27 / 0, community 94 / 5 / 1. All 29 BLOCKs are a young owner's source repository, a finding the 20 September read judged real. They are not misreads; they are a policy, and the false-positive case is stated plainly on the S1 page: a first project on a new account cannot be told from a lure on the day of the install, and an override records the reason. That is why this page never prints "0 BLOCK" for the skill lists.

What we got wrong first. The first run, the same morning, read the official marketplace 49 / 256 / 5. The lookalike check fired on 250 official and 98 community plugins, because a catalogue entry was compared with itself; the credential-redirect check blocked on a variable's name; and every plugin key had answered "not assessable" in production, because the marketplace catalogue — 185,864 bytes for the official one — was read through a 64 KB per-file cap. Details on the malicious skills page.

ClawHub — 83 / 16 / 0

The platform's own top 100 by downloads, 25 September 2026, read from ClawHub itself — the platform answers 99 items for a limit of 100, and the population is what it returns: 83 PASS, 16 REVIEW, 0 BLOCK, 13 partial scans. Fifteen of the sixteen REVIEWs are accurate statements about what the listing publishes — a declared name outside the specification, a SKILL.md whose frontmatter does not parse, a name identical to a known skill under another owner. One is a misread: 1 of 99. A skill that vets other skills carries a "reject immediately if you see" list, and the credential-read shape matched a line of that list. It is the same defence-vocabulary gap as on the official marketplace, pinned as a known misread and not yet fixed. ClawHub's own moderation flagged none of the 99.

The agent check-up — 0 critical false positives on 310 official plugins

The agent check-up reads what an AI-agent installation already has and sends RepoGates identifiers, never contents. It was measured by installing each plugin of Anthropic's official marketplace — all 310 — and the first 100 of the community marketplace, one at a time, into a synthesised home, and running the product on each. The gate was 0 critical false positives on the official catalogue, or the rule is downgraded before launch.

On 21 September 2026, after calibration: no critical and no high finding on either catalogue, and no run a block. Official: allow 287, warn 23. Community: allow 90, warn 10. Of the 26 + 10 medium findings, 34 are real — every one a launch line that starts a package with no version, so whatever the registry serves next is what runs — and 2 are the same shared defence-text misread as above, one per catalogue.

Check (severity)Fires on the official catalogueMisreads
Credential redirect (critical)00 — the first run: 58 fires, 58 false
Hidden Unicode (critical)00 — the first run: 1 false, an emoji joiner
MCP server launch line runs a fetched script (critical)0 of 67 launch lines—
Hook runs a fetched script (critical)0 of 181 command hooks— the community hundred's first run: 1 false, a || read as a pipe
MCP package unpinned (medium)250 — 25 real
Instruction override (medium)11 of 1 — defence text

What we got wrong first. The first run on the same catalogue found 59 critical findings, all 59 false: 58 were the credential-redirect variable named in a plugin's text, one an emoji sequence read as hidden Unicode. 307 of the 308 runs warned, and one blocked. The fixes are the ones described above — an assignment, not a name; an emoji joiner is not hidden text; a double bar is the shell's or, not a pipe — and the last one was an engine fix that the GitHub, skill and check-up rosters had all shared. A deliberately hostile fixture, run beside the catalogues, still scores 0 and still blocks, which is how we know the calibration did not blind the real shapes. The check-up matches these shapes on configuration it reads; what a plugin's code does when it runs is the deep scan's question.

Where a catch rate stands

A false-positive run tells you what a roster costs a developer on ordinary code. It does not tell you what the roster catches, and nothing on this page should be read as if it did. Where the record tried to measure the other side:

  • VS Code Marketplace. All 40 Microsoft-Marketplace exemplars named in the primary write-ups were removed before we measured; the roster meets each as "not assessable" — a warning, never a pass — and the campaign blocklist is what refuses a re-registered name. The live volume of that campaign family is on Open VSX, which RepoGates does not assess.
  • Docker Desktop. No malicious Docker Desktop extension has been publicly reported, so there is nothing to measure a catch rate against. The page is a threat model.
  • GitLab. No gitlab.com campaign of the FakeGit shape exists in the 2024–2026 record. No catch rate exists to publish.
  • Skills and plugins. Of 19 named incident keys re-measured on 25 September 2026, 13 were still there to assess: 9 PASS, 3 REVIEW, 1 BLOCK — the one BLOCK Snyk's own fixture repository. Several of those PASS are listings that may no longer be what they were when they were reported, and a tier that reads twelve files is not a reading of the whole package: that is the deep scan's job.
  • ClawHub. Nothing malicious was in the population, and there is no labelled ClawHub corpus to measure against; the named campaign listings are gone.

What this doesn't mean

A low false-positive rate on popular listings is what a roster should have, and it is not evidence that a PASS is a clearance. These runs read metadata, file lists and a capped number of text files; the code of a repository, a package or an image is not read on this tier, and a PASS says the provenance and the declared surface raised nothing — never that the code was read. Each figure is dated, and each will move as the rosters and the platforms do: when a roster changes a rule that moves a number, the number is measured again and this page says which date it is.