What the checks flag on the most popular listings of every source they read · each figure from its dated measurement record · none of them is a catch rate
What this is. Every roster RepoGates runs was measured, before it shipped, on the most popular listings of its own platform — things nobody has reported as malicious. What a roster flags on that population is its false-positive side: the warnings and blocks a developer would meet on ordinary, well-known code. This page puts those runs side by side, with the date each was measured, the misreads we found when we read the flagged rows by hand, and what we changed because of them.
None of these figures is a catch rate. A catch rate needs a labelled set of malicious listings that are still there to be read. On most of these platforms there is no such set: the published VS Code campaigns were removed before we measured, no malicious Docker Desktop extension has been publicly reported, gitlab.com has no campaign of the FakeGit shape on the record, and the named ClawHub listings are gone. A PASS below means the checks on that tier found nothing to flag — never that the code was read, and never that the listing is harmless.
A verdict is PASS, REVIEW or BLOCK (what a verdict means). A REVIEW is not a false alarm by default: most REVIEWs on a popular population are true statements — a last release over a year old, an executable committed to the tree, an agent config file — that a developer may reasonably decide not to care about. So each run below gives two things: the verdict split as measured, and, where the record read the flagged rows by hand, how many were misreads — the check matched something it was not built to match. Where a record did not grade a row, this page does not grade it for it.
Three caveats hold for every row. The populations are the most popular listings, which are old and established; a roster flags a young, obscure listing far more often, by design. Several rosters were calibrated on the first run and measured again on the same population, so the second figure is not a held-out number — both are printed. And every run was made through the same code that answers the verdict API, from a laptop, on the date shown; a listing changes, and a verdict is only as old as its scan.
| Source | Population | Measured | PASS / REVIEW / BLOCK | Misreads found | Record |
|---|---|---|---|---|---|
| VS Code Marketplace | 100 most-installed extensions | 18 Sep 2026 | 99 / 1 / 0 | 0 BLOCK; the one REVIEW is a true finding | docs/32 |
| Docker Desktop | all 50 listed extensions — the whole Marketplace | 18 Sep 2026 | 50 / 0 / 0 | none to find | docs/33 §6 |
| GitLab | 100 most-starred gitlab.com projects | 19 Sep 2026 | 14 / 84 / 2 | not graded; see below | docs/34 §5 |
| GitHub (the control) | 100 most-starred repositories | 19 Sep 2026 | 0 / 96 / 4 | 3 of the 4 BLOCKs came from a rule changed since; see below | docs/34 §5 |
| skills.sh | 189 skill keys | 20 Sep, restated 27 Sep 2026 | 153 / 8 / 28 | 0 of 36 REVIEWs on 20 Sep; the 28 BLOCKs are a product decision | docs/36 §9, §11 |
| Anthropic's official plugin marketplace | all 310 plugins | 20 Sep, restated 27 Sep 2026 | 283 / 27 / 0 | 8 of 310, every one a REVIEW | docs/36 §9, §11 |
| Anthropic's community plugin marketplace | first 100 plugins | 20 Sep, restated 27 Sep 2026 | 94 / 5 / 1 | 2 of 100, both a REVIEW | docs/36 §9, §11 |
| ClawHub | the platform's top 100 by downloads (it answers 99) | 25 Sep 2026 | 83 / 16 / 0 | 1 of 99, a REVIEW | docs/36 §10 |
| Agent check-up | all 310 official plugins, installed one at a time | 21 Sep 2026 | run action: allow 287 / warn 23 / block 0 | 0 critical; 1 of 26 medium findings | docs/38 §14 |
| Agent check-up | first 100 community plugins | 21 Sep 2026 | run action: allow 90 / warn 10 / block 0 | 0 critical; 1 of 10 medium findings | docs/38 §14 |
The records are RepoGates' own measurement files, kept with the code; every figure on this page is quoted from the one named, and the platform's intelligence page prints the same figures with their context. VS Code Marketplace and Docker Desktop verdicts are answered through the MCP server and the Claude Code plugin, not in the browser. The Hugging Face roster's run on the 100 most-downloaded models (12 September 2026) is published on its own page and has no numbered record of this series, so it is not restated here.
The 100 most-installed extensions, 18 September 2026, through the
same code as /v1/vsx/score: 99 PASS, 1 REVIEW,
0 BLOCK, no partial scans, every listing read, median 1.1 s. The
REVIEW is abusaidm.html-snippets — an established
listing whose publisher has no verified domain and declares no
source repository, a medium finding under V5 that is true as
stated.
What we got wrong first. The first run of the same day read
87 / 13 / 0. Thirteen REVIEWs came from three shapes that are
ordinary on this Marketplace: executables inside a dependency under
node_modules/ (four extensions), closed-source extensions
from domain-verified publishers such as Copilot and Claude Code
(four), and agent config files shipped inside the package (five),
which an agent never opens as a workspace. Each was regraded to a
note on an established listing and kept as a finding on a new one —
the 2025 campaign hid its JavaScript under
node_modules/, so the signal was narrowed, not dropped.
Details on the VS Code
page.
Every one of the 50 extensions in Docker's Marketplace index, 18
September 2026, anonymously against the registry, no image pulled:
50 PASS, 0 REVIEW, 0 BLOCK, 0 errors, 0 partial scans. That is
not the roster being silent. 21 extensions declare host binaries, 29 a
backend inside Docker Desktop's VM (12 of them mounting the Docker
socket), 39 come from a publisher without a verified badge and 33
were not pushed in a year — and every one of those is recorded as a
note, because each is listed in the Marketplace and at least 853 days
old. The same shapes on an image that is not in the Marketplace are
findings: the control library/nginx, a Hub image that is
no extension, reads REVIEW 64. Docker paused new Marketplace
submissions on 16 June 2026, so this population is closed; the run is
the whole of it. Details on
the Docker Desktop
page.
The 100 most-starred gitlab.com projects, 19 September 2026, anonymously, through the same 22 checks RepoGates runs on GitHub: 14 PASS, 84 REVIEW, 2 BLOCK, 12 partial scans, 0 errors. The same engine the same day on GitHub's 100 most-starred repositories, the control: 0 PASS, 96 REVIEW, 4 BLOCK. The REVIEW rate belongs to the roster, not to either host: on the largest, oldest, busiest projects the 22 flag execution surface that is ordinary there — submodules, wrapper JARs, committed binaries, stale releases, agent config files. On GitHub the star-velocity note alone fires on 99 of the 100.
The record does not grade these rows true or false, and neither does
this page. The two GitLab BLOCKs are gitlab-org/cli and
gitlab-org/gitaly, each carrying a committed
.git directory in its test fixtures — the
nested-repository
class, CRITICAL because a recursive clone can run its hooks. On
GitHub, three of the four BLOCKs were the credential-redirect pattern
inside an AGENTS.md and one was an owner account under
the critical age.
What changed since, and what is not re-measured. On 21 September 2026 the credential-redirect check (C21) was changed to fire on an assignment of a base-URL variable to another host, not on the variable's name alone — the name-only match was the largest single source of false BLOCKs in the skill and check-up runs below. Three of the control's four BLOCKs came from that rule. The GitHub control has not been run again since, so this page prints the 19 September split and no after-figure.
What we got wrong first. The first GitLab run of the day
scored gitlab-org/gitlab — 136,997 tree entries — a
clean PASS 100: the file list comes in pages, the first page was full,
and a truncated list was recorded as a note rather than as a partial
scan. Twelve of the 100 had more than one page. A full page is now a
partial scan, cached briefly and never a clean PASS; four of the 14
PASS above are partial, and say so. There is no GitLab catch rate:
the record holds no gitlab.com campaign to measure one against.
Details on the GitLab page.
Three catalogues, measured 20 September 2026 on the 15-check roster after its first calibration: skills.sh (189 keys) 153 / 36 / 0, Anthropic's official plugin marketplace (all 310) 282 / 28 / 0, the community marketplace's first 100 94 / 6 / 0. No false BLOCK on any of the three — that was the gate, and it was met under the rule of that day. 159 of the rows were partial scans: this tier reads at most twelve files per key at 64 KB each, and eight skills.
Every REVIEW was read by hand. On skills.sh none of the 36 was a
misread: 28 were five young-owner repositories and their listings
riding on the source repository's own verdict, two were archives
beside a SKILL.md, six were one author's name
mismatches. On the official marketplace 8 of 310 were misreads
— an installer from the vendor's own short domain, a sentence
describing what a tool blocks, defence text telling the agent
to ignore instructions it finds, a placeholder URL read as a
destination. On the community hundred, 2. Every one a REVIEW,
never a BLOCK; each is recorded with the rule that would close it, and
none is fixed yet.
What changed on 27 September. The owner decided that a skill is its source repository: when the repository's own 22-check verdict is BLOCK, so is the skill's (S1). Re-measured that day, the three lists read skills.sh 153 / 8 / 28, official 283 / 27 / 0, community 94 / 5 / 1. All 29 BLOCKs are a young owner's source repository, a finding the 20 September read judged real. They are not misreads; they are a policy, and the false-positive case is stated plainly on the S1 page: a first project on a new account cannot be told from a lure on the day of the install, and an override records the reason. That is why this page never prints "0 BLOCK" for the skill lists.
What we got wrong first. The first run, the same morning, read the official marketplace 49 / 256 / 5. The lookalike check fired on 250 official and 98 community plugins, because a catalogue entry was compared with itself; the credential-redirect check blocked on a variable's name; and every plugin key had answered "not assessable" in production, because the marketplace catalogue — 185,864 bytes for the official one — was read through a 64 KB per-file cap. Details on the malicious skills page.
The platform's own top 100 by downloads, 25 September 2026, read from
ClawHub itself — the platform answers 99 items for a limit of 100, and
the population is what it returns: 83 PASS, 16 REVIEW, 0 BLOCK,
13 partial scans. Fifteen of the sixteen REVIEWs are accurate
statements about what the listing publishes — a declared name outside
the specification, a SKILL.md whose frontmatter does not
parse, a name identical to a known skill under another owner. One is a
misread: 1 of 99. A skill that vets other skills carries a
"reject immediately if you see" list, and the credential-read shape
matched a line of that list. It is the same defence-vocabulary gap as
on the official marketplace, pinned as a known misread and not yet
fixed. ClawHub's own moderation flagged none of the 99.
The agent check-up reads what an AI-agent installation already has and sends RepoGates identifiers, never contents. It was measured by installing each plugin of Anthropic's official marketplace — all 310 — and the first 100 of the community marketplace, one at a time, into a synthesised home, and running the product on each. The gate was 0 critical false positives on the official catalogue, or the rule is downgraded before launch.
On 21 September 2026, after calibration: no critical and no high finding on either catalogue, and no run a block. Official: allow 287, warn 23. Community: allow 90, warn 10. Of the 26 + 10 medium findings, 34 are real — every one a launch line that starts a package with no version, so whatever the registry serves next is what runs — and 2 are the same shared defence-text misread as above, one per catalogue.
| Check (severity) | Fires on the official catalogue | Misreads |
|---|---|---|
| Credential redirect (critical) | 0 | 0 — the first run: 58 fires, 58 false |
| Hidden Unicode (critical) | 0 | 0 — the first run: 1 false, an emoji joiner |
| MCP server launch line runs a fetched script (critical) | 0 of 67 launch lines | — |
| Hook runs a fetched script (critical) | 0 of 181 command hooks | — the community hundred's first run: 1 false, a || read as a pipe |
| MCP package unpinned (medium) | 25 | 0 — 25 real |
| Instruction override (medium) | 1 | 1 of 1 — defence text |
What we got wrong first. The first run on the same catalogue found 59 critical findings, all 59 false: 58 were the credential-redirect variable named in a plugin's text, one an emoji sequence read as hidden Unicode. 307 of the 308 runs warned, and one blocked. The fixes are the ones described above — an assignment, not a name; an emoji joiner is not hidden text; a double bar is the shell's or, not a pipe — and the last one was an engine fix that the GitHub, skill and check-up rosters had all shared. A deliberately hostile fixture, run beside the catalogues, still scores 0 and still blocks, which is how we know the calibration did not blind the real shapes. The check-up matches these shapes on configuration it reads; what a plugin's code does when it runs is the deep scan's question.
A false-positive run tells you what a roster costs a developer on ordinary code. It does not tell you what the roster catches, and nothing on this page should be read as if it did. Where the record tried to measure the other side:
A low false-positive rate on popular listings is what a roster should have, and it is not evidence that a PASS is a clearance. These runs read metadata, file lists and a capped number of text files; the code of a repository, a package or an image is not read on this tier, and a PASS says the provenance and the declared surface raised nothing — never that the code was read. Each figure is dated, and each will move as the rosters and the platforms do: when a roster changes a rule that moves a number, the number is measured again and this page says which date it is.