Writing

The Worst Bugs Said Nothing: Building a Tool That Checks AI Plugins

What fourteen rounds of adversarial testing taught me while building neckbeard, a tool that reads a Claude Code plugin before you install it. The dangerous failures were never wrong answers. They were silences.

A Claude Code plugin can put instructions into every session you start, and often into every subagent those sessions spawn, for as long as it stays installed. Earlier this month I nearly installed one that would have added over a thousand tokens to every session. I asked an LLM to compare its rules with mine, and it gave me an answer, but I had no good way to tell whether the answer was right.

So I built neckbeard to make that comparison something you can check. It lists everything a plugin injects and when, whether it reaches subagents, and what it writes to disk. Then it sorts every rule the plugin carries against the rules you already run (duplicate, conflict, or new) and cites the rule on each side, so every call can be checked by hand. It never edits a rule file. It writes a one-page verdict and stops.

A bearded developer under a desk lamp, arms crossed, scowling at a terminal full of struck-out lines.

Making it trustworthy took fourteen rounds of adversarial testing, and the repo publishes every one of them, including what each round cost the tool. This post is about what those rounds taught me, because most of it applies well beyond plugins.

The worst failures were silences

I expected wrong answers, like a hook counted twice. Those were easy. The failures that mattered said nothing.

In one round, neckbeard couldn’t read a hook that gated every shell command, and the report didn’t mention it: the warning only fired if no hook could be read. The fix made it worse. It skipped a popular plugin’s hook entirely, and told the model such hooks had “nothing resolvable to inject”.

A false alarm costs five minutes. A silence costs your trust, and you never know it happened. neckbeard now has a third answer besides yes and no: I could not read that.

The builder cannot audit the builder

For eight rounds I tested neckbeard myself. Then a separate Claude session, told to treat every claim in the docs as a hypothesis, found about a dozen problems I had missed.

The worst was in the output. neckbeard copied my private rules into its working file, about 90 KB of them, one git add from a public commit. I had checked what it read, never what it wrote.

Four published claims were false in the same way: I fixed the one case I’d built, then wrote a general sentence. Widen the claim to match the fix, or narrow it to match the test.

Every reviewer was an AI agent, not a human auditor. It still found more in one pass than I had in eight.

A guard nobody has seen fail is only a claim

Before going public, the self-check was all green. A reviewer inverted the line that decides whether to read your rules, then put back a bug we had just fixed. The suite stayed green both times: every test checked a helper, and none ran the code that decides.

Now every guard gets broken on purpose, and a check that stays green through its own sabotage counts as a finding.

The half that resists testing goes untested

neckbeard has two halves. A script reads files, and a filesystem can prove it wrong. The instructions tell the model how to sort each rule, and nothing on disk can check that. For twelve rounds, only my say-so did.

So a separate session built sealed test cases and graded the runs. I have never seen the answers. On a later re-run, deliberately weakened instructions scored as well as the real ones: held-out tests wear out as models improve, and weakened controls are how you notice. On a harder second set, the instructions scored 30 of 31 where the model without them scored 16.

Who it is for

neckbeard is open source under MIT, and you can install it as a plugin or as a plain skill. The repo has a gallery of real reports on eight public plugins and skill packs, including one of neckbeard vetting itself, which marks it down for its own report format. I left that finding standing.

My collaborator Selva tried the release on a skill from a large public pack. His verdict: “Instead of prose, this quickly lets me evaluate a skill.” His next message was a feature request for a security-focused scan, and he is building it.

If you install AI plugins, read them first. If you build tools that check things, break your own checks before you trust them, and pay most attention to the answers that come back empty.