Case study
neckbeard
A Claude Code plugin that reads another plugin before you install it, then tests its own claims harder than most tools test anything.
- Role
- Author
- Period
- Sep 2026–present
- Status
- Open source (MIT)
The problem
Claude Code plugins inject instructions into every session, and often into every subagent, for as long as they stay installed. I nearly installed one that would have added 1,300 tokens to every session permanently. I asked an LLM to compare its rules with mine, and it gave me an answer, but I had no good way to tell whether the answer was right.
neckbeard makes that comparison something you can check. It lists everything a plugin injects and when, whether it reaches subagents, and what it writes to disk. Then it sorts every rule the plugin carries against the rules I already run: duplicate, conflict, or new. It never edits a rule file and never uninstalls anything. It writes an evaluation and stops.
How it works
The tool has two halves, and they are tested differently.
The mechanical half is a Python script with no dependencies. It reads the plugin’s manifest, skills, commands, agents and hooks, follows what each hook injects where a static read can, and discovers the rules already in force across every settings layer. A filesystem can prove it right or wrong.
The judgment half is a set of instructions to the model: how to classify each rule, what to cite, when a plugin’s carrying cost outweighs what it adds. No filesystem can check that, which is exactly why it needed a different kind of test.
How it was tested
The repo publishes fourteen rounds of adversarial testing, with what each one cost the tool and every published claim that turned out false.
- Real plugins, picked to break it. A large popular plugin, one picked at random from 183, Anthropic’s own security plugin, and a sweep of all 39 first-party plugins.
- Every reviewer was an AI agent. Each was a separate Claude session that had not written the code, which is the honest meaning of “independent” here. The first such review found eleven problems and proved four of the tool’s published claims false, including a path that could have leaked tens of kilobytes of private rules into a public repo.
- Two sealed held-out fixtures for the judgment half, each built and graded by a separate session that pre-registered its rubric. I have never seen the answers. The first stopped telling the instructions from the model once models improved, and the write-up says so. On the second, the instructions scored 30 of 31 where the same model without them scored 16.
- Sabotage. Every guard in the self-check has been broken on purpose and seen to fail. Before that rule, fifteen passing checks would have let a fixed bug return unnoticed.
What the project shows
- A tool that asks you to trust its measurements should publish its own, including the ones it got wrong.
- The dangerous failures in AI tooling are silences: an empty result that reads like an all-clear.
- A guard nobody has seen fail is a claim, not a guard.
- Held-out evaluation wears out as models improve, and controls are how you notice.
Read the write-up: what fourteen rounds of testing taught me.