Inside mattpocock/skills
Inside mattpocock/skills, the 25 agent skills 270,000 GitHub users starred
Matt Pocock keeps the agent skills he uses every day in one public repo, and about 270,000 people on GitHub have starred it. Twenty five of those skills ship as a Claude Code plugin, and every one of them is a Markdown file you can read in a minute. This is a walkthrough of what's in the repo and how the pieces fit together. One thing to say up front: it is one engineer's workflow, opinionated on purpose, and the README says so.
The agent built the wrong thing
The agent built the wrong thing, so the first skill interviews you before it builds
❓ Q1 - Where do sessions live? In the browser only, or synced to the server so a second device can pick one up? ➡️ Browser only for now. Sync is a separate slice, and it would block everything else. --- ❓ Q2 - What happens on a duplicate name? Reject, auto-suffix, or overwrite? ➡️ Reject, with the existing item linked. Silent overwrite is the bug you can't undo.
You know this moment. You describe a feature, the agent works for twenty minutes, and what it built is not what you meant. The README calls this the most common failure in software, and quotes The Pragmatic Programmer: no one knows exactly what they want. His fix is a grilling session, where the agent interviews you before it writes a line. A round looks like this. Every question the agent can ask right now, numbered, each with a recommended answer, and it waits for yours before the next round. Finding facts is the agent's job. The decisions are yours.
1 · The set, and how it installs
The set, and how it installs
First, what is in the set and how it installs.
Subscribe or fork
Two install routes, and they are two philosophies: subscribe, or fork
A read-only bundle from Claude Code's official marketplace. It updates when he ships.
Editable copies written into your repo, for Codex and every other agent. Pull his changes when you want them.
There are two ways in, and the README is careful to say they are two philosophies. The Claude Code plugin is in the official marketplace, so one command installs the whole set as a managed bundle you don't edit, and it updates when he ships. The skills.sh route copies the skill files into your project as ordinary files you own. Hack on them, and pull his changes when you want them. That is also the route for Codex and every other agent. Pick one. Installing both leaves you with every skill twice.
Who may invoke a skill
Every skill is either yours to type, or the agent's to reach for
---
name: to-spec
disable-model-invocation: true
---/grill-with-docs · /to-spec · /to-tickets · /implement · /triage · /wayfinder · /ask-matt
---
name: tdd
description: Test-driven development.
Use when the user wants to build
features or fix bugs test-first…
---/tdd · /diagnosing-bugs · /code-review · /prototype · /research · /wizard · /grilling
Twenty five skills ship, eighteen for engineering and seven for general productivity, and the one line that splits them is who may invoke them. A user-invoked skill sets one flag in its frontmatter, and after that only a human typing its name can fire it. Those are the orchestrators: /to-spec, /to-tickets, /implement. A model-invoked skill keeps a rich description, with the trigger phrases in it, so the agent reaches for it on its own when the task fits. Those hold the reusable discipline: /tdd, code review, diagnosing bugs. The rule underneath: a user-invoked skill may call model-invoked ones, and never another user-invoked one.
2 · The main flow, idea to ship
The main flow, from an idea to a shipped change
Now the main flow, the route most work travels, from an idea to a shipped change.
Four skills, one context window
Grill, spec, split into tickets, implement, and keep the first three in one context window
Keep grilling, spec and tickets in one unbroken window. Then /clear before every /implement.
The router skill, /ask-matt, draws the whole map, and most work travels one road. You grill the idea with /grill-with-docs. You turn the thread into a spec with /to-spec. You split the spec into tickets with /to-tickets. Then /implement builds each ticket, driving /tdd inside it and closing with a code review before the commit. The rule that holds it together is about context. The first three steps stay in one unbroken window, so the spec and the tickets build on the same thinking, and every implement starts fresh from its ticket, so the last one's context is disposable.
A shared language in CONTEXT.md
The interview leaves a glossary behind, and the glossary makes the agent concise
"There's a problem when a lesson inside a section of a course is made 'real' (i.e. given a spot in the file system)"
"There's a problem with the materialization cascade"
/grill-with-docs is the same interview you saw at the start, with one difference: it is stateful. As it learns a term it writes it into a glossary file, CONTEXT.md, and it records a hard-to-reverse decision as an ADR. The README shows why that pays off with a line from his own course repo. Twenty words become three, because the agent and the developer now share a name for the thing. Variables and files get named the same way, the codebase gets easier to navigate, and the agent spends fewer tokens thinking. He calls it possibly the single coolest technique in the repo.
Tickets with blocking edges
Each ticket cuts a full vertical slice, fits one context window, and names what blocks it
# 02: Save a session to the browser
What to build: a signed-in user closes the tab, reopens it, and the session is still there, with its name and its messages.
Blocked by: 01: Name a session
Status: ready-for-agent
- [ ] Reopening the tab restores the last session
- [ ] A renamed session keeps its new name- A slice runs through every layer: schema, API, UI, tests.
- Finished, it is demoable on its own.
- Sized to one fresh context window.
- A wide refactor is the exception: expand, migrate in batches, contract.
/to-tickets breaks the spec into tracer bullets. On a local tracker that is one Markdown file per ticket, like this one: what to build, from the user's point of view, the ticket that blocks it, and acceptance criteria. On GitHub or Linear the blocking edges become native links, so any ticket whose blockers are done can be grabbed. The rules for a slice: it cuts through every layer at once, schema to tests, it works on its own as a demo, and it fits one fresh context window. The one exception is a wide refactor, a rename that touches thousands of call sites. That gets sequenced as expand, migrate in batches, then contract, so the build stays green in between.
Red before green, at agreed seams
Tests live only at seams you agreed on first, and refactoring is not part of the loop
- Red before green: a failing test first, then only enough code to pass it.
- One slice at a time: one seam, one test, one minimal implementation.
- No test is written at an unconfirmed seam.
- Implementation-coupled: the test breaks on a refactor when behaviour didn't change.
- Tautological: the assertion recomputes the answer the way the code does.
- Horizontal slicing: all tests first, then all implementation.
/implement is four lines long. Use /tdd where possible at pre-agreed seams, run the type checker and single test files often and the full suite once, then run code review and commit. The weight is in the /tdd skill it calls. A seam is the public boundary you test at, and before any test is written the agent has to list the seams and get your yes. Red comes before green, one slice at a time, and refactoring is moved out of the loop entirely into the review. It also names the tests to refuse: one that breaks on a refactor when nothing changed, one whose assertion recomputes the answer the same way the code does, and writing every test up front, which tests the shape of things instead of behaviour.
3 · Bugs, fog, and the steps only you can take
Bugs, fog, and the steps only a human can take
Three on-ramps that generate work and then merge onto that road: a bug, a foggy effort, and a step only a human can take.
No theory until one command goes red
The agent may not theorise about a bug until it has one command that goes red on it
- Red-capable: it asserts your exact symptom, and goes green once fixed.
- Deterministic: same verdict every run. A flaky bug gets its reproduction rate raised first.
- Fast: seconds, not minutes.
- Agent-runnable: unattended, or a human driven by a script.
## Redact Redact every secret first: write <REDACTED> in its place. Build loops against env vars, so the credential stays in the environment.
Diagnosing bugs is model-invoked, so the agent reaches for it when you say something is broken, and its whole discipline is in phase one. Before any hypothesis, the agent has to name one command it has already run that goes red on this exact bug: a failing test, a curl, a replayed trace, a bisection harness. If it catches itself reading code to build a theory before that command exists, the skill tells it to stop. A flaky bug isn't reproduced cleanly, its reproduction rate is raised until it is debuggable. The latest release, one point two point three, added a redact section, because the skill has the agent show commands and captured output, and those carry auth headers.
Decisions, not deliverables
A huge foggy effort becomes a map of decision tickets, and the map produces decisions, not deliverables
## Destination <what reaching the end of this map looks like> ## Decisions so far - [Sessions live in the browser](link): sync is a later slice - [Duplicate names are rejected](link): with the existing item linked ## Not yet specified ## Out of scope
- The map is one issue labelled
wayfinder:map; each ticket is a child issue holding one question. - Blocking uses the tracker's native links, so the frontier is visible in its own UI.
- When the map clears, it hands off to
/to-spec. It doesn't build.
Wayfinder is for the idea too big to hold in one session, where the way from here to the destination isn't visible yet. He calls it the most cognitively demanding flow in the repo. It charts a shared map on your issue tracker, one issue with a destination and a running index of decisions, and every child ticket holds one question sized to a single session. Blocking uses the tracker's own links, so you can see what is takeable without opening the map. And the rule he repeats is plan, don't do: the pull to just start building is the signal you've reached the edge of the map, and the map hands off to /to-spec rather than building anything itself.
A bash wizard for the clicks only you can make
For the steps only a human can take, the agent writes a bash wizard instead of a numbered list in the chat
- Provisioning infrastructure, creating credentials, walking an unfamiliar third-party dashboard, a one-off migration.
- It opens each URL, says what to click, captures each value, and writes it into
.envand GitHub Actions secrets. - Work an agent can do, an agent should do. The wizard is for the clicks and approvals you would not hand to one.
stage "Create the Stripe webhook" open_url "https://dashboard.stripe.com/webhooks" say "Add endpoint → paste the URL above" capture STRIPE_WEBHOOK_SECRET --secret env_upsert .env STRIPE_WEBHOOK_SECRET gh_secret STRIPE_WEBHOOK_SECRET
Every project has steps a human has to do by hand: create the webhook in a dashboard, paste a secret, approve a cutover. The usual result is a numbered list dumped into the chat that you follow and hope you didn't miss one. The wizard skill generates an interactive bash script instead. It opens each URL, tells you what to click, captures the value with hidden entry, and writes it into your env file and your GitHub Actions secrets. The bundled template already handles progress, confirmation gates and idempotent writes, so the agent only authors the stages. It is model-invoked with an explicit non-trigger: if the agent could do the step itself, it should, and the wizard is only for the clicks you would not hand to it.
What he says it can't do yet
The repo names its own limits: no native Codex plugin, a survey that won't rescue a codebase, and a listing that hides half the set
- No native Codex plugin. Codex's manifest takes one skills path and drops symlinks on install, so the promoted subset can't be expressed. Codex users take the skills.sh route.
- /improve-codebase-architecture is a survey, not a rescue. On an old codebase it finds candidates; it won't untangle the mud for you.
- Claude's desktop and web surfaces drop user-invoked skills from the listing, which is why /wizard was made model-invoked.
The repo is unusually clear about what it can't do, and the limits are written down where he decided them. The Claude Code plugin ships and the Codex plugin is deferred, because Codex's plugin manifest takes a single skills path and drops symlinks when it copies the plugin, so there is no way to ship only the promoted set. The ADR records both escape hatches he tried. The architecture survey finds deepening candidates and hands them to you. He says plainly it won't untangle an old codebase. And Claude's desktop and web apps drop user-invoked skills from the listing, which is the reason the wizard skill was made model-invoked when it graduated.
Where to read the files
Run the setup skill once per repo, then the interview from the first slide is one command away
MIT · v1.2.3 · 25 promoted skills, every one a Markdown file
claude plugins install mattpocock-skillsOne caveat before the link. The engineering skills assume you ran the setup skill once in the repo, so they know which issue tracker to write to and which triage labels you use. Skip that and /to-spec will stop and tell you to run it. Do that, and the grilling round from the first slide, the one that asks before it builds, is one command away. Every skill is a Markdown file, so read the ones you'd use before you install them. That is what is inside Matt Pocock's skills repo, the twenty five agent skills that 270,000 people on GitHub starred.
























