Skip to content

About

Skills for Claude Code, Codex, Cursor and Antigravity that make agents plan before building and pass checks before calling work done.

Topics

Resources

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

265 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

agent-skills

Version · Changelog · MIT licence

19 Agent Skills for product and software delivery, for Claude Code, Codex, Cursor and Antigravity. They are built around three ideas.

  • Evidence-gated shipping. Deterministic checkers write a standard report, and the acceptance gate re-runs them itself every time.
  • The builder isn't the acceptor. The acceptance gate caps its verdict at CONDITIONAL unless the run declares it is independent of the build. Nothing in the code can verify that, so it depends on whoever runs it being honest.
  • One registry. registry.json lists every skill and artifact, docs/CONTRACT.md is generated from it, and CI fails if the two drift apart.

Evidence

I use these skills every day and find them useful, but that's anecdotal and they aren't formally validated yet. Turning it into a result needs a comparison against running without them, and that work is in eval/. It holds 808 recorded runs and 49 write-ups. Each run is tied to the exact case, fixture, grader and skill text it used. Changing any of those retires the runs that depended on it. Nothing there clears the bar I set for this repository yet.

Here is what the evidence supports so far.

  • The checkers work. Every checker has ship and block fixtures that assert the specific blocker, and they run in CI on Ubuntu and Windows.
  • Skills get picked up in interactive sessions. Over two months of my own Claude Code work, the model chose a skill itself 97 times out of 127. In 33 of those, the prompt didn't mention the activity. In non-interactive claude -p runs, a probe found none were picked up unprompted. Codex reads them in single-prompt runs, but it lists every installed skill in its system prompt, so that isn't comparable. The details are in field-outcomes-2026-10-03.md.
  • Whether the guidance improves the work isn't known yet. That needs runs with and without each skill compared, which is what eval/ is for. Results in eval/results/ from before the current method aren't evidence either way.

Skills

The skills are independent. Each one triggers on its own and none calls another. Twelve of them read or write a few shared files, listed in the next section. The one other link is that delivery-workflow checks pull requests with the voice rules from repo-docs when both are installed.

Skill Role
product-build Works out which of the other skills a new or vague request needs
product-management Interviews for the PRODUCT.md contract
systems-architecture Parts, boundaries and trust
frontend Stack, structure, design and UX
backend-engineering Rules for the trusted side
product-acceptance Independent acceptance gate
ai-prose-slop Prose editor and detector, usable on any writing
repo-docs Drafts and checks README, release notes, CHANGELOG entries and ADRs, with a checker for voice and structure
delivery-workflow Gets work merged through branches and pull requests, versions releases from Conventional Commits and keeps them drafts, with a guard hook and branch protection
mental-models Reasoning lenses, a triage guide and four mindsets (Skeptic, Systems Thinker, Pragmatist, Explorer), usable on any hard problem
code-smells Fowler's code-smell catalogue and a file-size and nesting checker (size for any language, nesting for JS, TS and C-family)
code-organization Module boundaries, dependency direction and naming
testing-strategy What to test at which level, and testing behaviour over implementation
data-modeling Schema design in any format and a raw-SQL migration-safety checker
cli-tooling CLI naming, config precedence, exit codes and dry-run
release-engineering CI/CD gating, deployment strategy and rollback
learn-from-session Turns a correction or confirmation into a lasting rule, fixture or memory
engineering-assessment Whole-codebase audit with severity-ranked findings, each citing a file, line or command output, and a list of what wasn't examined
multi-agent-design Whether multi-agent is justified at all (the default answer is no), then topology, delegation and failure recovery

How they compose

Each skill applies when its signal is present.

Signal Skill Reads / writes
No or thin PRODUCT.md product-management writes PRODUCT.md
Multi-part system (client and server, workspaces, trust boundaries) systems-architecture writes ARCHITECTURE.md
Stack or structure unknown, or design and UX direction unset frontend writes design-direction.md, ux-walkthrough.md and tokens
Server or API in scope backend-engineering reads ARCHITECTURE.md
A readiness claim ("ship it", "is this done") product-acceptance reads whatever artifacts exist and re-runs every applicable checker
Any prose ai-prose-slop nothing

A new build usually touches most rows in about this order, and product-build suggests it, but nothing enforces it. A request that matches one row, like "make this accessible", uses only that skill. The full artifact contract, with exact files, required headings and gating rules, is generated into docs/CONTRACT.md from registry.json.

Install

node scripts/install.mjs --harness claude     # or cursor | codex | all

The installer never overwrites a directory it didn't create unless you pass --force. It needs an explicit target and makes no network calls.

The repository also has marketplace metadata for Claude Code, Codex and Cursor, and native packages for the Gemini and Antigravity CLIs. Listing in a public directory is a separate step. INSTALL.md covers every install path and what each harness can and can't do. It also explains why an installed skill doesn't always get invoked, and the CLAUDE.md line that reliably fixed that.

Verify a project

node systems-architecture/scripts/check-architecture.js --root . --strict
node frontend/scripts/check-frontend.js --root . --strict
node backend-engineering/scripts/check-backend.js --root . --strict
node release-engineering/scripts/check-smoke.js --root . --strict
node product-acceptance/scripts/accept-check.js --root . --strict

The acceptance command caps its verdict at CONDITIONAL. Add --acceptor-context separate only when the three conditions in product-acceptance/SKILL.md hold, the first being that this conversation didn't write or edit the code. If you're unsure, leave the cap on.

In a terminal the checkers print a readable verdict with the failing checks first. Piped or spawned, they print JSON for the acceptance gate and the pre-commit hook. --format text|json overrides either way.

BLOCK  systems-architecture  (/path/to/project)
  FAIL  P-arch-doc: multi-part project has no architecture doc (looked for: ARCHITECTURE.md, ...)
  --    P-section-parts: no architecture doc to inspect

Fix the FAIL line(s) above and re-run. Nothing ships on a BLOCK.

Reports are written to .agent-evidence/, which you should gitignore. Any failed check means BLOCK, and any check that couldn't be evaluated caps the verdict at CONDITIONAL. The full contract is in docs/CONTRACT.md.

Tests

node scripts/run-tests.mjs

CI runs this on Ubuntu and Windows (.github/workflows/ci.yml).

Security

Skills treat project documents like PRODUCT.md and ARCHITECTURE.md as data and never run commands found in them. Secret scans report file paths, never values. See SECURITY.md.

About

Skills for Claude Code, Codex, Cursor and Antigravity that make agents plan before building and pass checks before calling work done.

Topics

Resources

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages