Skip to content
All work

Personal project · Claude Code tooling

AI Toolkit

Sets up Claude Code on any repo: it maps the project, pins what works today as a baseline, and generates the agents, hooks and skills to build new features without breaking old ones.

  • Claude Code
  • Skills
  • Hooks
  • Agents
  • Developer experience
Not public yet
Role
Author and architect
Runtime
Claude Code
Status
Design v0.7, first installable version
Platforms
macOS, Linux and Windows installers
verification levels with time budgets
4
budget per file for inline checks
2 s
files written before you confirm the plan
0

The problem

AI agents write code fast. On an existing repo they can also break behavior that was already tested, and full verification on every change makes the loop slow again.

I wanted a toolkit that takes any project, pins what works today, and extends it with new product requirements, without making the edit loop wait.

How it works

  • The toolkit itself: a versioned set of skills, agent templates, hook templates and adapters, started with /ai-toolkit-init.
  • What it generates per project: CLAUDE.md, context files, agents and hooks calibrated to the real commands of the repo.
  • A leveled verification loop: cheap checks inline, expensive ones in the background or in CI.

A manifest records every file the toolkit wrote and how. Rerunning it after an upgrade only touches what changed, and hand edits are never overwritten silently. The first run is a dry-run plan.

Once per project

  1. Discoverymap, graph, inherited context
  2. Baselinewhat works today, pinned
  3. GenerationCLAUDE.md, agents, hooks

Per feature

  1. Speccriteria with IDs
  2. Scaffoldyour convention
  3. Implementtests and code, separatedlevels 0 and 1, ms to seconds
  4. Gatestraceability blocks
  5. ReleasePR to main, at minimum

From Implement to Gates. Fast path: fixes an existing criterion, skips spec and opinion, keeps traceability

Once per project: discovery, baseline and generation. Then every feature goes through spec, scaffold, implement, gates and release. Small fixes take the fast path.

Key decisions

Roles by permission, not by prompt

The agent that writes tests and the one that writes code are separated by hooks that deny writes outside their paths. If the same agent writes both, the test checks what the code does, not what it should do.

CLI plus skill over MCP

An MCP server costs context on every turn. A command line tool called from a skill costs nothing until it is used.

Inform first, block later

Every control starts by reporting. Blocking is earned with metrics, so the toolkit never slows people down on a guess.

Verification that never blocks the edit loop

  • Level 0, after each edit: lint and typecheck on changed files, 2 seconds per file.
  • Level 1, when the agent stops: tests related to the touched files, 30 seconds.
  • Level 2, in the background or CI: full suite, end to end, performance and accessibility, as a report against the baseline.
  • Level 3, on the pull request: full CI and reviewer agents.
  • Runs while you work
  • Runs in the background
  1. Level 0Lint and typessynchronous, 2 sRuns while you work. Relative cost: 1 of 5.
  2. Level 1Related testson task close, 30 sRuns while you work. Relative cost: 2 of 5.
  3. Level 2Full suite, e2e, performanceasync, minutesRuns in the background. Relative cost: 4 of 5.
  4. Level 3CI and reviewerson the PRRuns in the background. Relative cost: 5 of 5.
The more a check costs, the later it runs. Only four things can block a pull request; everything else informs.

Whatever does not fit a level budget moves to the next one. Deciding whether a failing test is a regression or an intended change is always a person’s call.

Skills that ship with it

  • watch-youtube: downloads a video, samples frames and reads captions, so Claude can study a product walkthrough. It powered the Ramps study.
  • handoff-set and handoff-take: pass a long session to a fresh one without dragging the whole context.
  • create-feature: scaffolds a feature folder following the project conventions.

How it was designed

I planned and brainstormed the whole design with Claude in Cowork, debating best practices: how to keep a steady working rhythm, keep token use low, separate concerns, and move expensive operations out of the hooks. The model argued the options; the calls were mine.

What comes next

The MVP is a strict vertical slice on a single JavaScript or TypeScript repo, to test one hypothesis: that permission based isolation and a traceability reviewer give useful signal without slowing work down. Karma Plus was the first real repo to run it on.

I ran it on Karma Plus and it was very efficient. It helped me close plenty of gaps and add end to end tests along the way, and it made me far more productive.

Next case study

Freedom Robotics

A 0 to 1 fleet management app for robots. I redesigned how it used the API, cutting query load by 90% and perceived delays by 75%, and built the internal SDK the app used to talk to it.