- Introduction
- The Setup Is Code
- Standards Before Prompts
- Small Skills for Models That Keep Getting Better
- Spec-Driven Development, the Way I Already Worked
- I Stay in the Loop
- Context Is a Budget
- Room to Act, Hard Boundaries
- What Carries Over to Building AI Products
- Conclusion
Introduction
My first few months with Claude Code went like this. I opened a session, described what I wanted, and got code back that mostly worked. Then I read it and found default exports where I use named ones, any sprinkled through the types, comments restating what the line below already said, and a caret range in package.json. So I typed the correction, the agent fixed it, and the next day I typed the same correction again in a fresh session.
The model was not the problem. The agent simply had no way to know my standards, because I had never written them down anywhere it could read. Every session started from zero, and I spent part of it re-explaining how I write code.
Over time I stopped treating the agent as a chat window and started treating it as a tool I configure. In this post, I will share what I learned along the way and the setup I ended up with: what I write down, what I deliberately leave out, and how a feature moves from an idea to a commit.
The Setup Is Code
My Claude Code configuration lives in its own git repository: the global CLAUDE.md, the settings.json, my skills, and two small shell scripts for the status line and notifications. A make install symlinks everything into ~/.claude. When I change how the agent behaves, I change it in the repository, and the change is live in the next session. On a new machine, one command sets up everything.
This turned out to matter more than I expected. Because the setup is a repository, every change to how the agent behaves has a diff, a commit message, and a history. When the agent starts doing something odd, I can look at what I changed recently. When something starts working well, I can see why. Without that history, I would be tuning the setup by memory.
Standards Before Prompts
The base layer is the global CLAUDE.md, an instruction file that Claude Code reads at the start of every session in every project. It holds the rules I would otherwise repeat forever: how I name things, how I type things, how I handle dependencies, and what the agent may and may not do. A few lines from it:
- Avoid `any`
- Use named exports wherever possible; no default exports
- Minimize comments. Comment only non-obvious *why* (tradeoff, workaround, gotcha), never *what*.
- When installing new dependencies, pin exact versions (no `^`/`~` ranges)
- NEVER push, force-push, or open a PR unless explicitly asked.
Two principles shaped that file more than anything else.
Write the rule at the point of failure. A rule like “prefer clean code” does not change anything. A rule that names the exact thing I keep correcting changes the next file that gets written. Most lines in that file exist because I once corrected the same thing twice.
Rules about behavior matter as much as rules about syntax. The style rules make the output look like my code. The git rules, like never pushing on its own, are what allow me to leave the agent running without watching it.
Small Skills for Models That Keep Getting Better
The instruction file says how to write code. It does not say how to do a specific job. That is what skills are for: short procedures the agent loads when a task calls for them. I have skills for commit messages, for my REST API conventions, for setting up linting and git hooks, for pinning and updating dependencies, and for writing new skills.
All of them are small. Many are a single paragraph. Here is the complete commit message skill:
---
name: gcc
description: Write a commit message for staged git changes. Use whenever a commit is about to be made: when the user asks to commit, asks for a commit message, invokes /gcc, or you are about to commit your own work.
---
Read the staged changes (`git diff --cached`) and stop if there are none. Read `git log --oneline -20` and follow whatever format those commits use; fall back to conventional commits if there is no clear one. Add a scope, body, or footer only when the subject genuinely cannot carry the meaning alone. Commit with the message only if committing was asked for; otherwise present it and stop.
Keeping them this short is a deliberate choice, and it comes from watching the models change. Every few months the model knows more than it did before. It already knows how to write a good commit message, how to set up ESLint, and how to structure a test. A skill that spells all of that out was useful once. With a newer model it is mostly noise, and in the worst case it pins the model to an older, worse way of doing something it would do better on its own. Long skills age with the model they were written for.
So a skill holds only the difference between general competence and my situation: my conventions, my exact commands, and the traps I already fell into. The model brings the rest, and that part keeps improving without me touching a file.
One more decision matters for every skill: who may start it. Some skills should trigger on their own. The commit message skill runs whenever a commit is near, including when the agent commits its own work, so I do not have to think about commit format anymore. Other skills should only run when I ask for them, which Claude Code supports with disable-model-invocation: true in the skill’s frontmatter. Setting up tooling, for example, is a deliberate step at the start of a project, not something the agent should decide to do halfway through fixing a bug.
Spec-Driven Development, the Way I Already Worked
The individual skills are useful on their own, but the workflow they form together is what changed how I work the most.
Think about how a feature reaches a developer on a team. You get a ticket that says what to build and why. You think about how to approach it in general. Then you get technical: which modules, which data, which decisions. You break that down into steps, and then you implement them one at a time.
That routine is not new, and it is not specific to AI. It is how I worked long before I used an agent. So when I set up spec-driven development, I did not want a new process. I wanted the one I already had, with the agent doing most of the typing.
Why Not an Existing Framework
I tried the established options first. I used Matt Pocock’s skills for a while, and I tried GitHub’s Spec Kit and OpenSpec. All three are good work, and I learned something from each of them.
But they all felt heavyweight for how I work: more steps, more files, and more ceremony than a feature usually needs, plus more text the agent has to read before it writes a single line. Every extra artifact is another thing to keep in sync, and every extra instruction is another thing a better model might handle well without it.
So I kept the idea and dropped most of the machinery. My version consists of four small skills that follow my own routine.
From Ticket to Commit
Each step writes one document into the project, and each document is the input for the next step:
- Specify, the ticket. The feature description, or the conversation so far, becomes a spec: the problem, prioritized user stories with acceptance criteria, requirements, edge cases, and what is out of scope. It holds only the what and the why, no implementation details. Every user story has to be shippable on its own, so implementing only the first one still leaves something to demo.
- Plan, the general approach. The agent reads the spec, explores the codebase, and decides how the feature fits in. The plan has to describe this codebase, not a generic one. It prefers extending what exists over introducing something new, and every design decision names the alternative it rejected.
- Tasks, the breakdown. Spec and plan become a task list, grouped by user story, where every task is small enough to finish and verify in one sitting. Each story ends with a checkpoint.
- Implement, one step at a time. The agent works through the list top to bottom, test-first, and runs the tests as it goes. At every checkpoint it verifies that story on its own. At the end it reviews its own work before anything is committed.
Each step has a specific purpose. The spec moves knowledge out of a conversation, which is temporary, and into a file, which is permanent. The plan forces the agent to read the codebase before it decides anything. The task list cuts the work into pieces a fresh session can finish. And the implementation relies on types and tests as a feedback loop, so the agent can tell whether it is done without asking me.
I Stay in the Loop
Spec-driven development does not mean I hand over a ticket and come back to a finished pull request. The workflow is built around the points where my judgment matters.
I review between the steps. After the spec, I check the stories and the assumptions. After the plan, I check the design decisions. After the task list, I check the granularity. Correcting the direction at this stage is cheap, because it is still a paragraph of text and not yet a pile of code.
The task list is the boundary. Work the agent discovers along the way gets added as a new task and shown to me. It is never done silently, and it is never skipped silently.
Tests are agreed on before they are written. The plan names the seams where the feature is tested: the public boundaries where behavior can be observed without reaching inside. The agent does not invent its own. Expected values come from the spec, not from the code under test. That rules out the kind of tests agents tend to write, the ones that are coupled to implementation details or pass by construction because they compute the expected result the same way the code does.
I own the result. The agent writes most of the code, but the architecture, the trade-offs, and the quality bar stay with me. Everything that gets committed goes out under my name, so I read it as if I had written it myself.
Context Is a Budget
The one piece of tooling I would recommend to anyone is a custom status line that shows how much of the context window is used, as a bar that goes from green to yellow to red.
It is a small script, but watching that bar all day changed how I work. Context is a limited resource, like memory or a rate limit, and the quality of the output drops well before you actually run out. When the bar turns yellow, I stop starting new things and focus on finishing the current one. When it turns red, I finish, commit, and open a fresh session.
That is also why the spec, the plan, and the tasks live in files. A fresh session reads them and continues exactly where the last one stopped, without depending on a conversation that no longer fits into the context.
Room to Act, Hard Boundaries
Claude Code runs in auto mode for me, so it can work without asking me to approve every single step. That is only reasonable because of the git rules in my CLAUDE.md. Nothing leaves my machine without me, so the worst case is a diff I throw away. If you loosen the permissions, tighten the boundaries first.
Since the agent can work on its own for a while, I do not sit and watch it. A notification hook sends a desktop notification when it is done or when it actually needs my input, and nothing else. A notification that fires all the time is one I would quickly learn to ignore.
What Carries Over to Building AI Products
Looking back, most of what I learned while building this setup is not specific to coding agents. It applies just as much to building products on top of language models:
- Writing a skill is writing a tool definition. It has a name, a description that decides whether the model picks it correctly, and a body that gives just enough context. I have debugged skills that never triggered because the description was too vague. That is the same bug as a tool the model never calls.
- Keeping skills small is designing for a moving model. Anything you build around a model’s current weaknesses becomes a liability once the next model no longer has them.
- Cutting work into tasks a fresh session can finish is context engineering. You decide what the model needs to see, what it does not, and where a job has to be split because it will not fit.
- Specify, plan, tasks, implement is orchestration. Decompose, sequence by dependency, execute in isolated contexts, and verify at checkpoints.
- Permissions plus boundaries is guardrail design. Give the component room to act, then keep the blast radius small enough that a mistake is cheap.
The interesting work has moved for me. It is less about generating code quickly and more about designing the system the model runs inside: what it knows, what it may do, how the work is cut, and how the output gets verified.
Conclusion
If you take one thing from this post, let it be the first step: put your agent configuration into a git repository, even if it starts as a single, almost empty instruction file.
Everything else can grow from there, and in my experience it grows from friction. If you correct the same thing twice, the rule belongs in the instruction file. If you explain the same procedure twice, it belongs in a skill. If you notice the agent doing something you did not sanction, a boundary belongs in the rules.
Keep the skills small and let the model handle the rest. And before you adopt someone else’s process, write down the one you already follow. It may be all the framework you need.