
I build software in my free time. Lately that means my homelab: when there’s an app I’d normally pay for, I build my own version and run it at home. So far that’s a document archive that reads and files whatever I scan, a household finance app, a lockbox for the numbers a family shouldn’t keep in a spreadsheet, and a pile of smaller tools around them.
I don’t write most of the code. One Claude Code session running in my terminal does. It acts like a project lead, handing pieces of the work to other AI agents and putting the results back together. My job is to make decisions and approve things.
I’ve used the same setup for a product launch, for client work and for the homelab, and the rules turned out to be the same every time. Here’s how it’s set up, and the rules that keep it from going off the rails.
The shape of it
There are three layers:
- Me. I set the goal, answer questions, and approve anything that touches real money, real data or other people.
- One main session. This is the Claude Code instance I’m talking to. It reads the project’s rules, checks my task list, does small things itself, and hands big things off.
- Subagents. When there’s real work to do, the main session spins up separate agents and runs them in parallel. Each gets a written brief, its own copy of the repo (a git worktree) and its own branch. When it’s done, it reports back.
That’s one project. Once you have several, each project gets its own main session, and the sessions can message each other. That turned out to matter more than I expected. This week the homelab session settled on a new look for my dashboard. I told the session for a different project to use it. It asked the homelab session where the design lived, got back a written list of the colors, the layout and the choices it had decided against, and restyled every page of that project to match. The same day another session needed a list of to-dos filed for me. It asked the session that already had access to my task list, and that one filed them and reported back what it had created.
I didn’t relay any of that. I said what I wanted, and the sessions worked out who knew what.
Rule 1: batch the work, release once
Left alone, every agent opens its own pull request, and every push kicks off a full build and test run. If you pay for build minutes, that adds up. Even if you don’t, five half-finished changes landing separately is five chances to break something.
So the rule is:
- Agents push branches but don’t open pull requests.
- The main session merges them locally, runs the full test suite on my machine, and pushes once.
- One pull request, one build, then one deploy.
One thing to add: have it clean up after itself. Every agent’s private copy of the repo stays on disk until something removes it. When I finally looked at one project, there were 34 of them taking up 6.7 GB, nearly all for work that had already merged.
Rule 2: the real thing is where the truth comes out
Passing tests tell you the code does what the tests expect. They don’t tell you it works. The bugs that matter show up when someone, or some agent, uses the real thing the way a person would.
My document archive is the best example I have. None of these came from a test. They came from using it:
- A photographed statement came back a quarter unreadable. Three pages, one of them sideways. Reading each page more than once, and turning the sideways one, got that down to 6%.
- The text reader was fighting itself. It sized itself for every processor in the machine while the app was only allowed one, so a page took 9.3 seconds when it should have taken 3.2. Run two scans at once and a page took so long that it timed out, and the document got filed with no text at all.
- Uploads from a phone died at exactly 60 seconds. On a slow connection the upload simply took longer than a default timeout in front of the app. From the outside it looked like a flaky app. It was one setting.
The dashboard had its own version of this. Its health widget showed a green dot for an app that had never been deployed, because a “page not found” answer counted as healthy. Nothing flagged it, because nothing was failing.
So before anything is called done, an agent has to exercise it for real: click through the deployed site, upload the actual file, run the flow end to end with test data. Then it says go or no-go.
Rule 3: the AI never makes the outward-facing decisions
This one is non-negotiable, and it’s written into a rules file the AI reads at the start of every session:
- It never sends email. It writes drafts into my Gmail, and I send them. (That rule exists because of a near miss where someone almost got the same email twice.)
- It never deletes a backup. Only I do.
- It backs up before it changes real data, and it logs the change in a work journal: what changed, why, the backup it took, and the numbers before and after.
- Anything live needs my OK. It does a dry run, shows me exactly what will happen, and waits.
- Some things it never sees. The lockbox holds passport and account numbers. Its rule is one line: none of it is ever given to an AI model.
The same thinking applies to automation that runs when I’m not watching. I have a small script on my router that restarts the VPN connection when it goes slow. It has to see fifteen minutes of slowness before it acts, it checks that my internet isn’t the real problem, and after three restarts in a day it stops and tells me, because at that point restarting isn’t the fix. Something that can act by itself needs a point where it gives up and asks.
Rule 4: make it check its own claims
AI is confident. That’s not the same as right, and the fix is to check claims against evidence instead of asking how sure it is.
The document archive taught me this one too. It uses a model to read each document and pull out the date, the sender and the amount. On a handwritten childcare invoice it reported 90% confidence, and every value was invented. It read “$20” as “$ZO” and gave a service period starting on February 29 in a year that doesn’t have one. Its confidence was only ever about which category the document belonged in.
So the archive doesn’t go by the model’s confidence anymore. A document gets flagged for me to review if a value it pulled out doesn’t actually appear in the text, if a date can’t be true, or if the scan itself was hard to read. Plain checks, with no model involved.
The same goes for what it writes about its own work. My homelab notes said an old service had been removed from every build pipeline. When an agent checked, 12 pipelines still loaded it. The dashboard listed which machine ran which app, and three of four were wrong. Now the rule is to check the running system, not the notes, and to fix the notes when they disagree.
Rule 5: write decisions down where people will find them
Small decisions pile up fast. If they only live in a chat, the next session starts from zero and makes a different call.
- Every decision becomes a task or a note in Todoist, which is my source of truth.
- Rules that should hold next time go into the project’s rules file, with the reason. “Never delete a backup” is a rule. The story of what happened the one time it did is what makes the next session take it seriously.
- The homelab keeps a map of what exists and a short list of how new things get built there, so a new app starts from the same conventions as the last one.
- At the end of a long day, I have the main session go through the task list and close what shipped, with notes on where each thing landed.
What I’d tell you if you try this
- Keep one session in charge of each project. One orchestrator with a clear picture beats five independent chats stepping on each other.
- Give agents narrow briefs and their own branches. Parallel work is great until two agents edit the same file. Separate copies plus one integration step fixes that.
- Release in batches. Merge locally, test once, push once.
- Trust the real thing, not green tests. A real scan from a real phone found what the tests didn’t.
- Check claims against evidence. That includes the AI’s confidence and its own notes.
- Keep the outward-facing actions for yourself. Email, live changes, production deploys and deleting anything stay with a human.
It isn’t hands-off. I make a lot of calls. But I spend my time deciding instead of typing, and that’s the trade I want.
If you have questions about the setup, feel free to reach out!














