Category: Technology

  • One person, one terminal, a team of AI agents: how I build software now

    One person, one terminal, a team of AI agents: how I build software now

    Diagram: me, one Claude Code session, the subagents it spins up, the systems it can reach, and the batched release pipeline
    How the work flows: me, one Claude Code session, and the agents it spins up

    I build software in my free time. Lately that means my homelab: when there’s an app I’d normally pay for, I build my own version and run it at home. So far that’s a document archive that reads and files whatever I scan, a household finance app, a lockbox for the numbers a family shouldn’t keep in a spreadsheet, and a pile of smaller tools around them.

    I don’t write most of the code. One Claude Code session running in my terminal does. It acts like a project lead, handing pieces of the work to other AI agents and putting the results back together. My job is to make decisions and approve things.

    I’ve used the same setup for a product launch, for client work and for the homelab, and the rules turned out to be the same every time. Here’s how it’s set up, and the rules that keep it from going off the rails.

    The shape of it

    There are three layers:

    • Me. I set the goal, answer questions, and approve anything that touches real money, real data or other people.
    • One main session. This is the Claude Code instance I’m talking to. It reads the project’s rules, checks my task list, does small things itself, and hands big things off.
    • Subagents. When there’s real work to do, the main session spins up separate agents and runs them in parallel. Each gets a written brief, its own copy of the repo (a git worktree) and its own branch. When it’s done, it reports back.

    That’s one project. Once you have several, each project gets its own main session, and the sessions can message each other. That turned out to matter more than I expected. This week the homelab session settled on a new look for my dashboard. I told the session for a different project to use it. It asked the homelab session where the design lived, got back a written list of the colors, the layout and the choices it had decided against, and restyled every page of that project to match. The same day another session needed a list of to-dos filed for me. It asked the session that already had access to my task list, and that one filed them and reported back what it had created.

    I didn’t relay any of that. I said what I wanted, and the sessions worked out who knew what.

    Rule 1: batch the work, release once

    Left alone, every agent opens its own pull request, and every push kicks off a full build and test run. If you pay for build minutes, that adds up. Even if you don’t, five half-finished changes landing separately is five chances to break something.

    So the rule is:

    • Agents push branches but don’t open pull requests.
    • The main session merges them locally, runs the full test suite on my machine, and pushes once.
    • One pull request, one build, then one deploy.

    One thing to add: have it clean up after itself. Every agent’s private copy of the repo stays on disk until something removes it. When I finally looked at one project, there were 34 of them taking up 6.7 GB, nearly all for work that had already merged.

    Rule 2: the real thing is where the truth comes out

    Passing tests tell you the code does what the tests expect. They don’t tell you it works. The bugs that matter show up when someone, or some agent, uses the real thing the way a person would.

    My document archive is the best example I have. None of these came from a test. They came from using it:

    • A photographed statement came back a quarter unreadable. Three pages, one of them sideways. Reading each page more than once, and turning the sideways one, got that down to 6%.
    • The text reader was fighting itself. It sized itself for every processor in the machine while the app was only allowed one, so a page took 9.3 seconds when it should have taken 3.2. Run two scans at once and a page took so long that it timed out, and the document got filed with no text at all.
    • Uploads from a phone died at exactly 60 seconds. On a slow connection the upload simply took longer than a default timeout in front of the app. From the outside it looked like a flaky app. It was one setting.

    The dashboard had its own version of this. Its health widget showed a green dot for an app that had never been deployed, because a “page not found” answer counted as healthy. Nothing flagged it, because nothing was failing.

    So before anything is called done, an agent has to exercise it for real: click through the deployed site, upload the actual file, run the flow end to end with test data. Then it says go or no-go.

    Rule 3: the AI never makes the outward-facing decisions

    This one is non-negotiable, and it’s written into a rules file the AI reads at the start of every session:

    • It never sends email. It writes drafts into my Gmail, and I send them. (That rule exists because of a near miss where someone almost got the same email twice.)
    • It never deletes a backup. Only I do.
    • It backs up before it changes real data, and it logs the change in a work journal: what changed, why, the backup it took, and the numbers before and after.
    • Anything live needs my OK. It does a dry run, shows me exactly what will happen, and waits.
    • Some things it never sees. The lockbox holds passport and account numbers. Its rule is one line: none of it is ever given to an AI model.

    The same thinking applies to automation that runs when I’m not watching. I have a small script on my router that restarts the VPN connection when it goes slow. It has to see fifteen minutes of slowness before it acts, it checks that my internet isn’t the real problem, and after three restarts in a day it stops and tells me, because at that point restarting isn’t the fix. Something that can act by itself needs a point where it gives up and asks.

    Rule 4: make it check its own claims

    AI is confident. That’s not the same as right, and the fix is to check claims against evidence instead of asking how sure it is.

    The document archive taught me this one too. It uses a model to read each document and pull out the date, the sender and the amount. On a handwritten childcare invoice it reported 90% confidence, and every value was invented. It read “$20” as “$ZO” and gave a service period starting on February 29 in a year that doesn’t have one. Its confidence was only ever about which category the document belonged in.

    So the archive doesn’t go by the model’s confidence anymore. A document gets flagged for me to review if a value it pulled out doesn’t actually appear in the text, if a date can’t be true, or if the scan itself was hard to read. Plain checks, with no model involved.

    The same goes for what it writes about its own work. My homelab notes said an old service had been removed from every build pipeline. When an agent checked, 12 pipelines still loaded it. The dashboard listed which machine ran which app, and three of four were wrong. Now the rule is to check the running system, not the notes, and to fix the notes when they disagree.

    Rule 5: write decisions down where people will find them

    Small decisions pile up fast. If they only live in a chat, the next session starts from zero and makes a different call.

    • Every decision becomes a task or a note in Todoist, which is my source of truth.
    • Rules that should hold next time go into the project’s rules file, with the reason. “Never delete a backup” is a rule. The story of what happened the one time it did is what makes the next session take it seriously.
    • The homelab keeps a map of what exists and a short list of how new things get built there, so a new app starts from the same conventions as the last one.
    • At the end of a long day, I have the main session go through the task list and close what shipped, with notes on where each thing landed.

    What I’d tell you if you try this

    • Keep one session in charge of each project. One orchestrator with a clear picture beats five independent chats stepping on each other.
    • Give agents narrow briefs and their own branches. Parallel work is great until two agents edit the same file. Separate copies plus one integration step fixes that.
    • Release in batches. Merge locally, test once, push once.
    • Trust the real thing, not green tests. A real scan from a real phone found what the tests didn’t.
    • Check claims against evidence. That includes the AI’s confidence and its own notes.
    • Keep the outward-facing actions for yourself. Email, live changes, production deploys and deleting anything stay with a human.

    It isn’t hands-off. I make a lot of calls. But I spend my time deciding instead of typing, and that’s the trade I want.

    If you have questions about the setup, feel free to reach out!

    Don’t miss an update

  • We’re BACK – Fantasy Football 2026-27

    Back by literally no demand will be my Fantasy Football series! New team, new draft agent, new tech to support the team. If you are new here, last year I did a whole series about using AI to manage my Fantasy Football team. I was even . interviewed about it.

    During the NFL offseason I retooled the drafting agent. Some highlights:

    1. New data backbone using https://gridirondata.com that has historical season context, player news, and season actuals (once the season starts)
    2. Newer AI models – Last year I used Anthropic’s Haiku 3.5 model. This year, we have Haiku 4.5. The model reasoning is now more roster aware in the sense that it wants to add starters first and then focus on bench seats.
    3. A real GUI – last years draft was done all via CLI. This year I have a full web interface for drafting, seeing the AI recommendations as well as “best available” players.

    I’m still retooling the “Coach” but don’t worry, I will still force the AI to adopt a Dan Campbell persona.

    So, I present to you for the first time… The Hallucination Hail Mary’s!

    If you think we are going to win the league this year comment below. Otherwise just let me know what you think!

  • An agentic Kanban Workflow

    An agentic Kanban Workflow

    I had this idea the other day and started building it into a project I’m working on. We all hate Jira but, the idea of Kanban is still a useful way to track projects.

    As we think about this in the AI era we could easily integrate the Jira MCP into a workflow but, once again we hate Jira. So that led me down a path of reinventing the wheel, mostly for my own purposes. I came up with this simple diagram:

    We will forever want to keep a human in the loop so a web interface is still likely necessary. However, a text based interface could also be cool…. Maybe in the future!

    What we end up with is a system of three agents:

    1. Developer agent – this agent writes and builds code in a sandboxed container based on specs written either by a human or by bugs found by the QA agent
    2. Build Agent – This agent monitors our build pipelines and if there is a failure it diagnosis why and opens a bug accordingly for the developer agent to fix
    3. QA Agent – Arguably the most important agent. This one will execute testing as close to simulating a human interaction with the software as possible. Upon finding bugs it would be able to log them back into the Kanban for the developer agent.

    Now we have a full DevOps life cycle with three agents. If I build this out, I become the user who is simply entering specs as features or bugs for the developer agent to work through. The code is still stored inside of some git based repository, build failures can utilize my already coded Build Failure Agent. Claude Code or Codex could function as the developer agent or we could run the whole thing on AWS Bedrock.

    Proof of concept coming some day when I have time!

    Don’t miss an update

  • What I’ve Built in 2026

    What I’ve Built in 2026

    Q1 is done. Feels like time is flying by this year already. We’ve already hit that gross humid stage in Chicago. I’m sure all the AI data centers aren’t contributing to climate change.

    Anyway, I’ve been hard at work both at the day job as well as hammering out personal projects. I’ve been spending a lot more time setting up systems for 45Squared, my web development company that I started back in 2016 crazy to think it will be turning 10 years old this year!

    I built a custom CRM for monitoring and managing incoming leads. The frontend lives in my home lab but reaches out to a DynamoDB table where I have leads being scraped using Google Places through an automated Lambda. Some other features:

    • Email sending and templating
    • Email tracking and unsubscribe
    • Incorporated SEOScoreAPI so I can get quick scores
    • Gemma4 Integration for suggesting which leads to contact
    • Custom scraper for enriching leads with phone numbers, addresses publicly available emails
    • Apollo.io for email validation
    • Hunter.io for email searching
    • AWS Cost agent integrated
    • Stripe and Quickbooks Revenue being tracked

    So far i’ve gotten 3 unsubscribes and sent over 100 emails! Success!

    SEOScoreAPI is continuing to slowly grow with over 1000 audits. I built in ADA compliance checking after reading about so many frivolous lawsuits on Reddit. I also added GEO scoring so that you can gauge how well your site is readable by LLM’s.

    My Fantasy Football pipeline is working and I think it’s ready for the next season! I’ll still probably lose but at least this year I will be able to blame technology again.

    The build remediation pipeline is working swimmingly. I’ve been integrating it into every pipeline I setup so that I can continue to build data for training a true DevOps AI agent on at some point.

    I’m still active on UpWork and taking on new work there as it comes to me. Bidding on jobs with UpWork is an absolute nightmare. My last engagement was handling background noise with AWS Lex.

    My AWS Cost Agent now supports multiple accounts as well as tracking AI/Bedrock usage costs. I will be uploading a video about how this works soon.

    Seems like there is more and I’m sure there is. If you want to hear more about any individual project feel free to reach out to me at anytime!

    Don’t miss an update

  • Evolution of my build failure agent

    Evolution of my build failure agent

    I’ve written in the past about troubleshooting build pipelines with AI. While all of this is a great step in speeding up your development and reducing the amount of troubleshooting the DevOps team needs to do in the enterprise, it is NOT the end goal.

    The end goal would be to have the AI fix the problem for you.

    I’m rebranding my Jenkins Sentinel to just be Sentinel. This workflow allows you to automate remediation for your pipelines while still retaining human in the loop security.

    The other primary feature is storing your build failures and remediations in a database that you can view, update, analyze for custom model training.

    Originally we had the dispatch layer that would notify us of build failures and possible resolutions. The new addition is the cluster of “workers”. Running on AWS Fargate, this team of developers works with the LLM on Bedrock to resolve the failure.

    1. The task spins up in the cluster
    2. The build logs identify the repository and branch
    3. The repository is cloned, and branch checked out
    4. The code fix is implemented
    5. The task generates its reasoning and updates the database accordingly
    6. Code is committed to a new branch and a pull request is opened.
    7. The task cleans up and shuts down

    Dispatch still remains the same and the developer is notified accordingly. I need to implement developer specific notifications so that channels are not flooded or email lists abused.

    The other major thing I wanted to see was the cost per fix.

    This screenshot is from the dashboard which shows the compute spend and the LLM spend. For this simple Terraform fix you can see the was a little around $0.02. Assuming your code bases are more complex this value could increase proportionally.

    I also included a stats page which shows the totals for the entire organization.

    This is all real data from my testing project. The build agent is successfully troubleshooting pipelines for:

    • Python
    • Terraform
    • Java
    • Typescript
    • Docker
    • Kubernetes
    • Go
    • Cloudformation

    I plan to continue to add more supported platforms and languages as time allows. The other major integration that I am working on is support for GitHub Actions. Once I complete that integration and put this into all of my pipelines I expect that my troubleshooting and development time will decrease rapidly.

    Other future plans include:

    • Ingestion of bugs through sources like Jira, ToDoist (my favorite), or another ticketing system.
    • Discord Dispatching
    • Teams Dispatching – although this is really hard to develop for without a paid account
    • Custom model – using the build failure data to train a model

    Anyway, this project has been super fun. If you want to implement it on your own infrastructure feel free to reach out!

    Don’t miss an update

    PS: the featured image was generated and setup through my Nano Banana WordPress plugin

  • Fantasy Football and AI – The next season

    Fantasy Football and AI – The next season

    If you aren’t familiar with my Fantasy Football series from last year, I built an AI assisted drafting agent based on data I collected from previous seasons. I then iterated over that setup to build an in season coach that determined my line up each week. AI and I did pretty well but ultimately we are still chasing the championship.

    So now what? How do I improve the system? Well, I started over.

    This off season I took all of the stats I could get from 2020-2025 seasons. I restructured my database and moved it all to my local home lab. The purpose of this was to be able to cost effectively train, re-train, throwaway, and train again multiple models using multiple strategies for each position until I could accurately predict the results of the 2025 season.

    The result is a CatBoost and LightGBM system of models that has the following features:

    • Rolling Stats – 3, 5, 8 week rolling average of fantasy points
    • Vegas lines – spread, over/under, team implied total, implied win probabilty
    • Next Gen Stats: snap counts, target share, CPOE, rush yards over expected
    • Play-by-play: EPA per play, weighted opportunity share, red zone target share, goal-line carry share
    • Defensive matchup: rolling 5-week EPA allowed by the opposing defense at each position
    • Injury status: game designation, practice participation, teammate injuries
    • Schedule context: home/away, dome/outdoor, rest days, primetime

    I tested XGBoost, LightGBM, CatBoost, NGBoost, LSTM, and Transformer architectures. CatBoost + LightGBM ensemble won — the gradient boosting models crushed the deep learning approaches on tabular data (MAE 2.95 vs 4.68 for the LSTM).

    Using this, I then build a brand new draft simulator that tests various selection strategies based on position in the draft, draft type, and number of teams.

    The infrastructure is setup for next season. I’ll be going into more detail on each aspect of this setup throughout the off season. If you are interested in using the model feel free to reach out!

  • Why I built my own WordPress Platform

    Why I built my own WordPress Platform

    I’ve been building websites for a long time. I remember learning HTML and Microsoft FrontPage as my first website builder. It was such a fun time to be creating horrific looking websites back in the early 2000s. As the internet progressed so did my skills and back in 2016 I formed my company 45Squared to build websites for small businesses. My whole goal is to be your trusted resource when it comes to being online.

    When I started the company I built WordPress websites of various shapes and sizes but they always ran on AWS. This helped me expand my AWS skills as well as provide robust infrastructure for my client’s websites to live on. I managed the website and the underlying infrastructure for a small monthly cost that beat the competition. The result, a bunch of paying customers a decent side hustle.

    As time went on, selling became harder and the race to zero for cost was apparent. So, as the AI boom is on, I decided that it was time to automate the site building process.

    I started documenting out how I would want this to work. Fully automated website deployments, design, content, custom domains, good SEO base and deployed FAST!

    Enter https://ai.45sq.net. This platform is fully automated. The customer can provide inputs and descriptions of what they want as well as photos or other graphical content. The workflow takes all of the inputs and builds a fully functional WordPress website hosted on AWS. The user can easily point their own domain to the server and setup automatic payments. They then get full administrative access to their website so they can expand and add features just like any other WordPress site.

    So why did I build this?

    If you contact a web designer now you will have to pay them to build up the initial design, work with their timelines, end up with something that needs revisions and your time to live will be in the weeks not minutes.

    The platform I built for 45Squared eliminates the need for the initial design fees and focuses on getting you online quickly. Its great for small businesses who are just getting started.

    So now when I get a request to build a site I can tell the customer that I have two options. First, fully custom. I’m still willing to sit with you and build out the picture perfect website. Or, two, you can launch your own and I will still support the website and help you with your online presence.

    So that’s it. An easy to use WordPress website launcher. Running on enterprise grade cloud. With content, design, layout and all the rest handled by the magic of Claude Opus.

    Try it out: https://ai.45sq.net. No contracts. No weird fees. Get online today.

  • I built a WordPress Plugin For Generating Images With Nano Banana

    An image of a blogger who takes himself way too seriously for his own good.

    AI is every where. Accept it. Anyway, I had a random thought last night about having a WordPress plugin that allows you to generate images on the fly for your posts. Pictures increase engagement on posts so, what if we just inline Nano Banana directly into Gutenberg?

    This morning I built this plugin which is a simple API call to Google’s Gemini AI Studio through a Gutenberg block.

    1. Type your prompt
    2. Choose your model
    3. Hit generate
    4. Insert

    Simple!

    Nano Banana Image Generator block

    Once the image is inserted into the post it turns the block into a standard image block so its as easy to manage as any other image.

    I submitted the plugin to the official WordPress repository but it takes a while to get approved. So, if you want to add it to your own WordPress instance feel free to message me and I’ll give you access to the repository!

    Don’t miss an update

  • Building in Public – The Automated WordPress Deployment Platform Part 2

    Building in Public – The Automated WordPress Deployment Platform Part 2

    A few days ago I wrote about building an automated WordPress deployment platform using Terraform and AI. Well, i’m happy to report that the entire platform is live and ready for you to explore and launch your own WordPress website.

    Introducing 45Squared’s WordPress deployment platform powered by Ubuntu and Claude. Try it out today at https://ai.45sq.net.

    Let’s talk about how this all works.

    The front end infrastructure that an end user will see is pretty straightforward. I am utilizing an ECS cluster and NextJS to deliver the end user experience. The second portion of the user experience is handled by an AWS API Gateway to manage all of the user credentials, payment processing, sit launch status. Authentication is handled by AWS Cognito. Hate on it all you want, Cognito works just fine when configured correctly.

    Frontend architecture

    Behind the scenes, once a user transaction has completed successfully, the website is provisioned using another ECS task. This container runs through a sequence of steps to provision the AWS EC2 instance for the user to utilize. Each tenant instance is running a hardened Ubuntu image that is built using Packer. I will cover this in another post. Throughout the provisioning process, the task is updating the DynamoDB table so that the user gets a live look into how their website is progressing.

    Provisioning architecture

    Each tenant is given a subdomain as well as the ability to utilize a custom domain name. Each tenant is also given a Cloudfront CDN for global static content distribution. And of course, each tenant receives their own SSL certificate for both their custom domain and their subdomain.

    Each site can be managed by SSM which will eventually be linked into an AI agent for management through Slack or another messaging platform.

    I don’t intend to use this platform to compete with the large players. 45Squared’s vision has always been to serve the small to medium size businesses who want personalized support while still receiving an amazing product. This platform gives them the ability to quickly launch a website and get their company on the world wide web within 10 minutes.

    If you are interested in building out a website using the platform the first few users can receive 50% using code “BETATESTER50”. There are limited redemption so be sure to get going quickly!

    Don’t miss an update

  • How I’m Using AI Agents to Run My Side Projects

    I think the promise of AI is to handle the work flows that maybe you don’t have time for. Or, maybe something that’s slightly out of your realm of expertise.

    I read an article the other day that had this quote:

    I keep waiting for someone to walk into my office and tell me what problems I should be solving with AI. Nobody’s come. (link)

    This mind set, in my opinion, is fundamentally incorrect. If you are looking for problems to solve there are a vast number of them and you’ll be drowned in possibility. For me, AI has always been about gaining efficiency or adding capability.

    If you haven’t noticed, I have a lot of side projects. I hope some day one of them hits the jackpot and allows me to retire to an island with all my family and friends. It’s unlikely, but hey, I’m allowed to dream.

    I realized long ago that I’m not great at marketing. I don’t understand lead generation, i’m not super amazing at SEO (so i built a tool for it https://seoscoreapi.com), I’m not in the least bit artistic (that went to my brother (https://mrbenny.co). But what I am good at is problem solving and process creation.

    I’ve been working on a concept of second brain for a while but I realized even a second brain needs to have tools. I started thinking about how to manage all of my side projects and how to interface with them through my preferred platform of Slack. (Sponsor me?).


    I’ve come up with a business development agent that can handle the things that I don’t particularly specialize in. On the backbone of Claude Sonnet 4.6, I created an API Gateway that takes input from my Slack instance and can handle a variety of tasks. 

    • SEO – Using my SeoScoreAPI it can handle generating SEO reports
    • Lead Generation – Using a variety of 3rd party API’s I have it looking for businesses that don’t have websites so I can pitch them some design services. Or, if their SEO is bad I can assist in fixing it
    • Lead nurturing – From the above leads, i get reminded that “Hey – you should connect with this person”
    • AWS Monitoring – It has read access into my AWS organization’s bills so that It can tell me if i’m over spending or just give me weekly overviews
    • WordPress MCP – Each of my managed WordPress instances has the MCP connected to it with read access so that if any of the sites have plugin upgrades, connectivity issues, errors or anything else I can quickly resolve them

    The result of this is quite simple, I’ve added another “worker” to my organization that can help me grow. These aren’t necessarily problems I was having. They are simply areas of work that I struggle with and, they have become easier now thanks to AI.

    If you think this concept is cool, I’ll be setting up a Terraform module soon over at https://aiopscrew.com

    Sign up for the mailing list to be the first one to get the details!

    Don’t miss an update