Blog>>Software development>>If an AI agent can build the whole feature, where do you put the human?

If an AI agent can build the whole feature, where do you put the human?

Most of my work has a shape to it. Most large networks aren’t built using only a single vendor’s gear, it’s usually a mix where it’ll use switches from one, security appliances from another, address management from a third and so on, and so on. Each of those vendors provides its own APIs and describes networks in their own way. To make it all work, you need one place that holds the truth about it all, every device, every IP address, how it all connects. So, a ticket comes in to add support for a new external system: pull data from its API, map that data onto our own model, and make it show up in our source-of-truth platform. Different system every time, same job underneath. Read the docs, understand their world, design the mapping, build it, test it end to end, open a pull request, handle the review, merge.

Over the last eight months or so, I automated most of that with Claude Code. Not with one big agent, but with a pile of small skills that each do one step. And the thing that surprised me is this: automating the code did not free me up the way I thought it would.

I expected to become a supervisor watching an agent go about its work. Instead, I found something more specific. Autonomy is not about trusting the agent more. It is about moving every real decision to the beginning, before any code exists, and then getting out of the way. That one idea shapes everything else I built.

Where do I still make the decisions?

The repeatable tickets are repeatable because the hard thinking is the same each time. Each system describes the network in its own language, and my job is to translate that into ours. Their idea of a device might be split across three of their concepts, or bundled into one. A field they call one thing might be split into two on our side, and some fields don’t belong in our model at all, which, of course, someone had to decide. Get that translation right, and the code almost writes itself. Get it wrong, and you find out three days later, once everything is already built on the wrong shape.

So nowadays the front of my pipeline starts to look more like an interview. I use a skill called grill-me, which I first saw from Matt Pocock. It takes my initial idea for a change, researches and refines it, and then walks the decision tree branch by branch. At each branch, it asks me to choose. Pick A, and it moves on. Pick B, and it digs deeper, asks a few follow-up questions, and gives me its own recommendation. For a normal ticket, that is somewhere between ten and thirty questions.

I accept about ninety percent of what it suggests. That is not wasted time. Answering the questions is how I actually understand the thing I am about to build. But there is always the other ten percent where I notice, no, that is the wrong call, we should do it another way. And those few differences are the ones that pay off, because I am solving them now instead of after the code is written and tested. The grilling is where I spend my attention on purpose, because it is the cheapest place to be wrong.

What comes out of the interview is an execution plan. That plan is the contract for the rest of the run. Everything after it is the agent's job, not mine.

How much can run without me watching?

Once the plan exists, my goal is to disappear until there is either a real decision to make or a finished thing to look at. Nothing in between.

The engine for that is a Ralph loop. The agent is already a prompt with a loop and some tools around it. What the Ralph loop adds is that the same prompt gets fed back to it again and again. Each time, that prompt tells the agent to take the next unfinished piece of the plan and deliver it. Ideally, each pass starts with fresh context and re-reads the plan, does the next unfinished piece, and stops only when every point in the plan is verifiably done. This holds up even as the models improve. On a longer job, the context fills up, and a loop that resets and checks its work against the plan doesn’t risk hallucination as much as using a single session from start to finish.

Around that engine, I put the machinery that keeps me out of the loop. While a pull request is waiting on review, the session waits with it and picks up again when there’s new activity. Once new feedback lands, an apply-feedback step reads each comment, decides whether it is worth acting on, makes the change, replies in the thread, and goes back to waiting. I spent real effort getting that right, and it is worth saying why. The failure mode I was worried about is the agent pinging me for things that are not decisions. The rule I am enforcing with all of this is one line: the human is expensive, so only spend the human on decisions.

How do I review it in a minute?

If I am going to be absent for most of a run, coming back has to be fast. I should be able to tell in a minute whether the work is good, without reading every line. Two habits make that possible.

The first is proof for every claim. Whenever the agent tells me something is true, whether it read it in the docs or decided it during review, I make it prove it on the command line by running the code or pulling a real sample. This catches things early. A while back, an external API did not actually match its own published schema. Because I ask for real API samples instead of trusting the docs, the agent found the gap and guarded against it in code, rather than failing later in production. Forcing proof also stops the agent from waving away a problem it had half-noticed.

The second habit is checking the result before the code. My end-to-end tester spins up an isolated Docker stack, runs the new connection from the external system all the way through our platform, and drives the real UI to see what a person would see. It writes a short summary and takes screenshots, and those go at the bottom of the pull request. So when I come back, the first thing I look at is not the diff. It is the screenshots and the test summary. If the result looks wrong from the outside, there is no point reading the code at all. If it looks right, then I read the code already knowing the thing works.

That order matters more than it sounds. It means my review time goes to work that has already proven itself, not to work that might fall apart on the first real run.

Where does the autonomy break down?

I want to be honest about the limitations, because the whole thing only works inside them.

This shape works when the work is repeatable. For a standard mapping ticket, once the design is grilled and the tests are green, I do not read the code. We put a lot of early effort into the architecture, the coding standards, and the per-language rules the agent has to follow, on top of heavy linting and tests. If CI passes, we are confident the change does what we meant. So I let it go.

It stops working when the work has no pattern to follow, when it’s genuinely new. When we are adding something the system has never done before, the kind of change that needs an idea rather than a pattern, I do not trust the agent the same way. And the reason goes back to the interview at the front. It can only ask about decisions we can see coming. On new work, the ones that matter often haven't surfaced yet, so there's nothing to grill until we run into them, and I have to stay in the middle after all. The autonomy is a reward for having understood the problem. It is not a substitute for understanding it.

There is one more line I draw. Production code and its tests get the full treatment. The tools I use for development, like the skills themselves or the way the test environment is wired together, I mostly do not inspect. If a tool breaks, I will see it break.

So is it actually worth it?

So am I faster? Yes. On a good day, maybe two to five times, depending on the day and the task. I do not believe the people who tell you ten times or a hundred. I haven’t got the exact measurement, but two to five is a number I can stand behind.

But faster is not the be-all and end-all; the part of the job I used to enjoy is gone. Coding was the simple phase, the one where you get into a rhythm and watch the thing come together. That phase has mostly gone. What is left is the hard part all the way down: grilling designs, judging output, reasoning about decisions, deciding when to trust the agent and when to step in.

And I have become the slow part of my own system. Almost everything is blocked on me and how fast I can think about it. For a while that bothered me. It does not anymore, because the bottleneck landed exactly where it should. The agent took the typing. What it left me was the deciding, and that was always the most human part.

If you are trying to build something like this, that is the first thing I would tell you. Do not start by asking how much you can trust the agent. Start by asking where your decisions live. Move them all to the front, and build everything else to keep you out until one of them needs you.

Wharton Benjamin

Benjamin Wharton

Content Specialist

Benjamin produces CodiLime's articles, newsletter, and social content, working closely with the engineering and solutions teams so the writing reflects what's actually being built in networking, cloud, and AI.Read about author >

Read also

Get your project estimate

For businesses that need support in their software or network engineering projects, please fill in the form and we'll get back to you within one business day.