Skip to content

What to Look for in a Solo AI Builder

For business owners weighing whether a one-person AI Consultant + Builder can actually deliver.

The usual worry about hiring a one-person AI Consultant + Builder is headcount. Can one person really do what a team used to do? It’s a fair question, asked slightly wrong. The AI supplies labor a team used to supply: hours of mechanical work, done quickly. What it doesn’t supply is judgment: knowing what’s actually broken, in what order to fix it, and whether the AI’s own answer can be trusted before you act on it. That last part is the one most people skip, and it’s the one I’d want you to ask about before you ask about anything else.

Here’s what I’d actually look for, if I were on your side of this decision.

Does this person check the AI’s confidence, or just trust it?

Before I gave an AI agent the ability to recommend fixes to live cloud infrastructure, the kind of change that can take down a service if it’s wrong, I wanted to know something the vendors don’t publish: when the AI sounds confident, is it actually right? So I ran the same agent through 30 real network fault scenarios, across three different AI models, and scored two things separately: did it diagnose the problem correctly, and was its recommended fix actually safe to run. The diagnosis came out right 90% of the time. The fix it recommended was only safe to run 73% of the time. And the model sounded equally confident either way: in 29 of 30 runs, it reported “high confidence,” whether the answer was right or dangerously wrong.

That gap between sounding right and being right is the whole reason I don’t trust a model just because its answer reads smoothly. I wrote up the full findings here, and a companion piece on how to actually pick a model for work that matters, if you want the detail rather than the summary.

I won’t run a 30-scenario, 3-model evaluation for every client task. What carries forward from doing it once, properly, is knowing how to think about whether an answer can be trusted, and building that scrutiny into smaller decisions by default. It’s the same instinct behind the Agentic Safety Shell, a gate that stops an AI agent before it acts on something risky, because “it sounded confident” was never a good enough reason to let it proceed.

That habit didn’t start with AI. It’s twenty years of systems architecture before any of this: designing network protocols that had to work correctly the first time, work that turned into patents, work where a mistake doesn’t show up as a bug report but as a data center outage. AI didn’t teach me to be skeptical of my own conclusions. It just gave me a much faster way to apply that skepticism to a lot more problems.

Ask them: how do you know when it’s wrong, not just when it sounds confident?

Does the work outlast the engagement, or leave when they do?

Every fix I make ships with the root cause, the before-and-after evidence from your actual data, and a record of how it was found, not just the fix itself dropped in with no explanation. For clients running on Claude, that record becomes something more useful than documentation: a reusable skill their own team can run again, on the next instance of the same kind of problem, long after my engagement has ended. That’s how the engagement I’m currently running is structured: sequenced in phases, each validated before the next begins, with the client’s team able to see and use what was built at every step, not just at the end.

Ask them: if you disappeared tomorrow, could my team still run what you built?

Does their diligence show up when nobody’s checking?

I recently moved my own podcast’s website, a project with no client, no deadline, and nobody checking my work, from a Webflow subdomain to a domain I’d owned but never used. Nobody would have noticed if I’d cut corners on it. I didn’t. An AI-written script had a bug that would have broken the live site, and I caught it in a review pass before running it. The category pages came up empty after migration, and I tracked down the actual cause instead of guessing. The full account is here, warts included. It’s a small example, but it’s the same habit at zero stakes, which is exactly why it’s worth checking for.

Ask them: can you show me something you did carefully when nobody was paying you to?

What this means for you

Headcount tells you how many people will work on your problem. It doesn’t tell you whether any of them, including the AI, can tell the difference between a confident answer and a correct one, whether what gets built survives the engagement, or whether the diligence is real or just performed for the invoice. Those are the three questions worth asking, whoever you hire.

If this is the kind of judgment you’re looking for

I work with small and medium businesses that have a system, a process, or a function that isn’t working the way it should, and need someone to find out why and fix it properly. The engagement I’m currently running shows what that looks like in practice, phase by phase, as it happens.

If something in your operations fits that description, I’d be glad to talk: ranga.sampath@gmail.com