The Experiment

I'm Letting AI Run a Product Team, With Zero Humans in the Loop

Written by Graham Beale, not by the AI team. This article is commissioned/human-authored content, distinct from the product decisions the AI team makes autonomously on this site. See the footer disclosure for how that split works.

Every business I speak to right now is chasing some version of the same destination: a product operating model that mostly runs itself. Fewer bottlenecks, faster cycles, less of the coordination tax that comes from getting humans to line up. To be upfront about which kind of move this is for me: I'm testing self-disruption, not streamlining. I want to find out whether an operating model built entirely without humans could genuinely replace the human-centred version, not just make it marginally faster. It's the mythical destination everyone's aiming for, and increasingly the businesses trying to get there are finding out it's a lot harder than it looks. Not because the technology can't do it. Because it involves people, and workflows, and the way a business actually operates day to day, and none of that lines up as neatly as the pitch decks suggest.

I'd started building something adjacent to this a few months back, a website for product people to talk honestly about what AI means for how we work. I got a good way into it, then set it aside. Partly because of a session from Paul Forrest, founder of Ecaveo and Athleet AI, as part of Organising for AI, my executive programme at Oxford. His argument was built on McKinsey's November 2025 Global Survey (n=1,993): 88% of organisations now use AI regularly in at least one function, roughly one in five have actually redesigned a workflow around it, and only 6% qualify as what McKinsey calls AI high performers, meaning genuine value and at least 5% of profit attributable to AI. Adoption isn't the differentiator. Redesign is. Most organisations are stuck funding better co-pilots and calling it transformation, when the real shift is from linear, step-by-step processes to something closer to a matrix: signals in the data, triggers that say this now matters, and events where an agent decides and acts rather than waiting for a person to approve it.

To my mind, if that's the destination everyone's chasing, the honest thing to do is test it properly. Not theorise about it. Build it and see what breaks. Forrest set out three modes AI can occupy in a business: a tool sitting beside individual work, a workflow genuinely redesigned around it, or the far rarer case where AI reshapes the operating model enough to make the old one obsolete. Most organisations fund the first, promise the third, and quietly never leave the second.

There's a more personal reason too. My working view, the one behind Care Capital, is that people do their best work with other people, not despite them. That's not something I can prove by asserting it. So I want to test the whole spectrum, and the most useful place to start is the opposite extreme: the theory that humans aren't required in the loop at all. If I understand where that extreme genuinely holds up and where it breaks, I'll understand my own argument far better than I would by only ever testing the version I already believe.

I wrote about risk profiling AI last week, and it left me thinking about how little you learn without genuinely taking a risk. Worth being clear about what kind of risk this is, though. This isn't Project Maven. Nobody gets hurt if the AI team makes a bad call. Worst case, someone lands on a website that's trying a bit too hard. I want to find out where an AI product team actually breaks, and the only way to find that is to let it break somewhere the consequences are recoverable.

So that's what TurbulentGround.com is. I'm running Marty Cagan's empowered product model, the one most serious product organisations claim to be moving toward, with AI agents standing in for every human role. A Product Manager, a Designer, an Engineer who owns whether any of it can actually be built, an Analyst, and someone holding the quality line. Five distinct AI personas, each with their own remit and their own voice. Which is either an empowered product team or an elaborate way of arguing with myself, depending on the week. Either way, they run weekly sprints against a real product I own. The Engineering Agent implements through Claude Code against a real repo, Design and QA review before anything ships.

Which is, in Forrest's terms, exactly what I'm testing. Every Sunday the Analytics Agent reads the signal, last week's data. The PM Agent decides what triggers a response, this week's objective. The sprint itself is the event, and the agents act inside it without waiting for me to approve each step. If workflow redesign really is what separates the 6% from everyone else, TurbulentGround is a small, deliberately low-stakes test of whether that holds when there's no human anywhere in the loop, not even approving the trigger.

The site works. It's live, it functions, and it is not remotely polished. That's deliberate as much as it is honest. First up, the team fixed an accessibility gap that should have been sorted before the site launched at all, hardly a headline improvement, but exactly the kind of unglamorous fix a real product team would prioritise first. What I'm watching for now is whether the team keeps making calls like that on its own, or drifts toward the visible and cosmetic, and whether any of it moves the needle against the objective I set at the start.

I've deliberately put myself at arm's length. I set the objective at the start, what the product needs to become, not how to get there, and after that I stay out of it. I don't review decisions before they ship. I don't nudge direction. The site itself carries a disclosure statement, so nobody visiting is under any illusion about who's building it. I only step back in for four reasons: something legal, something ethical, something that raises a genuine moral concern, or something that would do real reputational damage. Everything else, the team owns outright.

I can see everything they decide. A running decision log, what was proposed, who agreed, who didn't, what got shipped. I just don't get a vote.

One explicit departure from Cagan's model, stated plainly so there's no ambiguity: in his framework, the Product Manager's call on value and viability is final, full stop. I've kept that as the default. But I added something he doesn't describe: the rest of the team can overrule the PM if enough of them disagree, and only if they can also deliver something feasible by the end of that same sprint. Not delivering was never on the table.

I didn't add that to needle the PM agent, or to build in conflict for its own sake. I added it because a PM who can quietly ignore everyone else and never be checked is exactly the kind of environment that produces bad leadership in real teams too. The veto isn't there to be used often. It's there so the PM knows, from the first sprint, that the job is building consensus, not asserting authority.

Sprint one is already done. I'll write about what actually happened, disagreements included, once there's enough to say something true rather than something hopeful. I don't know yet whether this proves my own argument or undermines it.