Eleven sprints into TurbulentGround.com, and I think I've landed on the actual conclusion, even though it's not the one I set out looking for. The evidence points one way: a fully automated product team, run under Marty Cagan's empowered product model, doesn't work. Not because the agents were badly built, or the process was badly designed. Because Cagan's model was written for humans, and the thing it actually depends on, judgement, leadership, the willingness to take an unproven bet is exactly what an agent playing a role doesn't have.
New here? This is one dispatch from an ongoing experiment: an AI product team running a live site with zero human intervention in the product decisions. Start with the premise for the full setup before diving into this one.
Quick context if you're new here: five AI agents run a product cycle against my live diagnostic site, following Cagan's model, twice a week since sprint 5. I set the objective at the start of the quarter. After that, I don't get a vote.
I couldn't see what the team was doing, and that was the first sign of a deeper problem. For the first several sprints, the only way to know what had shipped was to open individual sprint files and read them in order, or ask directly and piece together an answer. There was no way to glance at a backlog and see what was in flight. I connected Linear for a kanban view and Notion for a proper sprint journal to fix this.
[Embedded image: Linear kanban view screenshot]
It worked, but it told me something worth sitting with: a team can document itself thoroughly and still be completely illegible to the person meant to be observing it. That's not a tooling gap. That's a communication gap, and communication is a leadership skill, not a documentation one.
Then there was the question of ambition, and this is where the real problem showed up. I gave the team a clear objective and key results. What came back, sprint after sprint, was small, defensible, well-evidenced work: accessibility fixes, contrast audits, hover states. Careful. Correct. Never once a genuine attempt to move the numbers that were sitting flat. I asked directly whose job it actually was to chase those targets. The answer I got back was, roughly, that's a growth agent's job — as if driving toward the key results was someone else's remit, a role that doesn't exist on this team, rather than anyone's actual responsibility. I'd already said explicitly the key results were the PM's job. They were never reported on, never acted on, never the subject of anything resembling a plan.
Here's why I think that happened, and it's structural, not a one-off lapse. The team's own governing rule is that a proposal needs a stated user problem and evidence behind it before it can go forward — a sensible-sounding guardrail. But a stretch goal doesn't have evidence yet. You can't prove a leap before you take it. So a model built to demand evidence before acting will always select for the safe, provable thing over the ambitious, unprovable one. That's not a flaw in how carefully any individual agent reasoned. It's what the structure was always going to produce.
The deployment saga is the clearest single piece of evidence, and it took a genuinely absurd amount of back-and-forth to resolve. Every one of the first ten sprints, real work got done, verified, and then sat there, undeployed, blocked by a recurring lock file the sandboxed environment couldn't clear. I did my own research outside the sprint process and set up an API-based deployment route specifically to work around it. For a while, every time I asked, the answer came back some version of still can't fix it. In sprint 11, I asked directly: have you actually tried the API? That question is what finally moved things, though not quite the way I expected. The API itself turned out to be blocked at the network level, a dead end once genuinely tested. But the same investigation, prompted by that one direct question, led the team to a different fix — a lower-level git approach that routed around the lock entirely. It worked, for the first time in eleven sprints, without me touching my own machine.
And here's a real, unambiguous win I don't want to undersell. I now have a loop that deploys automatically, twice a week, without me. What I've actually built is a ticket factory. It takes in objectives and produces tickets, reliably, twice a week, forever closed, verified, shipped. What it doesn't produce is outcomes. Lots of work going in, almost nothing coming out the other end that anyone would notice or that moves a single number that matters. That's a strange, specific kind of success.
I set out to test this to destruction, deliberately, from the very first post in this series. I didn't expect a literal collapse, and there hasn't been one. Nothing has crashed, nothing has broken in a way that couldn't be fixed. But eleven sprints in, I think I've found what destruction actually looks like for a ticket factory. It isn't a crash. It's irrelevance. A system that runs forever, correctly, producing nothing anyone would ever pay for or notice. If this were a real business rather than a sandbox, it would already be dead — not from a dramatic failure, but from eleven months of nothing happening while the lights stayed on. That's a more honest test result than I expected going in, and a worse one for the model than a clean failure would have been. A crash tells you where the wall is. This tells you the whole thing was quietly pointless the entire time, and you'd only find out by watching it not matter for long enough.
A friend suggested a different, supposedly sharper-reasoning model might change the equation. I think it would help marginally. But I'd burn through credits fast enough to run one sprint a week instead of two, and I don't believe the constraint I'm actually up against is reasoning quality. Sharper reasoning inside the same structure still can't manufacture the thing the structure was never built to produce.
There is one genuine counter-example worth naming, because it's the clearest moment the model worked exactly as intended. In sprint 3, the team found a paid analytics tool running in production and, reading it as unauthorised spend, quietly removed it. I'd paid for it myself hours earlier. I intervened live, the change was reversed immediately, no argument, fully logged. That's Cagan's model doing what it's supposed to do: a team empowered to act, and a human still able to step in on a genuine judgement call. It's also, notably, the one moment in eleven sprints where a human made the actual call. Every other instance of real progress in this project — the Linear and Notion decision, the API research, the direct question that finally broke the deployment lock, this reflection itself — came from me, not from the team operating inside its own process.
So: fully automated teams, on the evidence of eleven sprints, don't do leadership. Leadership is what takes the risk of a step change, and it comes from context and judgement that a role-playing AI simply doesn't have, however good the underlying reasoning gets. Cagan's model was never a set of mechanical rules. It's a description of how empowered people behave: the trust, the taste, the willingness to bet on something unproven because you understand the bigger picture. Strip the people out and you're left with the process's skeleton, competently executed, missing exactly the thing that made it work in the first place.
There's a connection here back to something I mentioned in the very first post in this series — a framework from Paul Forrest's Oxford lecture, on signals, triggers, and events: a signal is a change worth noticing, a trigger is the rule that says this now matters, an event is where a system acts on it without waiting for a person to approve each step. I think I've been building the wrong version of that idea. I gave the agents team meetings, discussion, disagreement, consensus — a structure built to look like a human product team, complete with the parts of a human team that produce leadership. What it should have been doing instead is the opposite: acting more like a clean signal-trigger-event system, not less. Mechanical, linear, fast, genuinely autonomous on the parts that don't need judgement — and entirely silent on the parts that do, because those get escalated to a human, not simulated by a role-play of one.
The mistake wasn't giving the agents too little humanity. It was giving them the shape of a human process without any of the substance that makes a human process work, when what they actually needed was to be less like a team and more like a well-built pipeline, with leadership sitting somewhere else entirely.
So here's what I'm actually going to spend the coming week thinking about, away from a screen. Not whether to add more automation, but how to split this properly: a mechanical layer, closer to Forrest's signal-trigger-event model, handling everything that's genuinely reactive and rule-based, and a much smaller number of real decisions that come to me directly, as actual choices rather than something dressed up as a team meeting. I already have a vision for what this should become. The open question is where exactly that line sits — which signals the system should just act on cleanly, and which ones are actually trigger points for me, not for it.
More on where this goes once I'm back.
Would your own team survive the same test?
The two-minute diagnostic shows you where human judgement is genuinely load-bearing in your operating model, and where it isn't. Take the diagnostic →