Home Learnings Take the diagnostic →
Illustration for I'm Going to Make My AI Team Take Risks. I Don't Know If It Can.

The Experiment

I'm Going to Make My AI Team Take Risks. I Don't Know If It Can.

Written by Graham Beale, not by the AI team. This article is commissioned/human-authored content, distinct from the product decisions the AI team makes autonomously on this site.

Seventeen sprints and my AI product team has never taken a risk. Not one. Not a speculative bet, not an unproven idea, not a single thing it couldn't justify with evidence before starting.

New here? This is one dispatch from an ongoing experiment: an AI product team running a live site with zero human intervention in the product decisions. Start with the premise for the full setup before diving into this one.

It has also never removed anything. Not a page, not a sentence, not a claim.

I'm about to give it explicit permission to gamble, and I genuinely don't know whether that produces anything resembling judgement, or just randomness with better paperwork.

Both of those, the caution and the question, are my fault, and they have the same cause.

The team operates under a governing constraint that sounds unarguable. A proposal needs a stated user problem and evidence behind it before it can go forward. No speculative work, no vanity projects, no building things because they seem interesting.

I wrote that rule. I'd defend it in most contexts. And I now think it's quietly responsible for almost everything that's gone wrong.

Why the rule fails

A stretch goal has no evidence. That's what makes it a stretch. You can't prove a leap before you take it, so a system that demands proof before acting will never take one.

Deletion has it worse. "Remove this paragraph" has no user problem behind it that any metric will demonstrate. Addition is always defensible, because you can point at the thing you added and describe who it helps. Subtraction can't be argued for on the same terms.

At one point two parallel sessions independently added copy saying the same thing in different words, and the site now carries three different descriptions of the same promise, all live, all trying to say one thing.

The Design agent spotted two of the three contradicting each other. It logged the tension and moved on, correctly identifying that resolving it needed judgement it had no mandate to exercise.

That's the rule working exactly as written.

What a human would do here

Any product person worth hiring takes bets they can't justify. They cut things that technically work because the whole reads better without them. They act on taste, on pattern recognition, on a hunch, and they accept that some of it will be wrong.

None of that survives an evidence-first rule. What's left is the fraction of good product work that happens to be provable in advance, which turns out to be the small, safe, incremental fraction.

So the question I'm actually testing next is whether you can put that back deliberately. Whether an agent given explicit permission to gamble will produce anything resembling judgement, or whether it will just produce randomness with better paperwork. Taste might be the thing that doesn't transfer. I don't know yet.

The three rules I'd need to write

Permission to remove. An explicit carve-out saying the team may propose removing something without a stated user problem, on the grounds that coherence is itself a user problem even when no metric can show it. This is the one I'm most confident about and least comfortable with, because it means handing an autonomous team permission to delete things. I haven't written it yet.

One unevidenced bet per sprint. A mandated slot for something the team believes might work but cannot prove. Not optional, because anything optional under an evidence-first culture gets deferred. Required, and logged as a bet rather than a decision, so it's judged on what it taught rather than whether it worked. Writing this one means admitting the day-one rule has been shaping every outcome since, and not in the direction I intended. That's harder to write than it sounds.

Kill criteria, defined in advance. This one is for me, not the team. I've just committed to daily sprints, a second provider and a new role, and I haven't written down what success looks like. Without that I'll be in the same position in three weeks, several thousand tokens down, unable to say whether the attempt failed or simply needed longer. Deciding afterwards what would have counted as working is not evaluation, it's storytelling. Writing it down first means accepting in advance that I might have to stop, and I've been running this long enough to have got attached to it continuing.

The part that bothers me

None of these three ideas came from the team. In seventeen sprints it has never proposed a change to its own rules. It has investigated symptoms with genuine care, dozens of times, and it has never once asked why the same shape of problem keeps recurring.

I spent four sprints looking at individual outputs without asking whether the instructions producing them were even the ones I'd written, and I only noticed because I went away for a week. The team's blind spot and mine turned out to be the same blind spot.

So I still don't know if it'll actually take the risk. But I know why neither of us thought to ask.

Would your own team survive the same test?

The two-minute diagnostic shows you where human judgement is genuinely load-bearing in your operating model, and where it isn't. Take the diagnostic →