When to Stop Testing New AI Tools and Just Pick One
There's a mathematical answer. It's smaller than the famous one, and here's why.
Mathematicians have a name for your problem. It's called optimal stopping, and it's been studied since the 1960s: options arrive one at a time, you can't evaluate them all at once, and at some point you have to commit.
The famous answer is 37%. Look at roughly 37% of your options without committing, then take the next one that beats everything you've seen. Gilbert and Mosteller proved this optimal in 1966, and remarkably the win probability stays near 37% whether you're choosing between a hundred options or a hundred million.
It's a genuinely elegant result. It's also the wrong tool for choosing AI software — and the academic literature on the rule says so more bluntly than most people who quote it realise.
So here's the number that actually applies, why it's lower, and the stopping rule that beats any percentage.
📋 In This Guide
The 37% Rule and Its Fine Print
The strategy is often summarised as "observe, then leap." Explore about a third of your options while committing to nothing, then jump on the first one that's better than everything so far.
Two details rarely make it into the popular version.
First, the rule only wins 37% of the time. That's not the cutoff percentage doing double duty by coincidence — the exploration fraction and the success probability both converge on the same number. Even playing perfectly, you fail to land the single best option nearly two-thirds of the time. Which is oddly freeing: the optimal strategy is mostly wrong.
Second, the rule depends on five assumptions:
- You know how many options exist
- They arrive in random order
- They can be ranked unambiguously
- Each decision is immediate and irrevocable
- You want the single best one
Read those against your situation. Two of them are badly wrong for AI tools, and they pull in opposite directions.
Why It Breaks for AI Tools
Your decisions aren't irrevocable
The classic problem assumes a rejected candidate is gone forever. AI tools don't work that way — trials are free, most have no contract, and a tool you dismissed in March is still there in September, probably improved.
Work on the recall variant of the problem shows that when you can go back, the optimal strategy shifts toward more exploration, not less. Every option you've seen becomes a fallback with genuine option value, and that value justifies searching longer.
Which is uncomfortable, because that's precisely the behaviour trapping you. The mathematics, on this dimension alone, endorses your tool-hopping.
You don't need the single best one
This is the assumption that matters, and it's the one the rule's critics attack hardest.
One analysis of the generalised problem argues the 37% rule is misleading precisely because most people aren't holding out for the very best. Its authors show that when your goal is a good outcome rather than the optimal one, the correct cutoff is far smaller — on the order of the square root of the number of options rather than a fixed third. They frame it explicitly as a challenge to the 37% rule.
And behavioural work backs this up from the other side: in practice, people tend to stop earlier than 37% anyway. Not because they're irrational, but because rejecting good options in pursuit of a perfect one is a bad trade when "good" is genuinely sufficient.
Do you need the world's best AI writing tool? Or one that writes well enough, that you know how to drive, and that isn't about to shut down? That's not a compromise. It's the actual objective.
⚡ So One Force Says Search Longer, the Other Says Stop Sooner
Reversible decisions push toward more exploration. Not needing the best pushes toward much less.
The second one wins, decisively — and the numbers below show by how much.
The Number That Actually Applies
Take the square root of your realistic candidate pool instead of a third of it. The difference is dramatic.
| Candidate tools | 37% rule says test | Square-root rule says test |
|---|---|---|
| 9 | 3 | 3 |
| 16 | 6 | 4 |
| 36 | 13 | 6 |
| 100 | 37 | 10 |
Notice how the two diverge as the pool grows. That's the point — the 37% rule scales badly precisely where AI tools live, because there are dozens of options in every category and you don't need the best of them.
One honest caveat: define your pool sensibly. If you're picking an AI writing tool, your candidate pool isn't every AI tool in existence. It's the ones that genuinely fit your budget, your language, your platform and your use case. That's usually somewhere between six and twenty, which means test three to five and commit.
The Shortlist Trick (90%+)
There's a variant of the problem that fits your situation even better, and its result is the most encouraging thing in this article.
If instead of one irrevocable pick you're allowed to keep a shortlist and decide at the end, the odds improve enormously. With a shortlist of just three to five candidates, the probability of selecting the best exceeds 90%.
That's the situation you're actually in. Nobody forces you to accept or reject a tool the moment you try it. You can run three trials concurrently, compare them side by side, and then choose — which is a fundamentally easier problem than the one the 37% rule solves.
So the practical procedure is:
- Define the realistic pool — tools that genuinely fit your constraints
- Pick three to five as your shortlist
- Test them against the same real task, not against each other's demos
- Choose, and stop looking
The Stopping Rule That Beats Any Percentage
Percentages help you plan the search. They don't help when a shiny new tool appears six months after you committed. For that you need a threshold, not a fraction:
Only switch when the improvement is larger than your switching cost.
Switching between tools is expensive in ways that don't appear on any pricing page — rebuilding prompts, rewiring automations, and a productivity dip lasting weeks while you rebuild habits. Software migration analysis consistently finds these costs running to multiples of the savings that motivated the move.
Which gives you a clean test. A new tool that's 10% better is not worth switching to. One that's twice as good, or that does something your current tool simply can't, is. The threshold sits somewhere in between, and it's much higher than curiosity suggests.
This rule also handles the case the percentages can't: it lets you look at new tools without feeling obliged to adopt them. You're not refusing to explore. You're refusing to switch for marginal gains.
The Cost the Math Doesn't Model
Optimal stopping theory assumes each option's value is fixed and observable on inspection. AI tools break that badly, and it's the strongest argument for stopping sooner than any formula suggests.
A tool's value to you compounds with use. Month one you're fighting the interface. Month three you know its failure modes and how to work around them. Month six you have a prompt library that took a hundred iterations to develop. That accumulated skill is worth more than most capability differences between competing tools — and it resets to zero every time you switch.
Perpetual testing guarantees you never reach month three with anything. You spend permanently in the phase where every tool is at its worst, then conclude none of them are good enough. The testing itself produces the dissatisfaction that drives more testing.
There's a corollary worth noticing. Someone who picked a mediocre tool eighteen months ago and mastered it is now outproducing someone who has tried fourteen excellent tools. Not because the tool is better. Because they stopped.
Frequently Asked Questions
How many AI tools should I test before choosing one?
Three to five, for most categories. The famous 37% rule assumes you need the single best option; when good enough is genuinely sufficient, work on the generalised problem points to a cutoff nearer the square root of your candidate pool. And keeping a shortlist of three to five rather than deciding one at a time pushes your odds of picking the best above 90%.
What is the 37% rule?
A result from optimal stopping theory: examine about 37% of your options without committing, then take the next one that beats everything seen so far. Gilbert and Mosteller proved it optimal in 1966. It selects the single best option roughly 37% of the time, regardless of how large the pool is.
Why doesn't the 37% rule work for choosing software?
It assumes decisions are irrevocable and that you want the single best option. Neither holds — AI trials are free and reversible, and nobody needs the objectively optimal tool. Critics of the rule argue it is misleading precisely because most people aren't holding out for the very best, and that the correct cutoff in those cases is substantially smaller.
When should I switch away from a tool I already use?
Only when the improvement exceeds your switching cost — which includes rebuilding prompts, rewiring automations, and weeks of reduced output while habits rebuild. A tool that's slightly better fails that test. One that does something yours genuinely can't, or that removes a constraint you keep working around, passes it.
Am I missing out by not trying every new AI tool?
You're missing more by not mastering one. A tool's value compounds with use — knowing its failure modes and having a tested prompt library is worth more than most capability gaps between competitors, and it resets every time you switch. Perpetual testing keeps you permanently in the phase where every tool performs at its worst.
The Takeaway
The mathematics of when to stop looking is real, and the popular version of it is too conservative for this problem. You don't need the best AI tool; you need a good one you know how to use.
Define the handful of tools that genuinely fit your constraints. Shortlist three to five. Test them against the same real task rather than against demos. Pick one. Then apply the only rule that matters afterwards: switch when the improvement exceeds the switching cost, and not for anything smaller.
Even the mathematically optimal strategy is wrong about two-thirds of the time. That's your permission to stop optimising and start producing.
Simple AI Tools
AI guides for people who'd rather ship than shop.
👇 💬 Drop your comment below and let us know your thoughts! ✨
Comments
Post a Comment