Startup & Growth
Stop Scoring Ideas on Vibes
"We should test this" isn't a strategy — it's how most experimentation backlogs turn into whoever argued loudest in the meeting. The fix elite growth teams use is a four-variable score, not a gut check.
3 min read · 7/7/2026 · Inspired by Reforge
"We should test this" is not a growth strategy. It's also, according to research on how elite teams actually operate, the reason most experimentation stalls: ideas get greenlit by whoever argues loudest in the meeting, not by any consistent measure of what's worth testing.
The fix that growth teams at scale actually use is embarrassingly simple. Score every test idea 1–10 on three variables — Impact (how much would this move the metric if it wins), Confidence (how sure are you it will), Ease (how fast and cheap is it to ship) — multiply them, and rank the backlog by the result.[1] It's called ICE, and Sean Ellis-era growth teams built entire prioritization systems on nothing more complicated than that.
The gap most teams miss: ICE alone still lets a team optimize for busywork. A test that's easy, confident, and mildly impactful can outscore a test that would actually move revenue. The fix growth architects add is a fourth variable — Revenue Weight — multiplying the ICE score by how directly the test connects to CAC payback or LTV, not just a vanity metric like click-through rate.[2] A homepage color test and a checkout-flow fix might score similarly on ICE. They should not score similarly once revenue weight is in the formula.
The second number worth tracking isn't win rate. It's learning yield — the percentage of tests that produce a reusable, actionable insight, whether the variation won or lost. Benchmarks from elite growth teams put this at 70% or higher, running 8–12 tests a month, with roughly half of new tests directly informed by a prior test's learnings — meaning the backlog compounds instead of resetting every quarter.[2]
What this looks like on a real account: on the QuickLets rebuild, the highest-leverage tests weren't the ones with the flashiest UI — they were the search and saved-property flow changes tied directly to enquiry volume, prioritized ahead of cosmetic changes that would have scored fine on Impact and Confidence alone but carried no revenue weight.
Budget size doesn't predict acquisition efficiency nearly as well as test velocity does.[3] A team running ten small, revenue-weighted tests a month will usually out-learn a team running one large campaign relaunch a quarter — regardless of which one has the bigger budget.
Sources
Related client work
QuickLets Redesign
Modern Malta-wide rental platform — 62,000+ properties, 435 specialists, advanced search, and QLZH ecosystem integration. Migrating to quicklets.com.mt.
62,451 Properties