Experimentation System and Integrated Growth Engine
Module 8
This module assembles the full experimentation operating system: ICE prioritisation, the eight-column backlog, pre-registered decision rules, vanity metric audits, capacity allocation under constraint, and the 90-day growth roadmap that integrates the whole engine.
Impact, Confidence, and Ease (ICE) Prioritisation
Prioritisation is the fundamental problem for growth teams when ideas outnumber execution bandwidth. Growth planning typically starts with tens or hundreds of ideas, but teams can realistically ship only 5 to 7 experiments per month, or a maximum of 20 per quarter. Under 10 percent of executed experiments actually yield a positive impact on core business metrics. Without a structured framework, teams risk selecting incorrect ideas based on opinions rather than evidence. The Impact, Confidence, and Ease (ICE) framework solves this problem by using a transparent, repeatable score on a 1 to 10 scale to discipline executive judgment.
ICE Framework Components
Definitions and key questions for the three dimensions of the ICE framework are structured as follows:
| Dimension | Core Definition | Key Scoring Question |
|---|---|---|
| Impact | The estimated magnitude of movement that an experiment will produce on the core metric. | How much will this experiment optimize the metric that is tied directly to the primary growth constraint? |
| Confidence | The level of certainty that the predicted outcome will occur as hypothesized. | How sure is the team that the predicted outcome will happen, and what evidence supports this certainty? |
| Ease | The speed and effort required to ship, measure, and analyze the experiment. | How fast can the team build, launch, and extract usable data from this experiment? |
The overall ICE score is computed as the average of the three individual scores. Backlog items are sorted by their ICE scores in descending order to determine execution priority.
Impact Scoring Discipline
Impact scores must be anchored strictly to the primary growth constraint. A high score on a non-constraint metric does not count toward business growth, regardless of how much that metric improves.
| Impact Score | Movement on Constraint | Metric Tier Affected | Criteria and Guidelines |
|---|---|---|---|
| 9 to 10 | 20% or more shift. | Primary Constraint. | Shifts the primary constraint directly by a significant margin. Extremely rare. |
| 6 to 8 | Meaningful shift. | Secondary Metric. | Creates a strong impact on a secondary metric or a moderate shift on the constraint. |
| 3 to 5 | Modest shift. | Supporting Metric. | Moves a supporting metric with minimal impact on the primary constraint. |
| 1 to 2 | Negligible shift. | Vanity Metric. | Moves a vanity metric only, which fails to alter business health or performance. |
If an experiment does not touch the primary growth constraint, it cannot receive a high impact score. For example, if the primary constraint is activation, an A/B test of a pricing page designed to lift conversion from free trial to paid by 20% is downstream from activation. If only 34% of users activate, the conversion lift is applied to a thin slice of users. Thus, the pricing experiment should be scored only a 5 or 6 on impact, preserving scores of 9 for experiments that directly target the core activation rate.
Confidence Scoring Discipline
Confidence is the most commonly abused score in growth frameworks, frequently inflated by founders or team leaders to justify preferred ideas. Scoring must be bound to objective evidence.
| Confidence Score | Quality of Evidence | Required Proof / Criteria |
|---|---|---|
| 9 to 10 | Strong, direct evidence. | Based on prior experiments conducted in the same product, on the same metric, or a near-identical pattern at a competitor cited by name or report. |
| 6 to 8 | Indirect evidence. | Competitor SaaS pattern with a comparable ICP, plus qualitative signals or interviews from own users indicating the same friction point. |
| 3 to 5 | Logical hypothesis. | Sound reasoning and logical hypothesis, but lacks direct product data or past results. |
| 1 to 2 | Pure guess. | Brainstormed ideas that feel right but lack supporting numbers, data, or competitive case studies. |
The core discipline check requires that if the team cannot state the evidence for a confidence score in a single sentence, the score is capped at a maximum of 5. Statements like "it just feels right" or "everyone is doing this" automatically cap the score at a maximum of 4 or 5.
Ease Scoring Discipline
Ease counts the full cycle of building, measuring, and analyzing, not just the technical build time.
| Ease Score | Timeframe to Usable Signal | Team / Infrastructure Resources Required | Examples |
|---|---|---|---|
| 9 to 10 | 1 day or under. | Shipped instantly, measurements already in place with zero custom engineering. | Content change, ad copy change, email subject line test, simple price tweaks. |
| 6 to 8 | 1 to 2 weeks. | Light instrumentation, minor tool integrations. | New onboarding screen, basic API tool integration, tracking a new product event. |
| 3 to 5 | 3 to 6 weeks. | Collaborative work across multiple departments (e.g., product, marketing, engineering), requires new metrics. | Launching a new referral mechanic, full pricing page redesign, building a small feature. |
| 1 to 2 | 2 months or more. | Major engineering hours, custom infrastructure, platform architecture changes. | Platfrom rebuild, deep database changes (should not be in an experimental backlog). |
The core trap of ease scoring is ignoring the time required to collect data. An experiment that takes 2 days to build but requires 6 weeks to yield a statistically credible signal cannot be scored a 9. It must be scored a 6 at best, as time-to-usable-result is what dictates prioritization.
ICE Prioritization backlogs (Clairo vs. Zoko)
The ICE framework adapts to different product funnels by anchoring to brand-specific constraints.
| Parameter | Clairo Case Study | Zoko Case Study |
|---|---|---|
| Primary Constraint | Activation: 66% of free sign-ups never record a meeting. | Activation: 62% of first-time starter kit buyers do not repurchase within 90 days. |
| Underlying Friction | Funnel drops off before users experience core value. | Customers do not use products consistently enough to see visible results. |
| Top Experiment (ICE 8.0) | Cut onboarding from 4 steps to 2 (Impact: 9, Confidence: 8, Ease: 7). | Send Day 7, 14, and 21 WhatsApp photo prompts (Impact: 9, Confidence: 8, Ease: 7). |
| Supporting / High-Ease Experiment | Default auto-record toggle to on (Impact: 8, Confidence: 7, Ease: 6, ICE: 7.0). | Refer and Glow card inside unboxing kit (Impact: 7, Confidence: 7, Ease: 9, ICE: 7.7). |
| Low-Priority / Out-of-Constraint Ideas | Homepage hero copy (Impact: 5, Confidence: 5, Ease: 9, ICE: 6.7); Slack integration (Impact: 5, Confidence: 5, Ease: 2, ICE: 4.0). | Influencer myth-busting videos, TikTok creator partnerships, Amazon subscription toggle (ICE: 4.0 to 5.0). |
| Framework Logic | Homepage copy is easiest but ranks fourth because it fails to target the activation constraint directly. | Acquisition and revenue plays rank low because they fail to address the core activation leak. |
Clairo's third-ranked idea, prebuilt meeting templates by type of sales call, also targets activation and scores an ICE of 7.3, which is why the homepage hero copy test (ICE 6.7) ranks only fourth despite being the easiest idea on the list. The framework protects the team from confusing speed with importance: the easiest experiment to ship is not automatically the most important one to run.
Failure Modes in ICE Execution
Teams often degrade the ICE framework into an opinion list through four primary execution errors:
| Pitfall | Symptoms | Root Cause | Team Remedy |
|---|---|---|---|
| Confidence Inflation | Every experiment score on confidence is written as a 9 or 10. | Individuals seek to justify their personal ideas. | Apply the single-sentence evidence constraint, capping unproven ideas at a maximum score of 5. |
| Scoring in Isolation | Backlog scores drift based on individual interpretations. | Different team members have varying definitions of what a score of 7 or 8 represents. | Score the first three ideas as a team to calibrate scoring definitions. |
| Ease Discrepancy | High ease scores are assigned to complex engineering tasks. | Leaders assume implementation is quick without consulting the engineering team. | Score ease strictly from the perspective of the executing team who builds, measures, and analyzes. |
| Static Scores | Backlog rankings remain unchanged for months. | Teams treat scores as permanent instead of dynamic. | Rescore related backlog items whenever an experiment fails or new customer data emerges. |
A spread diagnostic check determines team scoring honesty: a real backlog sheet contains a diversified range of ICE scores (some 8 to 9, some 4 to 6, some 2 to 3). If the entire backlog is clustered between 7 and 9, the framework has failed.
Growth Backlog Design
An experiment backlog is not just a list of ideas; a list is a repository where ideas accumulate and are eventually abandoned. A backlog is a system with defined rules, columns, and rituals that determines what runs, what stops, and what learnings are captured.
The typical list-to-graveyard trajectory follows a predictable timeline: in Month 1 the team opens a fresh "growth backlog" sheet and accumulates around 30 ideas within 2 weeks; by Month 3 the list holds 80 ideas but only 4 have been run and half the team has stopped opening the sheet; by Month 6 the sheet is abandoned entirely. If the backlog cannot answer "what runs, what stops, and what did we learn" within 30 seconds, it is a graveyard, not a system.
Lifecycle Stages of an Experiment
Every experiment in a functioning backlog must sit in one of five distinct stages:
| Stage | Definition | Required Status Criteria |
|---|---|---|
| Idea | A written thought or proposal. | Unscored and undiscussed. |
| Scored | ICE inputs are fully completed. | Average score calculated, ranked against the backlog, but not running. |
| Active | Currently running in production. | Single owner assigned, decision rule pre-registered, and data collection active. |
| Deciding | Target sample size or time limit reached. | Data collection complete, pending application of the decision rule to scale, kill, or iterate. |
| Closed | Experiment is finalized. | Outcome captured, decision executed, and learnings documented. |
Experiment movement is strictly unidirectional. An experiment cannot return to the idea stage; whether it fails or succeeds, it must move forward to the closed stage with captured learnings. If a row cannot be assigned to one of these five stages, it does not belong in the backlog.
Eight Columns of a Universal Backlog System
The universal backlog system relies on eight columns to turn a spreadsheet into an operating tool:
| Column | Data Type | Operational Definition and Rule |
|---|---|---|
| 1. Experiment | Text | A single-line description of the specific test. |
| 2. Hypothesis | Text | Standardized format: If we do X, then Y will change because of Z, ensuring the hypothesis is falsifiable. |
| 3. Metric | Number | A single number that determines if the experiment succeeded. |
| 4. ICE Score | Formula | Dynamic formula calculating the average of the impact, confidence, and ease inputs. |
| 5. Stage | Dropdown | One of the five defined lifecycle stages: Idea, Scored, Active, Deciding, or Closed. |
| 6. Owner | Text | Exactly one named team member; multiple names or team labels are prohibited. |
| 7. Decision Rule | Text | Pre-registered thresholds (scale, kill, iterate) locked before the experiment goes active. |
| 8. Learning | Text | Mandatory single-sentence summary of results, cohorts affected, and findings; required to close the row. |
These 8 columns work for B2B and B2C without modification: the structure is universal, and only the brand-specific thinking inside the cells changes.
Clairo Working Backlog Sheet (Template Walkthrough)
The reference implementation is a Google Sheets workbook with 12 experiment rows, 8 columns, and 2 tabs (the backlog plus a decision scorecard tab listing the top 5 experiments by ICE score with their pre-registered thresholds):
- Live ICE formula: The ICE cell computes the average of the impact, confidence, and ease inputs. Changing one input (for example, confidence from 7 to 4) instantly updates the score and the rank order, so a disputed score is resolved by editing one cell rather than debating the total.
- Conditional formatting: The top rows are highlighted automatically when the ICE score exceeds 7.5; highlights move as the team rescores.
- Stage filter as the WIP view: Filtering the stage column to "active" shows only 4 rows (the WIP limit) while 18 scored experiments wait in line, so on Monday morning the team sees what is running, not the full idea pile.
- Example pre-registered rule: "Scale if the sign-up to record rate hits 45% or above, kill below 36%, iterate in between", agreed before the experiment went active.
- Example learning entry: "Reduced onboarding from 6 steps to 3. Activation rose from 34% to 41%. Effect held in week 2 and week 4 cohorts." Six months later, a proposed onboarding redesign builds on this row instead of re-running the test.
Built correctly, the decision-making becomes structural: the team does not have to remember to do the right thing, because the sheet's structure does the remembering.
WIP Limits and Rescoring Discipline
Work in Progress (WIP) limits prevent the backlog from becoming an unmanageable list. Successful growth teams establish a strict WIP limit of 3 to 5 active experiments at one time.
Exceeding 5 simultaneous active experiments introduces three operational complexities:
- Signal Noise Compounds: Multiple simultaneous experiments run on the same cohort, making clean attribution of metric movement impossible.
- Attention Spreads: Team focus degrades, leading to delayed implementation and missed details.
- Learning Rates Fall: The team executes more tests but lacks the bandwidth to extract deep, actionable insights from each.
Rescoring must be completed regularly as confidence is a dynamic variable. However, active rows must never be rescored midway through execution; experiments must run to their pre-registered decision rules to avoid bias.
Growth Rituals Cadence
The backlog is driven by a weekly ritual consisting of two short meetings:
| Meeting | Timing | Time Commit | Core Agenda and Actions |
|---|---|---|---|
| Monday Planning | Start of week. | 45 to 60 minutes. | 1. Apply decision rules to completed experiments and move to closed.<br>2. Document learnings.<br>3. Rescore confidence on remaining ideas based on new evidence.<br>4. Select the next batch of experiments from the top of the ranked list respecting the WIP limit.<br>5. Assign one human owner per active row. |
| Friday Status Check | End of week. | 15 to 30 minutes. | 1. Quick check-in on experiment status.<br>2. Each owner provides a single-sentence status update.<br>3. Flag any experiment that has hit its decision threshold early. |
No new commitments or experiments are allowed to be added midway through the week.
Backlog Design Pitfalls
Growth backlogs generally fail within the first eight weeks due to four common pitfalls:
| Pitfall | Symptoms | Core Root Cause | Structural Remedy |
|---|---|---|---|
| The Graveyard | Backlog grows to 80+ ideas, but only 4 are run; sheet is eventually abandoned. | No constraint on backlog growth, and no prioritization enforcement. | Enforce strict WIP limits of 3 to 5 active experiments at once. |
| No Owner | Active experiments are unassigned, or listed under a general department name. | Lack of individual accountability for experiment status and measurement. | A hard rule of exactly one human name per active row; no blank fields. |
| No Rule | Completed experiments are debated endlessly; success criteria are redefined after data arrives. | Failure to lock down metrics and thresholds before launch. | Pre-register and lock the decision rule before the experiment becomes active. |
| No Learning | Experiments are closed but the learning column remains blank. | Teams focus entirely on shipping rather than capturing knowledge. | Block experiments in the deciding stage; they cannot be closed without a written learning statement. |
Decision Rules
A decision rule is a contract written before an experiment runs that determines what its outcome means after the data is collected. Pre-registering decision rules prevents post-hoc negotiation of outcomes. Without a pre-registered rule, different departments will interpret the same data to support their personal preferences, resulting in indecision.
Experiment Outcomes Comparison
Every experiment outcome must map to one of three possible decisions:
| Decision | Definition | Core Threshold Criteria | Operational Action |
|---|---|---|---|
| Scale | Fully validate hypothesis. | Meets or exceeds the pre-registered success threshold. | Implement the experiment permanently as a feature, campaign, or message. |
| Kill | Invalidate hypothesis. | Drops below the pre-registered failure threshold. | Stop the experiment immediately, do not iterate, and document learnings. |
| Iterate | Unclear or marginal signal. | Lands within the band between the scale and kill thresholds. | Run a revised version (V2) with exactly one variable modified under a new rule. |
Excluding the iterate option forces premature closure of marginal experiments. Conversely, default iteration can become an excuse to avoid admitting failure. The rule must strictly define the iterate boundaries.
Four Components of a Decision Rule
Every pre-registered decision rule requires four locked components:
| Component | Definition | Alignment / Derivation |
|---|---|---|
| 1. Single Metric | A single number that determines the success or failure of the test. | Must align directly with the primary growth constraint. Multiple metrics are prohibited because they create conflicting signals. |
| 2. Scale Threshold | The numeric value that triggers a scale decision. | Anchored directly to the projected impact in the experiment's hypothesis. |
| 3. Kill Threshold | The numeric value that triggers a kill decision. | Usually a hurdle below which the cost (engineering, infrastructure, time) of running and maintaining the change exceeds its business benefit. |
| 4. Time / Sample Bound | The limit that triggers the decision review. | Anchored to statistical power calculations (fixed sample size) or a strict timeline (fixed window) set in advance. |
Worked derivations of the two thresholds:
- Scale threshold from the hypothesis: If the hypothesis projects activation rising from 34% to 45%, the scale threshold sits at or slightly below the projection (for example, 42% to 45%). If the threshold must be set much lower than the hypothesis projected, the hypothesis itself is wrong and should be rewritten before the experiment runs.
- Kill threshold from cost-benefit: If lifting activation only to 36% would generate roughly 2 lakhs of revenue benefit while the experiment costs roughly 2 lakhs in engineering and infrastructure, any result at or below that point is a net loss, so 36% becomes the kill threshold.
Case Studies Decision Rules (Clairo vs. Zoko)
Applying pre-registered contracts across B2B and B2C brands yields distinct decision paths:
| Parameter | Clairo Onboarding Redesign | Zoko WhatsApp Photo Prompt |
|---|---|---|
| Core Hypothesis | Reducing onboarding from 6 steps to 3 will increase the activation rate. | Sending photo prompts at Day 7, 14, and 21 will increase the 90-day repurchase rate. |
| Baseline | 34% activation rate (meeting recorded within 7 days). | 22% 90-day repurchase rate of first-time buyers. |
| Single Metric | % of free users recording a meeting within 7 days. | % of first-time buyers repurchasing in 90 days. |
| Scale Threshold | 42% or higher. | 30% or higher. |
| Kill Threshold | 36% or lower. | 24% or lower. |
| Iterate Band | Between 36% and 42%. | Between 24% and 30%. |
| Bound | 500 users per arm or 14 days, whichever occurs first. | 800 customers enrolled in the cohort, waited full 90 days. |
| Actual Result | 40.3% activation rate among 14,547 users. | 23.1% repurchase rate among 812 customers. |
| Outcome Decision | Iterate: Result falls directly within the iterate band. | Kill: Result falls below the 24% kill threshold. |
Constraints of Bounded Iteration
To prevent iteration from becoming an excuse for infinite unscientific testing, three constraints must be applied:
- Single Variable Change: Version 2 of the experiment must modify exactly one element of the first version. If multiple variables are changed, the team cannot attribute metric movement to a specific cause.
- New Pre-Registered Thresholds: Version 2 does not inherit the rules of Version 1. It must have its own thresholds and bounds written and locked before launch, and these are often stricter than the first version.
- Iteration Cap: A maximum of 2 iterations (up to Version 3) is permitted. If Version 3 still lands in the iterate band, the experiment must be killed because the underlying mechanism lacks sufficient strength.
Failed vs. Successful Kills
Killing an experiment is a positive outcome if it compounds team knowledge:
| Criteria | Failed Kill | Successful Kill |
|---|---|---|
| Stage & Learning State | Row moved to closed, but the learning column is empty or reads "it did not work". | Row moved to closed, and the learning column explicitly documents what was tested, what failed, and why. |
| Compounding Knowledge | Compounding learning is zero. | Future experiments avoid the dead mechanism entirely. |
| Downstream Impact | 6 months later, someone proposes the same idea, and the team wastes resources re-running it. | Compounding knowledge accumulates over quarters, increasing future experiment confidence. |
An example of a successful kill learning for Zoko is: "Photo prompt mechanism at this cadence does not move the 90-day repurchase rate; future retention experiments should test referral incentives or ritual habit cards before retesting prompt cadence variations".
Decision Rule Pitfalls
Four common execution errors break the integrity of pre-registered decision rules:
| Pitfall | Description | Core Root Cause | Remedy |
|---|---|---|---|
| Moving the Goalpost | Editing the thresholds after data is collected. | Teams want to claim victory for marginal results. | Enforce a strict pre-registration contract; no edits to rules are allowed after launch. |
| Cohort Cherry-Picking | Slicing the data post-hoc until a statistically significant segment appears. | Searching for positive signals in a failed test. | Slice only cohorts pre-registered in the baseline rules; treat all other post-hoc segments as noise. |
| The More Data Trap | Extending the run-time of a failed test hoping the numbers will improve. | Refusal to accept a negative signal. | Respect the pre-registered sample size or time bound; make the decision on schedule. |
| Iterate Default | Marking clearly failed tests as "iterate" because the team is married to the idea. | Emotional attachment to an experiment. | Strictly follow the kill threshold; if a metric lands below the boundary, force a kill. |
All four pitfalls are detectable with one question: did the rule exist in writing before the experiment went active? If yes, the rule wins and the decision follows it. If no, the team is negotiating outcomes after the fact, which is not experimentation.
Avoiding Vanity Metrics
A decision rule is only as effective as the metric it points to. If the metric is wrong, the rule fires backwards, optimizing numbers that fail to change business outcomes or profitability.
Vanity vs. Real Metrics Comparison
To identify metrics that represent actual business health, teams must distinguish between vanity and real indicators:
| Metric Type | Core Characteristics | Behavior in Decision Rules |
|---|---|---|
| Vanity Metric | 1. Can only go up by design (cumulative counts).<br>2. Moving it does not change team actions.<br>3. It is 2 or more hops away from a business outcome. | Dangerous; they provide true numbers but point the team in the wrong direction, wasting resources on empty growth. |
| Real Metric | 1. Is expressed as a rate rather than an absolute count.<br>2. Can move up or down based on product health.<br>3. Traces to revenue, retention, or margin in 1 or 2 hops. | Reliable; every change leads to a clear operational action and predicts business health. |
Five Tests for Metric Integrity
Every candidate metric must pass five tests before it is allowed into a decision rule:
| Test | Operational Requirement | Target Standard |
|---|---|---|
| 1. Denominator | Is it a rate and not just an absolute count? | Rates are comparable across time and cohorts; counts are not. |
| 2. Comparability | Can it compare performance across cohorts and time periods? | Must use a uniform computation across cohorts to compare apples to apples. |
| 3. Two-Directional | Can the metric go down? | A metric that can only rise by construction cannot show risk or real product health. |
| 4. Actionability | If this number changes, does the team know exactly what to do? | The team must agree on which lever caused the movement and the next step to execute. |
| 5. Business Outcome | Does it trace directly to revenue, retention, or margin in 1 or 2 hops? | More hops indicate a weak, unvalidated story rather than a true operational signal. |
A candidate metric must pass all five tests; passing only 4 out of 5 is insufficient.
Metric Audits (Clairo and Zoko)
Auditing reporting scoreboards reveals common vanity metrics and their real metric replacements:
| Brand | Vanity Metric (Failing Test) | Real Metric Replacement | Business Outcome Traced |
|---|---|---|---|
| Clairo | Total free signups (8,400) (Fails Two-Directional; cumulative count). | % activation by weekly signup cohort. | Standardized rate; compares user conversion quality week over week. |
| Clairo | Total meetings recorded (Fails Business Outcome; recording does not equal value). | % of users recording a meeting within 7 days. | Strong leading indicator of long-term customer retention. |
| Clairo | 94% transcription accuracy (Fails Actionability; marketing claim, not operating metric). | % of users with zero transcription errors per month. | Operational metric that can be moved directly by engineering work. |
| Zoko | Instagram followers (Fails Business Outcome; 3 hops away, unvalidated). | Instagram attributed first purchase rate. | Measures direct buyer conversion from the social channel. |
| Zoko | Total customers all time (6,200) (Fails Two-Directional; cumulative count, includes churned users). | Active subscribers + new subscribers monthly. | Fluctuates in real time, exposing subscriber churn and slowing acquisition. |
| Zoko | Average Order Value (AOV) (1,180) (Fails Business Outcome; ignores margin). | Contribution margin per order. | Protects company profitability; AOV can rise while actual margins drop. |
The AOV trap illustrated with numbers: if Zoko's AOV rises from Rs. 1,180 to Rs. 1,400 while the margin drops from 40% to 30%, the AOV dashboard shows improvement even though profitability has fallen. Vanity metrics are not false numbers; they are true numbers that provide incomplete information and point the team in the wrong direction.
The Two-Hop Rule
The two-hop rule mandates that any metric in a decision rule must trace to revenue, retention, or margin in at most two steps.
Strong Metric Chain (Clairo):
[Activation Rate] ---> (Hop 1) [Day 7 Retention] ---> (Hop 2) [LTV & Revenue]
(Validated in cohort data: users who activate also retain, and users who retain pay).
Vanity Metric Chain (Zoko):
[Instagram Followers] ---> (Hop 1) [Reach] ---> (Hop 2) [Website Visits] ---> (Hop 3) [Purchases]
(Unvalidated: requires three hops based on assumptions rather than direct data correlation).
If a team cannot draw the correlation chain on a whiteboard in under 30 seconds, the metric does not belong in a decision rule.
The Leading Indicator Trap
A leading indicator is a metric that moves before the lagging business outcome does. For example, activation rate is a leading indicator for retention.
Leading indicators fall into two categories:
- Earned Leading Indicators: Validated internally using the company's own cohort data, proving a statistically significant correlation to a lagging metric.
- Claimed Leading Indicators: Imported from competitor blog posts or external content without internal validation. These act as vanity metrics in disguise.
Hypotheses belong in the experiment column; only validated, earned metrics belong in the decision rules.
Four Patterns of Vanity Metrics
Immature growth teams frequently fall into four common metric patterns:
| Pattern | Definition | Operational Flaw |
|---|---|---|
| 1. World-Facing Metrics | Big, cumulative numbers presented to investors, media, or the public. | Cumulative signups or app downloads only increase, providing zero operational value. |
| 2. Input as Output | Measuring team effort or activities instead of results. | Running 30 A/B tests or answering 100 tickets is an input; it does not mean the business improved. |
| 3. Engagement Metrics | Measuring absolute site sessions, page views, or app opens. | Measures general activity without verifying if it translates into user value or revenue. |
| 4. Gameable Claims | Relying on qualitative surveys like NPS or CSAT ratings. | Highly sensitive to sample selection bias and survey design; easily manipulated. |
Resource Allocation Under Constraint
Prioritized backlogs often have more high-scoring ideas than a team has the capacity to execute. Capacity acts as the second filter after the ICE framework. If teams attempt to run all high-scoring ideas in parallel without resource limits, signals collapse, team bandwidth is overstretched, and experiments fail to reach a clean conclusion.
Three Currencies of Capacity
Capacity is a combination of three distinct operational currencies:
| Currency | Definition | Measuring Unit | Bottleneck Risk |
|---|---|---|---|
| Engineering | Available technical and development hours from the tech team. | Developer hours or engineering weeks. | High in custom-built software products. |
| Decision | The attention span and review time available from the CEO or department head. | Unblocked approval hours and meeting availability. | Ignored in tracking charts but causes delays if leadership becomes a bottleneck. |
| Audience | The volume of unique user cohorts available to absorb experiments. | Unique user traffic or cohort size. | High in low-traffic or early-stage subscription products; running overlapping tests pollutes data. |
The binding currency is whichever of these three runs out first, and teams must plan execution against it.
The 70/20/10 Allocation Rule
Binding capacity must be split across three distinct experiment categories:
| Category | Sizing | Target Confidence Score | Strategic Role |
|---|---|---|---|
| Known Winners | 70% of capacity. | ICE confidence of 7 or higher. | Optimizing and iterating on already validated mechanisms; most likely to scale. |
| New Bets | 20% of capacity. | ICE confidence of 5 to 7. | Testing new hypotheses with reasonable evidence; moves into the 70% column next quarter. |
| Exploration | 10% of capacity. | ICE confidence under 5. | Brand-new mechanisms with a high risk of failure, but high option value if they succeed. |
Memory hook: Capacity split 70/20/10: 70% known winners (confidence 7 or higher), 20% new bets (confidence 5 to 7), 10% exploration (confidence under 5). On top of that, keep a 20% to 35% slack budget of the binding currency uncommitted.
Capacity Constraint Audits (Clairo vs. Zoko)
Different brand structures hit different binding currencies, forcing distinct capacity layouts:
| Parameter | Clairo Capacity Plan | Zoko Capacity Plan |
|---|---|---|
| Binding Currency | Engineering weeks: 1 full-time engineer with 12 weeks of capacity. | Audience size: 1,100 active subscribers. |
| Non-Binding Currencies | Audience is highly plentiful; decision bottleneck is low. | Shopify/WhatsApp APIs require zero engineering; founder is highly available. |
| Constraint Calculation | Immature planning commits 5 top ICE ideas totaling 13 engineering weeks, immediately overrunning the budget. | A retention experiment requires 500 users per arm (1,000 total), limiting tests to 2 parallel cohorts. |
| 70% Allocation | Onboarding redesign and email timing test (5 engineering weeks). | Subscription onboarding experiment (660 subscribers). |
| 20% Allocation | Team invite mechanic (3 engineering weeks). | New referral mechanic (320 subscribers). |
| 10% Allocation | Founder-led sales outbound pricing test (0 engineering weeks, using sales calls). | Creator partnership pilot (110 subscribers). |
| Slack Budget | 4 engineering weeks (one-third of capacity) kept uncommitted. | 110 subscribers left untouched and unexposed. |
Slack Budget
A slack budget represents keeping 20% to 35% of the binding currency uncommitted. Slack is not waste; it is the operating buffer that allows plans to survive overruns, customer crises, or strategic reallocations when stronger mid-quarter evidence emerges.
Serial vs. Parallel Execution
Growth roadmaps must lay out experiments based on user interference:
| Execution Path | Core Rule | Interference Risk | Operational Example |
|---|---|---|---|
| Parallel | Run simultaneously when tests touch different parts of the funnel or distinct cohorts. | Low; cohorts do not overlap, multiplying team output. | Pricing test on new visitors and onboarding test on day 7 users. |
| Serial | Run sequentially when tests share a cohort or modify the same funnel step. | High; simultaneous testing on the same cohort makes the signal uninterpretable. | Running 2 onboarding flow tests on the same new user cohort. |
A blocking experiment occurs when Experiment 2 depends on the outcome of Experiment 1. For example, upgrade prompt tests are blocked by pricing tests because the price determines the definition of an upgrade. Running them in parallel is logically incoherent. Blocker experiments must be scheduled to run first. The practical check for each active experiment: "Does the outcome of any other experiment in the backlog change what this experiment is testing?" If there is even a hint of yes, the other experiment is a blocker and runs first.
Killing Active Work: Reactive Kill vs. Strategic Reallocation
The hardest discipline in capacity management is stopping an experiment that is currently running. Two scenarios carry very different difficulty levels:
| Scenario | Trigger | Difficulty | Rationale |
|---|---|---|---|
| Reactive Kill | The pre-registered kill threshold fires; the data tells the team to stop. | Easy. | The rule forces the decision; the team stops and moves on. |
| Strategic Reallocation | The active experiment is on track and nothing is failing, but stronger mid-quarter evidence emerges for a different, higher-ICE idea. | Hardest call in capacity management. | The data is not forcing the decision, so the team must voluntarily abandon running work. |
The sunk-cost logic for strategic reallocation: the weeks already spent are gone whether the experiment finishes or stops; the only real question is whether the next 3 weeks go to the diminishing experiment or the better one just discovered. The practical mechanic that makes this possible is a decision review every 2 weeks, not at the end of the quarter.
Capacity Failure Modes
Teams often waste up to 60% of their capacity due to four allocation failures:
| Failure Mode | Team Behavior | Root Cause | Cure |
|---|---|---|---|
| Everything is a Priority | Refusing the 70/20/10 split, running 5 to 10 ideas at once. | Lack of operational focus and prioritization discipline. | Enforce strict 70/20/10 ratio; filter backlog by binding currency. |
| Hidden Capacity Tax | Forgetting the overhead of design, analytics, write-ups, and meetings. | Underestimating non-engineering operational tasks. | Budget full cycle times; include analytical overhead in ease scoring. |
| Founder Bottleneck | Experiments stall waiting for leadership approval or review. | Ignoring decision currency constraints. | Formally budget leader review hours; schedule standing biweekly reviews. |
| No Slack Budget | Committing 100% of available engineering or audience capacity. | Assuming zero friction or overruns. | Keep 25% of binding currency uncommitted to absorb surprises. |
The discipline check to kill active work during strategic reallocation asks: "Would the team start this experiment today knowing what is currently known?". If the answer is no, the active test must be killed to redirect remaining resources to higher-leverage ideas.
90 Day Growth Roadmap
A backlog and a roadmap are distinct artifacts serving different operational purposes:
Backlog vs. Roadmap
The operational differences between the backlog and the roadmap are structured as follows:
| Attribute | Growth Backlog | 90-Day Growth Roadmap |
|---|---|---|
| Horizon | Stable, lives forever. | Quarterly, resets every 12 weeks. |
| Focus | Continuous collection of new ideas. | Weekly calendar of execution. |
| Sort Order | Sorted by ICE score. | Sorted by week. |
| Operational Job | Determines what is worth doing. | Determines what ships when. |
| Variables Tracked | Individual inputs (Impact, Confidence, Ease). | Pre-committed decision dates, dependencies, and slack budget. |
The backlog feeds the roadmap, and the roadmap drives the execution.
Quarterly Three-Month Logic
A standard execution quarter consists of 12 weeks, with each month having a distinct role:
| Execution Month | Primary Focus | Key Deliverable | Operational Standard |
|---|---|---|---|
| Month 1 | Validate Inputs. | Clean Infrastructure. | Tech-heavy setup: building experiments, confirming instrumentation works, verifying cohort segmentation. |
| Month 2 | Run Scaled Experiments. | Clean Data Signal. | Full scale execution of the 70% capacity allocated to known winners. |
| Month 3 | Decide and Reallocate. | Compounding Learnings. | Applying pre-registered decision rules to scale, kill, or iterate; pivoting uncommitted slack to next bets. |
Treating all 12 weeks of a quarter identically results in execution without learning.
Five Columns of a Growth Roadmap
Every quarterly roadmap contains 12 rows (one per week) and five columns to structure execution:
| Column | Data Tracked | Core Purpose |
|---|---|---|
| 1. Week | Numbers 1 through 12. | The vertical anchor and chronological spine of the roadmap. |
| 2. Active Experiments | Experiment names and status. | Displays what is running, tracking its lifecycle stage (building, active, deciding, closed). |
| 3. Capacity Committed | Volume of binding currency consumed. | Engineering weeks or subscriber count consumed per row; empty cells show visible slack. |
| 4. Decision Dates | Pre-committed review calendar dates. | Locked review meetings scheduled before experiments go active. |
| 5. Dependencies | Blocked experiments and waiting states. | Highlights critical path blockers using "waiting on Experiment X" syntax. |
Five columns across 12 rows produce roughly 60 cells that turn the experimentation system into a structured calendar. The roadmap is read in two directions: top to bottom it is a calendar (week by week, what is happening across the team), and left to right it is a lifecycle (where each experiment sits in its arc). Both views matter, because a problem with the plan usually shows up in only one of the two views, so both are checked every Monday.
Clairo Quarter 4 Roadmap Case
A case study of a 12-week roadmap built under an engineering capacity constraint is structured below:
- Week 1 to 3 (E1 Building): Experiment 1 (onboarding redesign) is in building phase, consuming 1 engineer week per row. Parallel founder-led pricing test launches in Week 1, consuming 0 engineering hours.
- Week 4 (E1 Active / E3 Building): Experiment 1 goes active on a new user cohort. The engineer is freed to start building Experiment 3 (email timing test). Parallel runs are fine because they do not share a cohort.
- Week 5 (E1 Deciding / E3 Active): Experiment 1 hits its sample bound. The pre-committed decision date ("E1 Decision Friday") fires.
- Week 6 to 7 (E3 Active): Experiment 3 is active, and its pre-committed decision date lands at the end of Week 7.
- Dependencies (E2 Blocked): Experiment 2 (team invite mechanic) is blocked by Experiment 1. The dependencies column for Weeks 1 to 5 reads "E2 waiting on E1". E2 building begins in Week 7 after E1 has closed, going active in Week 10.
- Capacity Sizing: 8 weeks show engineering committed, leaving 4 weeks empty as a visible slack budget.
- Review Cadence: Alternate Monday biweekly reviews are highlighted as standing meetings.
Pre-Committed vs. Reactive Reviews
The timing of experiment reviews determines whether a team maintains velocity:
| Attribute | Reactive Reviews | Pre-Committed Reviews |
|---|---|---|
| Scheduling | Scheduled post-hoc when the team feels the data is ready or ambiguous. | Locked on the calendar before the experiment even goes active. |
| Dates Behavior | Dates slip based on team busywork or data questions. | Dates do not move, ensuring strict operational discipline. |
| Under Incomplete Data | Open-ended execution; experiments run indefinitely. | Team forces the decision (scale, kill, iterate) using whatever data exists. |
| Outcome | Capacity remains locked to dead experiments; quarter ends with unfinished tests. | Slots are freed on schedule, permitting mid-quarter capacity reallocation. |
Roadmap Failure Modes
Growth roadmaps can collapse mid-quarter due to four common failure modes:
| Failure Mode | Symptoms | Root Cause | Cure |
|---|---|---|---|
| Over-Planned Month 1 | Weeks 1 to 4 are 100% committed with zero buffer. | No capacity budgeted for instrumentation bugs or cohort surprises. | Keep Month 1 lean; preserve empty slack cells in capacity tracking. |
| No Decision Dates | Experiments run open-ended; no reviews occur on schedule. | Relying on reactive review scheduling. | Pre-commit review dates on the calendar before activating tests. |
| Hidden Dependencies | Experiment 2 must be redesigned in Week 6 because it ran in parallel with blocker Experiment 1. | Failure to map blocking relationships during sequencing. | Walk the dependency graph before scheduling; schedule blockers first. |
| No Review Cadence | The plan written in Week 1 remains identical in Week 11 despite changes. | Lacking a standing biweekly strategic meeting. | Establish a biweekly alternate Monday review to ask "are we working on the right things". |
Integrating the Full Growth Engine
An integrated growth engine connects individual experimentation disciplines across the full AARRR funnel, compounding knowledge across quarters.
AARRR Funnel: Checklist vs. System
Viewing the funnel as a system rather than a siloed checklist dictates company-level impact:
| Siloed Checklist View | Closed-Loop System View |
|---|---|
| Funnel is treated as five separate stages (Acquisition, Activation, Retention, Referral, Revenue). | Stages are connected in a self-reinforcing master loop. |
| Each stage is assigned an isolated tactic (e.g., Acquisition gets ads, Activation gets onboarding). | Feedback loops inform other stages (e.g., referral feeds acquisition). |
| Sub-teams optimize their stages independently, ignoring cross-stage feedback. | Cross-stage experiment design compounds data (e.g., activation tests designed on retention data). |
| Local Optimization: Stage metrics improve, but company-level growth remains unchanged. | Compounding Growth: Experiment outputs from one stage become inputs for the next. |
The classic Dropbox referral program is an example of a system loop: referral drove acquisition, referred users brought friends, friends were activated by sharing folders, and the sharing loop drove retention. Referral, retention, and acquisition were not separate programs but one single mechanism.
The Six-Discipline Stack Hierarchy
The growth operating system is a vertical stack of six disciplines integrated in a single workbook:
| Stack Layer (Top to Bottom) | Operational Role | Interdependent Connection |
|---|---|---|
| 6. 90-Day Roadmap | Calendar Layer. | Scheduled based on capacity limits. |
| 5. Capacity Allocation | Sizing Layer. | Sized based on validated metrics, preventing over-commitment. |
| 4. Validated Metrics | Truth Layer. | Verifies metric integrity inside decision rules. |
| 3. Decision Rules | Deciding Layer. | Converts backlog experiment results into contracts. |
| 2. Backlog System | Storage Layer. | Houses experiment lifecycles, ranked by ICE. |
| 1. ICE Prioritisation | Scoring Layer. | Evaluates ideas to feed the backlog. |
Removing any layer collapses the entire stack.
The six-discipline stack and the AARRR loop are the same workbook read in two different ways: AARRR tells the team what to work on (which funnel stage holds the binding constraint), and the stack tells the team how and when to work on it. Every discipline applies to every AARRR stage; if next quarter's constraint shifts from retention to revenue expansion, the stack does not change, only the ideas being scored, ruled, and scheduled do.
A Monday Morning Running the System
A concrete 30-minute walkthrough of a growth lead operating the full system at Clairo, at the start of Week 5:
| Time | Action | Discipline Used |
|---|---|---|
| 9:00 | Open the quarter tab; see it is the start of Week 5. | 90-Day Roadmap. |
| 9:05 | The calendar shows E1 hits its decision date on Friday; the capacity column shows engineering still committed to E1. | Roadmap, Capacity Allocation. |
| 9:10 | Open the decision scorecard: E1's pre-registered rule reads scale at 42%, kill at 36% or lower, iterate in between. | Decision Rules, Validated Metrics. |
| 9:15 | E1 currently sits at 41.3%, inside the iterate band; the Friday review meeting is already on the calendar, so the call is made then. | Pre-Committed Reviews. |
| 9:20 | It is also a biweekly review Monday: quick check confirms E2 build is on schedule and E3 is building with no surprises. | Review Cadence. |
| 9:25 | A new idea surfaces from sales (an enterprise tier add-on); it is scored on ICE and added to the backlog, where it stays because it ranks below the active items. | ICE Prioritisation, Backlog System. |
| 9:30 | Done. Every discipline used in 30 minutes. | The system runs the team. |
Annual rhythm of Growth System Compounding
Running the integrated stack continuously over four quarters creates a compounding growth flywheel:
| Chronological Phase | Core Operational Focus | Compound Impact / Deliverable |
|---|---|---|
| Quarter 1 | Instrumentation and Calibration. | Learnings focus on the system: identifying metrics, cohort segmentation, and verifying the binding currency. |
| Quarter 2 | Pattern Emergence. | Q1 learnings feed Q2 hypotheses; first known winners appear, and ICE scoring becomes accurate. |
| Quarter 3 | Flywheel Compounding. | Genuinely validated mechanisms are shipped at scale; 70% capacity is allocated to verified winners with high confidence. |
| Quarter 4 | Flywheel Maturity. | Referral loops are measurable and revenue expansion mechanisms are validated, feeding the system continuously. |
Teams that restart the system every quarter without a structured backlog fail to compound learnings.
Integration Failure Modes
Four integration failures can occur even when individual disciplines appear mature:
| Failure Mode | Symptoms | Core Root Cause | Structural Cure |
|---|---|---|---|
| Siloed Disciplines | The workbook has six tabs, but they are treated as separate processes; decision rules never reference metrics. | Individual ownership of tabs with zero real-time cross-referencing. | Conduct a workbook audit to verify that tabs actively reference and update each other. |
| Stage-Level Optimization | Activation team optimizes activation while retention flatlines. | Sub-teams focus only on local metrics without touching stage boundaries. | Design experiments that cross stages (e.g., onboarding changes measured against both activation and 90-day retention). |
| No Quarterly Learning Transfer | Each quarter starts fresh; Q2 plans have no reference to Q1 learnings. | Teams fail to review the closed learning column during roadmap resets. | Enforce a mandatory 30-minute review of all Q1 closed experiments as the first item of Q2 planning. |
| Write-Only Workbook | The workbook is updated weekly but never read by founders, sales, or product leads. | Growth updates are isolated from cross-department communication. | Establish transparent, quick-sync communication and alignment on growth workbooks across all teams. |
Four Phased Rollout Steps for Scratch Setup
To avoid building six broken disciplines at once, teams must roll out the growth engine in phases:
| Rollout Step | Timeframe | Phased Focus | Key Deliverable |
|---|---|---|---|
| Step 1 | Week 1. | Focus on one core metric and cohort. | Prevents vanity metrics upfront. |
| Step 2 | Weeks 2 & 3. | Focus on backlog and ICE scoring. | Basic 2-column spreadsheet to score and rank ideas. |
| Step 3 | Weeks 4 & 5. | Focus on decision rules and roadmap. | Pre-register rules for top 3 ideas and build the 12-week calendar. |
| Step 4 | Weeks 6 to 12. | Focus on capacity allocation and metric audits. | The full integrated stack is active, preparing the first compounding cycle for Q2. |
Ultra-Quick Revision (Exam Essentials)
Key Concepts & Distinctions
The core growth architecture concepts and their distinctions are contrasted below:
| Concept Pair | Core Distinction | Crucial Exam Takeaway |
|---|---|---|
| ICE Score vs. Growth Constraint | ICE score ranks individual ideas; growth constraint is the systemic bottleneck of the funnel. | An experiment with a high ICE score is useless if its impact does not target the primary constraint. |
| Build Effort vs. Time to Usable Signal | Build effort is the time required to code or ship; time to signal is the window required to collect statistically valid data. | Ease scores must reflect both variables: a 2-day build that needs 6 weeks of data collection is capped at a maximum ease score of 6. |
| Success/Fail vs. Bounded Iteration | Success/fail is a binary output; bounded iteration is a third path for marginal results with strict caps. | Excluding the iterate band forces premature scaling of weak ideas or premature killing of valid hypotheses. |
| Vanity Metric vs. Validated Metric | Vanity metrics are cumulative absolute counts that only rise; validated metrics are rates that can move in both directions. | Real metrics must pass 5 tests: denominator, comparability, two-directional, actionability, and business outcome. |
| Backlog vs. Roadmap | The backlog is a stable, prioritized list of ideas ranked by ICE; the roadmap is a 12-week execution calendar. | The backlog answers "what matters"; the roadmap answers "when it ships". |
| Pre-Committed vs. Reactive Reviews | Pre-committed reviews have locked dates set before launch; reactive reviews are scheduled when data feels ready. | Pre-committed reviews force decisions on schedule even when data is incomplete, keeping capacity unblocked. |
| Checklist Funnel vs. Closed-Loop System | Checklist treats AARRR stages as siloed, independent projects; system treats stages as a self-reinforcing loop. | System engines design activation experiments based on retention data, producing compound growth. |
Must-Know Terms
Definitions for core growth terms are structured below:
| Term | Exam Definition | Key Operational Requirement |
|---|---|---|
| ICE Score | The average of Impact, Confidence, and Ease scored on a 1 to 10 scale. | Used to sort backlog items in descending order. |
| WIP Limit | A work-in-progress constraint, typically set at 3 to 5 active experiments. | Prevents attribution noise, attention spread, and learning degradation. |
| Unidirectional Lifecycle | The rule that experiments move only forward (from Idea to Closed) and can never go backward. | Failed or successful experiments must close with documented learnings. |
| Pre-Registered Decision Rule | A contract written and locked before an experiment goes active, defining success thresholds. | Prevents post-hoc data manipulation and departmental bias. |
| Scale Threshold | The numeric metric value that triggers permanent product implementation. | Derived directly from the projected impact in the experiment hypothesis. |
| Kill Threshold | The numeric metric value below which the cost of maintaining the change exceeds its benefit. | Triggers an immediate halt to execution and forces learning documentation. |
| Iterate Band | The numeric range between the scale and kill thresholds where the signal is positive but weak. | Triggers Version 2 with exactly one variable changed and a new pre-registered rule. |
| Failed Kill | An experiment that is stopped without documenting learnings in the backlog. | Wastes resources; the team risks re-running the same failed experiment 6 months later. |
| Successful Kill | An experiment stopped with a detailed backlog summary of what failed and why. | Compounds organizational knowledge, ensuring future experiments avoid the dead mechanism. |
| Two-Hop Rule | The requirement that a metric must trace to revenue, retention, or margin in at most 2 steps. | Eliminates unvalidated marketing stories from decision rules. |
| Earned Leading Indicator | A predictive metric validated internally using own cohort data. | Must prove statistical correlation to a lagging business outcome before use in a rule. |
| Binding Currency | The capacity resource (Engineering, Decision, or Audience) that runs out first. | All quarterly roadmap planning must be sized against this constraint. |
| Slack Budget | A portion of binding capacity (20% to 35%) kept uncommitted. | Crucial for absorbing overruns, customer crises, and mid-quarter reallocations. |
| Blocking Experiment | An experiment whose outcome determines the hypothesis of a downstream test. | Must be scheduled to run serially first to preserve logical coherence. |
| Three-Month Logic | Slicing the 12-week roadmap into Month 1 (Validate), Month 2 (Run), and Month 3 (Decide/Reallocate). | Prevents treating all weeks identically, ensuring learnings are captured. |
| Six-Discipline Stack | The vertical operating system: ICE -> Backlog -> Rules -> Metrics -> Capacity -> Roadmap. | Layers are strictly interdependent; pulling one layer out collapses the stack. |
| Stage-Level Optimization | An integration failure where sub-teams optimize individual funnel stages in silos. | Improves local metrics but fails to move company-level growth. |
| Flywheel Compound | The long-term state where today's experiment outputs become tomorrow's inputs. | Achieved by running the integrated system continuously across four quarters. |