Business Research and Growth Systems Architecture

Experimentation System and Integrated Growth Engine

Module 8

This module assembles the full experimentation operating system: ICE prioritisation, the eight-column backlog, pre-registered decision rules, vanity metric audits, capacity allocation under constraint, and the 90-day growth roadmap that integrates the whole engine.

Impact, Confidence, and Ease (ICE) Prioritisation

Experimentation System & Integrated Growth Engine: module overview infographic

Prioritisation is the fundamental problem for growth teams when ideas outnumber execution bandwidth. Growth planning typically starts with tens or hundreds of ideas, but teams can realistically ship only 5 to 7 experiments per month, or a maximum of 20 per quarter. Under 10 percent of executed experiments actually yield a positive impact on core business metrics. Without a structured framework, teams risk selecting incorrect ideas based on opinions rather than evidence. The Impact, Confidence, and Ease (ICE) framework solves this problem by using a transparent, repeatable score on a 1 to 10 scale to discipline executive judgment.

ICE Framework Components

Definitions and key questions for the three dimensions of the ICE framework are structured as follows:

DimensionCore DefinitionKey Scoring Question
ImpactThe estimated magnitude of movement that an experiment will produce on the core metric.How much will this experiment optimize the metric that is tied directly to the primary growth constraint?
ConfidenceThe level of certainty that the predicted outcome will occur as hypothesized.How sure is the team that the predicted outcome will happen, and what evidence supports this certainty?
EaseThe speed and effort required to ship, measure, and analyze the experiment.How fast can the team build, launch, and extract usable data from this experiment?

The overall ICE score is computed as the average of the three individual scores. Backlog items are sorted by their ICE scores in descending order to determine execution priority.

Impact Scoring Discipline

Impact scores must be anchored strictly to the primary growth constraint. A high score on a non-constraint metric does not count toward business growth, regardless of how much that metric improves.

Impact ScoreMovement on ConstraintMetric Tier AffectedCriteria and Guidelines
9 to 1020% or more shift.Primary Constraint.Shifts the primary constraint directly by a significant margin. Extremely rare.
6 to 8Meaningful shift.Secondary Metric.Creates a strong impact on a secondary metric or a moderate shift on the constraint.
3 to 5Modest shift.Supporting Metric.Moves a supporting metric with minimal impact on the primary constraint.
1 to 2Negligible shift.Vanity Metric.Moves a vanity metric only, which fails to alter business health or performance.

If an experiment does not touch the primary growth constraint, it cannot receive a high impact score. For example, if the primary constraint is activation, an A/B test of a pricing page designed to lift conversion from free trial to paid by 20% is downstream from activation. If only 34% of users activate, the conversion lift is applied to a thin slice of users. Thus, the pricing experiment should be scored only a 5 or 6 on impact, preserving scores of 9 for experiments that directly target the core activation rate.

Confidence Scoring Discipline

Confidence is the most commonly abused score in growth frameworks, frequently inflated by founders or team leaders to justify preferred ideas. Scoring must be bound to objective evidence.

Confidence ScoreQuality of EvidenceRequired Proof / Criteria
9 to 10Strong, direct evidence.Based on prior experiments conducted in the same product, on the same metric, or a near-identical pattern at a competitor cited by name or report.
6 to 8Indirect evidence.Competitor SaaS pattern with a comparable ICP, plus qualitative signals or interviews from own users indicating the same friction point.
3 to 5Logical hypothesis.Sound reasoning and logical hypothesis, but lacks direct product data or past results.
1 to 2Pure guess.Brainstormed ideas that feel right but lack supporting numbers, data, or competitive case studies.

The core discipline check requires that if the team cannot state the evidence for a confidence score in a single sentence, the score is capped at a maximum of 5. Statements like "it just feels right" or "everyone is doing this" automatically cap the score at a maximum of 4 or 5.

Ease Scoring Discipline

Ease counts the full cycle of building, measuring, and analyzing, not just the technical build time.

Ease ScoreTimeframe to Usable SignalTeam / Infrastructure Resources RequiredExamples
9 to 101 day or under.Shipped instantly, measurements already in place with zero custom engineering.Content change, ad copy change, email subject line test, simple price tweaks.
6 to 81 to 2 weeks.Light instrumentation, minor tool integrations.New onboarding screen, basic API tool integration, tracking a new product event.
3 to 53 to 6 weeks.Collaborative work across multiple departments (e.g., product, marketing, engineering), requires new metrics.Launching a new referral mechanic, full pricing page redesign, building a small feature.
1 to 22 months or more.Major engineering hours, custom infrastructure, platform architecture changes.Platfrom rebuild, deep database changes (should not be in an experimental backlog).

The core trap of ease scoring is ignoring the time required to collect data. An experiment that takes 2 days to build but requires 6 weeks to yield a statistically credible signal cannot be scored a 9. It must be scored a 6 at best, as time-to-usable-result is what dictates prioritization.

ICE Prioritization backlogs (Clairo vs. Zoko)

The ICE framework adapts to different product funnels by anchoring to brand-specific constraints.

ParameterClairo Case StudyZoko Case Study
Primary ConstraintActivation: 66% of free sign-ups never record a meeting.Activation: 62% of first-time starter kit buyers do not repurchase within 90 days.
Underlying FrictionFunnel drops off before users experience core value.Customers do not use products consistently enough to see visible results.
Top Experiment (ICE 8.0)Cut onboarding from 4 steps to 2 (Impact: 9, Confidence: 8, Ease: 7).Send Day 7, 14, and 21 WhatsApp photo prompts (Impact: 9, Confidence: 8, Ease: 7).
Supporting / High-Ease ExperimentDefault auto-record toggle to on (Impact: 8, Confidence: 7, Ease: 6, ICE: 7.0).Refer and Glow card inside unboxing kit (Impact: 7, Confidence: 7, Ease: 9, ICE: 7.7).
Low-Priority / Out-of-Constraint IdeasHomepage hero copy (Impact: 5, Confidence: 5, Ease: 9, ICE: 6.7); Slack integration (Impact: 5, Confidence: 5, Ease: 2, ICE: 4.0).Influencer myth-busting videos, TikTok creator partnerships, Amazon subscription toggle (ICE: 4.0 to 5.0).
Framework LogicHomepage copy is easiest but ranks fourth because it fails to target the activation constraint directly.Acquisition and revenue plays rank low because they fail to address the core activation leak.

Clairo's third-ranked idea, prebuilt meeting templates by type of sales call, also targets activation and scores an ICE of 7.3, which is why the homepage hero copy test (ICE 6.7) ranks only fourth despite being the easiest idea on the list. The framework protects the team from confusing speed with importance: the easiest experiment to ship is not automatically the most important one to run.

Failure Modes in ICE Execution

Teams often degrade the ICE framework into an opinion list through four primary execution errors:

PitfallSymptomsRoot CauseTeam Remedy
Confidence InflationEvery experiment score on confidence is written as a 9 or 10.Individuals seek to justify their personal ideas.Apply the single-sentence evidence constraint, capping unproven ideas at a maximum score of 5.
Scoring in IsolationBacklog scores drift based on individual interpretations.Different team members have varying definitions of what a score of 7 or 8 represents.Score the first three ideas as a team to calibrate scoring definitions.
Ease DiscrepancyHigh ease scores are assigned to complex engineering tasks.Leaders assume implementation is quick without consulting the engineering team.Score ease strictly from the perspective of the executing team who builds, measures, and analyzes.
Static ScoresBacklog rankings remain unchanged for months.Teams treat scores as permanent instead of dynamic.Rescore related backlog items whenever an experiment fails or new customer data emerges.

A spread diagnostic check determines team scoring honesty: a real backlog sheet contains a diversified range of ICE scores (some 8 to 9, some 4 to 6, some 2 to 3). If the entire backlog is clustered between 7 and 9, the framework has failed.

Growth Backlog Design

An experiment backlog is not just a list of ideas; a list is a repository where ideas accumulate and are eventually abandoned. A backlog is a system with defined rules, columns, and rituals that determines what runs, what stops, and what learnings are captured.

The typical list-to-graveyard trajectory follows a predictable timeline: in Month 1 the team opens a fresh "growth backlog" sheet and accumulates around 30 ideas within 2 weeks; by Month 3 the list holds 80 ideas but only 4 have been run and half the team has stopped opening the sheet; by Month 6 the sheet is abandoned entirely. If the backlog cannot answer "what runs, what stops, and what did we learn" within 30 seconds, it is a graveyard, not a system.

Lifecycle Stages of an Experiment

Lifecycle stages of a growth experiment

Every experiment in a functioning backlog must sit in one of five distinct stages:

StageDefinitionRequired Status Criteria
IdeaA written thought or proposal.Unscored and undiscussed.
ScoredICE inputs are fully completed.Average score calculated, ranked against the backlog, but not running.
ActiveCurrently running in production.Single owner assigned, decision rule pre-registered, and data collection active.
DecidingTarget sample size or time limit reached.Data collection complete, pending application of the decision rule to scale, kill, or iterate.
ClosedExperiment is finalized.Outcome captured, decision executed, and learnings documented.

Experiment movement is strictly unidirectional. An experiment cannot return to the idea stage; whether it fails or succeeds, it must move forward to the closed stage with captured learnings. If a row cannot be assigned to one of these five stages, it does not belong in the backlog.

Eight Columns of a Universal Backlog System

The universal backlog system relies on eight columns to turn a spreadsheet into an operating tool:

ColumnData TypeOperational Definition and Rule
1. ExperimentTextA single-line description of the specific test.
2. HypothesisTextStandardized format: If we do X, then Y will change because of Z, ensuring the hypothesis is falsifiable.
3. MetricNumberA single number that determines if the experiment succeeded.
4. ICE ScoreFormulaDynamic formula calculating the average of the impact, confidence, and ease inputs.
5. StageDropdownOne of the five defined lifecycle stages: Idea, Scored, Active, Deciding, or Closed.
6. OwnerTextExactly one named team member; multiple names or team labels are prohibited.
7. Decision RuleTextPre-registered thresholds (scale, kill, iterate) locked before the experiment goes active.
8. LearningTextMandatory single-sentence summary of results, cohorts affected, and findings; required to close the row.

These 8 columns work for B2B and B2C without modification: the structure is universal, and only the brand-specific thinking inside the cells changes.

Clairo Working Backlog Sheet (Template Walkthrough)

The reference implementation is a Google Sheets workbook with 12 experiment rows, 8 columns, and 2 tabs (the backlog plus a decision scorecard tab listing the top 5 experiments by ICE score with their pre-registered thresholds):

  • Live ICE formula: The ICE cell computes the average of the impact, confidence, and ease inputs. Changing one input (for example, confidence from 7 to 4) instantly updates the score and the rank order, so a disputed score is resolved by editing one cell rather than debating the total.
  • Conditional formatting: The top rows are highlighted automatically when the ICE score exceeds 7.5; highlights move as the team rescores.
  • Stage filter as the WIP view: Filtering the stage column to "active" shows only 4 rows (the WIP limit) while 18 scored experiments wait in line, so on Monday morning the team sees what is running, not the full idea pile.
  • Example pre-registered rule: "Scale if the sign-up to record rate hits 45% or above, kill below 36%, iterate in between", agreed before the experiment went active.
  • Example learning entry: "Reduced onboarding from 6 steps to 3. Activation rose from 34% to 41%. Effect held in week 2 and week 4 cohorts." Six months later, a proposed onboarding redesign builds on this row instead of re-running the test.

Built correctly, the decision-making becomes structural: the team does not have to remember to do the right thing, because the sheet's structure does the remembering.

WIP Limits and Rescoring Discipline

Work in Progress (WIP) limits prevent the backlog from becoming an unmanageable list. Successful growth teams establish a strict WIP limit of 3 to 5 active experiments at one time.

Exceeding 5 simultaneous active experiments introduces three operational complexities:

  1. Signal Noise Compounds: Multiple simultaneous experiments run on the same cohort, making clean attribution of metric movement impossible.
  2. Attention Spreads: Team focus degrades, leading to delayed implementation and missed details.
  3. Learning Rates Fall: The team executes more tests but lacks the bandwidth to extract deep, actionable insights from each.

Rescoring must be completed regularly as confidence is a dynamic variable. However, active rows must never be rescored midway through execution; experiments must run to their pre-registered decision rules to avoid bias.

Growth Rituals Cadence

The growth team ritual cadence

The backlog is driven by a weekly ritual consisting of two short meetings:

MeetingTimingTime CommitCore Agenda and Actions
Monday PlanningStart of week.45 to 60 minutes.1. Apply decision rules to completed experiments and move to closed.<br>2. Document learnings.<br>3. Rescore confidence on remaining ideas based on new evidence.<br>4. Select the next batch of experiments from the top of the ranked list respecting the WIP limit.<br>5. Assign one human owner per active row.
Friday Status CheckEnd of week.15 to 30 minutes.1. Quick check-in on experiment status.<br>2. Each owner provides a single-sentence status update.<br>3. Flag any experiment that has hit its decision threshold early.

No new commitments or experiments are allowed to be added midway through the week.

Backlog Design Pitfalls

Growth backlogs generally fail within the first eight weeks due to four common pitfalls:

PitfallSymptomsCore Root CauseStructural Remedy
The GraveyardBacklog grows to 80+ ideas, but only 4 are run; sheet is eventually abandoned.No constraint on backlog growth, and no prioritization enforcement.Enforce strict WIP limits of 3 to 5 active experiments at once.
No OwnerActive experiments are unassigned, or listed under a general department name.Lack of individual accountability for experiment status and measurement.A hard rule of exactly one human name per active row; no blank fields.
No RuleCompleted experiments are debated endlessly; success criteria are redefined after data arrives.Failure to lock down metrics and thresholds before launch.Pre-register and lock the decision rule before the experiment becomes active.
No LearningExperiments are closed but the learning column remains blank.Teams focus entirely on shipping rather than capturing knowledge.Block experiments in the deciding stage; they cannot be closed without a written learning statement.

Decision Rules

A decision rule is a contract written before an experiment runs that determines what its outcome means after the data is collected. Pre-registering decision rules prevents post-hoc negotiation of outcomes. Without a pre-registered rule, different departments will interpret the same data to support their personal preferences, resulting in indecision.

Experiment Outcomes Comparison

Every experiment outcome must map to one of three possible decisions:

DecisionDefinitionCore Threshold CriteriaOperational Action
ScaleFully validate hypothesis.Meets or exceeds the pre-registered success threshold.Implement the experiment permanently as a feature, campaign, or message.
KillInvalidate hypothesis.Drops below the pre-registered failure threshold.Stop the experiment immediately, do not iterate, and document learnings.
IterateUnclear or marginal signal.Lands within the band between the scale and kill thresholds.Run a revised version (V2) with exactly one variable modified under a new rule.

Excluding the iterate option forces premature closure of marginal experiments. Conversely, default iteration can become an excuse to avoid admitting failure. The rule must strictly define the iterate boundaries.

Four Components of a Decision Rule

Every pre-registered decision rule requires four locked components:

ComponentDefinitionAlignment / Derivation
1. Single MetricA single number that determines the success or failure of the test.Must align directly with the primary growth constraint. Multiple metrics are prohibited because they create conflicting signals.
2. Scale ThresholdThe numeric value that triggers a scale decision.Anchored directly to the projected impact in the experiment's hypothesis.
3. Kill ThresholdThe numeric value that triggers a kill decision.Usually a hurdle below which the cost (engineering, infrastructure, time) of running and maintaining the change exceeds its business benefit.
4. Time / Sample BoundThe limit that triggers the decision review.Anchored to statistical power calculations (fixed sample size) or a strict timeline (fixed window) set in advance.

Worked derivations of the two thresholds:

  • Scale threshold from the hypothesis: If the hypothesis projects activation rising from 34% to 45%, the scale threshold sits at or slightly below the projection (for example, 42% to 45%). If the threshold must be set much lower than the hypothesis projected, the hypothesis itself is wrong and should be rewritten before the experiment runs.
  • Kill threshold from cost-benefit: If lifting activation only to 36% would generate roughly 2 lakhs of revenue benefit while the experiment costs roughly 2 lakhs in engineering and infrastructure, any result at or below that point is a net loss, so 36% becomes the kill threshold.

Case Studies Decision Rules (Clairo vs. Zoko)

Applying pre-registered contracts across B2B and B2C brands yields distinct decision paths:

ParameterClairo Onboarding RedesignZoko WhatsApp Photo Prompt
Core HypothesisReducing onboarding from 6 steps to 3 will increase the activation rate.Sending photo prompts at Day 7, 14, and 21 will increase the 90-day repurchase rate.
Baseline34% activation rate (meeting recorded within 7 days).22% 90-day repurchase rate of first-time buyers.
Single Metric% of free users recording a meeting within 7 days.% of first-time buyers repurchasing in 90 days.
Scale Threshold42% or higher.30% or higher.
Kill Threshold36% or lower.24% or lower.
Iterate BandBetween 36% and 42%.Between 24% and 30%.
Bound500 users per arm or 14 days, whichever occurs first.800 customers enrolled in the cohort, waited full 90 days.
Actual Result40.3% activation rate among 14,547 users.23.1% repurchase rate among 812 customers.
Outcome DecisionIterate: Result falls directly within the iterate band.Kill: Result falls below the 24% kill threshold.

Constraints of Bounded Iteration

To prevent iteration from becoming an excuse for infinite unscientific testing, three constraints must be applied:

  1. Single Variable Change: Version 2 of the experiment must modify exactly one element of the first version. If multiple variables are changed, the team cannot attribute metric movement to a specific cause.
  2. New Pre-Registered Thresholds: Version 2 does not inherit the rules of Version 1. It must have its own thresholds and bounds written and locked before launch, and these are often stricter than the first version.
  3. Iteration Cap: A maximum of 2 iterations (up to Version 3) is permitted. If Version 3 still lands in the iterate band, the experiment must be killed because the underlying mechanism lacks sufficient strength.

Failed vs. Successful Kills

Killing an experiment is a positive outcome if it compounds team knowledge:

CriteriaFailed KillSuccessful Kill
Stage & Learning StateRow moved to closed, but the learning column is empty or reads "it did not work".Row moved to closed, and the learning column explicitly documents what was tested, what failed, and why.
Compounding KnowledgeCompounding learning is zero.Future experiments avoid the dead mechanism entirely.
Downstream Impact6 months later, someone proposes the same idea, and the team wastes resources re-running it.Compounding knowledge accumulates over quarters, increasing future experiment confidence.

An example of a successful kill learning for Zoko is: "Photo prompt mechanism at this cadence does not move the 90-day repurchase rate; future retention experiments should test referral incentives or ritual habit cards before retesting prompt cadence variations".

Decision Rule Pitfalls

Four common execution errors break the integrity of pre-registered decision rules:

PitfallDescriptionCore Root CauseRemedy
Moving the GoalpostEditing the thresholds after data is collected.Teams want to claim victory for marginal results.Enforce a strict pre-registration contract; no edits to rules are allowed after launch.
Cohort Cherry-PickingSlicing the data post-hoc until a statistically significant segment appears.Searching for positive signals in a failed test.Slice only cohorts pre-registered in the baseline rules; treat all other post-hoc segments as noise.
The More Data TrapExtending the run-time of a failed test hoping the numbers will improve.Refusal to accept a negative signal.Respect the pre-registered sample size or time bound; make the decision on schedule.
Iterate DefaultMarking clearly failed tests as "iterate" because the team is married to the idea.Emotional attachment to an experiment.Strictly follow the kill threshold; if a metric lands below the boundary, force a kill.

All four pitfalls are detectable with one question: did the rule exist in writing before the experiment went active? If yes, the rule wins and the decision follows it. If no, the team is negotiating outcomes after the fact, which is not experimentation.

Avoiding Vanity Metrics

A decision rule is only as effective as the metric it points to. If the metric is wrong, the rule fires backwards, optimizing numbers that fail to change business outcomes or profitability.

Vanity vs. Real Metrics Comparison

To identify metrics that represent actual business health, teams must distinguish between vanity and real indicators:

Metric TypeCore CharacteristicsBehavior in Decision Rules
Vanity Metric1. Can only go up by design (cumulative counts).<br>2. Moving it does not change team actions.<br>3. It is 2 or more hops away from a business outcome.Dangerous; they provide true numbers but point the team in the wrong direction, wasting resources on empty growth.
Real Metric1. Is expressed as a rate rather than an absolute count.<br>2. Can move up or down based on product health.<br>3. Traces to revenue, retention, or margin in 1 or 2 hops.Reliable; every change leads to a clear operational action and predicts business health.

Five Tests for Metric Integrity

Every candidate metric must pass five tests before it is allowed into a decision rule:

TestOperational RequirementTarget Standard
1. DenominatorIs it a rate and not just an absolute count?Rates are comparable across time and cohorts; counts are not.
2. ComparabilityCan it compare performance across cohorts and time periods?Must use a uniform computation across cohorts to compare apples to apples.
3. Two-DirectionalCan the metric go down?A metric that can only rise by construction cannot show risk or real product health.
4. ActionabilityIf this number changes, does the team know exactly what to do?The team must agree on which lever caused the movement and the next step to execute.
5. Business OutcomeDoes it trace directly to revenue, retention, or margin in 1 or 2 hops?More hops indicate a weak, unvalidated story rather than a true operational signal.

A candidate metric must pass all five tests; passing only 4 out of 5 is insufficient.

Metric Audits (Clairo and Zoko)

Auditing reporting scoreboards reveals common vanity metrics and their real metric replacements:

BrandVanity Metric (Failing Test)Real Metric ReplacementBusiness Outcome Traced
ClairoTotal free signups (8,400) (Fails Two-Directional; cumulative count).% activation by weekly signup cohort.Standardized rate; compares user conversion quality week over week.
ClairoTotal meetings recorded (Fails Business Outcome; recording does not equal value).% of users recording a meeting within 7 days.Strong leading indicator of long-term customer retention.
Clairo94% transcription accuracy (Fails Actionability; marketing claim, not operating metric).% of users with zero transcription errors per month.Operational metric that can be moved directly by engineering work.
ZokoInstagram followers (Fails Business Outcome; 3 hops away, unvalidated).Instagram attributed first purchase rate.Measures direct buyer conversion from the social channel.
ZokoTotal customers all time (6,200) (Fails Two-Directional; cumulative count, includes churned users).Active subscribers + new subscribers monthly.Fluctuates in real time, exposing subscriber churn and slowing acquisition.
ZokoAverage Order Value (AOV) (1,180) (Fails Business Outcome; ignores margin).Contribution margin per order.Protects company profitability; AOV can rise while actual margins drop.

The AOV trap illustrated with numbers: if Zoko's AOV rises from Rs. 1,180 to Rs. 1,400 while the margin drops from 40% to 30%, the AOV dashboard shows improvement even though profitability has fallen. Vanity metrics are not false numbers; they are true numbers that provide incomplete information and point the team in the wrong direction.

The Two-Hop Rule

The two-hop rule mandates that any metric in a decision rule must trace to revenue, retention, or margin in at most two steps.

Strong Metric Chain (Clairo):
[Activation Rate] ---> (Hop 1) [Day 7 Retention] ---> (Hop 2) [LTV & Revenue]
(Validated in cohort data: users who activate also retain, and users who retain pay).

Vanity Metric Chain (Zoko):
[Instagram Followers] ---> (Hop 1) [Reach] ---> (Hop 2) [Website Visits] ---> (Hop 3) [Purchases]
(Unvalidated: requires three hops based on assumptions rather than direct data correlation).

If a team cannot draw the correlation chain on a whiteboard in under 30 seconds, the metric does not belong in a decision rule.

The Leading Indicator Trap

A leading indicator is a metric that moves before the lagging business outcome does. For example, activation rate is a leading indicator for retention.

Leading indicators fall into two categories:

  1. Earned Leading Indicators: Validated internally using the company's own cohort data, proving a statistically significant correlation to a lagging metric.
  2. Claimed Leading Indicators: Imported from competitor blog posts or external content without internal validation. These act as vanity metrics in disguise.

Hypotheses belong in the experiment column; only validated, earned metrics belong in the decision rules.

Four Patterns of Vanity Metrics

Immature growth teams frequently fall into four common metric patterns:

PatternDefinitionOperational Flaw
1. World-Facing MetricsBig, cumulative numbers presented to investors, media, or the public.Cumulative signups or app downloads only increase, providing zero operational value.
2. Input as OutputMeasuring team effort or activities instead of results.Running 30 A/B tests or answering 100 tickets is an input; it does not mean the business improved.
3. Engagement MetricsMeasuring absolute site sessions, page views, or app opens.Measures general activity without verifying if it translates into user value or revenue.
4. Gameable ClaimsRelying on qualitative surveys like NPS or CSAT ratings.Highly sensitive to sample selection bias and survey design; easily manipulated.

Resource Allocation Under Constraint

Prioritized backlogs often have more high-scoring ideas than a team has the capacity to execute. Capacity acts as the second filter after the ICE framework. If teams attempt to run all high-scoring ideas in parallel without resource limits, signals collapse, team bandwidth is overstretched, and experiments fail to reach a clean conclusion.

Three Currencies of Capacity

Capacity is a combination of three distinct operational currencies:

CurrencyDefinitionMeasuring UnitBottleneck Risk
EngineeringAvailable technical and development hours from the tech team.Developer hours or engineering weeks.High in custom-built software products.
DecisionThe attention span and review time available from the CEO or department head.Unblocked approval hours and meeting availability.Ignored in tracking charts but causes delays if leadership becomes a bottleneck.
AudienceThe volume of unique user cohorts available to absorb experiments.Unique user traffic or cohort size.High in low-traffic or early-stage subscription products; running overlapping tests pollutes data.

The binding currency is whichever of these three runs out first, and teams must plan execution against it.

The 70/20/10 Allocation Rule

Binding capacity must be split across three distinct experiment categories:

CategorySizingTarget Confidence ScoreStrategic Role
Known Winners70% of capacity.ICE confidence of 7 or higher.Optimizing and iterating on already validated mechanisms; most likely to scale.
New Bets20% of capacity.ICE confidence of 5 to 7.Testing new hypotheses with reasonable evidence; moves into the 70% column next quarter.
Exploration10% of capacity.ICE confidence under 5.Brand-new mechanisms with a high risk of failure, but high option value if they succeed.

Memory hook: Capacity split 70/20/10: 70% known winners (confidence 7 or higher), 20% new bets (confidence 5 to 7), 10% exploration (confidence under 5). On top of that, keep a 20% to 35% slack budget of the binding currency uncommitted.

Capacity Constraint Audits (Clairo vs. Zoko)

Different brand structures hit different binding currencies, forcing distinct capacity layouts:

ParameterClairo Capacity PlanZoko Capacity Plan
Binding CurrencyEngineering weeks: 1 full-time engineer with 12 weeks of capacity.Audience size: 1,100 active subscribers.
Non-Binding CurrenciesAudience is highly plentiful; decision bottleneck is low.Shopify/WhatsApp APIs require zero engineering; founder is highly available.
Constraint CalculationImmature planning commits 5 top ICE ideas totaling 13 engineering weeks, immediately overrunning the budget.A retention experiment requires 500 users per arm (1,000 total), limiting tests to 2 parallel cohorts.
70% AllocationOnboarding redesign and email timing test (5 engineering weeks).Subscription onboarding experiment (660 subscribers).
20% AllocationTeam invite mechanic (3 engineering weeks).New referral mechanic (320 subscribers).
10% AllocationFounder-led sales outbound pricing test (0 engineering weeks, using sales calls).Creator partnership pilot (110 subscribers).
Slack Budget4 engineering weeks (one-third of capacity) kept uncommitted.110 subscribers left untouched and unexposed.

Slack Budget

A slack budget represents keeping 20% to 35% of the binding currency uncommitted. Slack is not waste; it is the operating buffer that allows plans to survive overruns, customer crises, or strategic reallocations when stronger mid-quarter evidence emerges.

Serial vs. Parallel Execution

Growth roadmaps must lay out experiments based on user interference:

Execution PathCore RuleInterference RiskOperational Example
ParallelRun simultaneously when tests touch different parts of the funnel or distinct cohorts.Low; cohorts do not overlap, multiplying team output.Pricing test on new visitors and onboarding test on day 7 users.
SerialRun sequentially when tests share a cohort or modify the same funnel step.High; simultaneous testing on the same cohort makes the signal uninterpretable.Running 2 onboarding flow tests on the same new user cohort.

A blocking experiment occurs when Experiment 2 depends on the outcome of Experiment 1. For example, upgrade prompt tests are blocked by pricing tests because the price determines the definition of an upgrade. Running them in parallel is logically incoherent. Blocker experiments must be scheduled to run first. The practical check for each active experiment: "Does the outcome of any other experiment in the backlog change what this experiment is testing?" If there is even a hint of yes, the other experiment is a blocker and runs first.

Killing Active Work: Reactive Kill vs. Strategic Reallocation

The hardest discipline in capacity management is stopping an experiment that is currently running. Two scenarios carry very different difficulty levels:

ScenarioTriggerDifficultyRationale
Reactive KillThe pre-registered kill threshold fires; the data tells the team to stop.Easy.The rule forces the decision; the team stops and moves on.
Strategic ReallocationThe active experiment is on track and nothing is failing, but stronger mid-quarter evidence emerges for a different, higher-ICE idea.Hardest call in capacity management.The data is not forcing the decision, so the team must voluntarily abandon running work.

The sunk-cost logic for strategic reallocation: the weeks already spent are gone whether the experiment finishes or stops; the only real question is whether the next 3 weeks go to the diminishing experiment or the better one just discovered. The practical mechanic that makes this possible is a decision review every 2 weeks, not at the end of the quarter.

Capacity Failure Modes

Teams often waste up to 60% of their capacity due to four allocation failures:

Failure ModeTeam BehaviorRoot CauseCure
Everything is a PriorityRefusing the 70/20/10 split, running 5 to 10 ideas at once.Lack of operational focus and prioritization discipline.Enforce strict 70/20/10 ratio; filter backlog by binding currency.
Hidden Capacity TaxForgetting the overhead of design, analytics, write-ups, and meetings.Underestimating non-engineering operational tasks.Budget full cycle times; include analytical overhead in ease scoring.
Founder BottleneckExperiments stall waiting for leadership approval or review.Ignoring decision currency constraints.Formally budget leader review hours; schedule standing biweekly reviews.
No Slack BudgetCommitting 100% of available engineering or audience capacity.Assuming zero friction or overruns.Keep 25% of binding currency uncommitted to absorb surprises.

The discipline check to kill active work during strategic reallocation asks: "Would the team start this experiment today knowing what is currently known?". If the answer is no, the active test must be killed to redirect remaining resources to higher-leverage ideas.

90 Day Growth Roadmap

A backlog and a roadmap are distinct artifacts serving different operational purposes:

Backlog vs. Roadmap

The operational differences between the backlog and the roadmap are structured as follows:

AttributeGrowth Backlog90-Day Growth Roadmap
HorizonStable, lives forever.Quarterly, resets every 12 weeks.
FocusContinuous collection of new ideas.Weekly calendar of execution.
Sort OrderSorted by ICE score.Sorted by week.
Operational JobDetermines what is worth doing.Determines what ships when.
Variables TrackedIndividual inputs (Impact, Confidence, Ease).Pre-committed decision dates, dependencies, and slack budget.

The backlog feeds the roadmap, and the roadmap drives the execution.

Quarterly Three-Month Logic

A standard execution quarter consists of 12 weeks, with each month having a distinct role:

Execution MonthPrimary FocusKey DeliverableOperational Standard
Month 1Validate Inputs.Clean Infrastructure.Tech-heavy setup: building experiments, confirming instrumentation works, verifying cohort segmentation.
Month 2Run Scaled Experiments.Clean Data Signal.Full scale execution of the 70% capacity allocated to known winners.
Month 3Decide and Reallocate.Compounding Learnings.Applying pre-registered decision rules to scale, kill, or iterate; pivoting uncommitted slack to next bets.

Treating all 12 weeks of a quarter identically results in execution without learning.

Five Columns of a Growth Roadmap

Every quarterly roadmap contains 12 rows (one per week) and five columns to structure execution:

ColumnData TrackedCore Purpose
1. WeekNumbers 1 through 12.The vertical anchor and chronological spine of the roadmap.
2. Active ExperimentsExperiment names and status.Displays what is running, tracking its lifecycle stage (building, active, deciding, closed).
3. Capacity CommittedVolume of binding currency consumed.Engineering weeks or subscriber count consumed per row; empty cells show visible slack.
4. Decision DatesPre-committed review calendar dates.Locked review meetings scheduled before experiments go active.
5. DependenciesBlocked experiments and waiting states.Highlights critical path blockers using "waiting on Experiment X" syntax.

Five columns across 12 rows produce roughly 60 cells that turn the experimentation system into a structured calendar. The roadmap is read in two directions: top to bottom it is a calendar (week by week, what is happening across the team), and left to right it is a lifecycle (where each experiment sits in its arc). Both views matter, because a problem with the plan usually shows up in only one of the two views, so both are checked every Monday.

Clairo Quarter 4 Roadmap Case

A case study of a 12-week roadmap built under an engineering capacity constraint is structured below:

  • Week 1 to 3 (E1 Building): Experiment 1 (onboarding redesign) is in building phase, consuming 1 engineer week per row. Parallel founder-led pricing test launches in Week 1, consuming 0 engineering hours.
  • Week 4 (E1 Active / E3 Building): Experiment 1 goes active on a new user cohort. The engineer is freed to start building Experiment 3 (email timing test). Parallel runs are fine because they do not share a cohort.
  • Week 5 (E1 Deciding / E3 Active): Experiment 1 hits its sample bound. The pre-committed decision date ("E1 Decision Friday") fires.
  • Week 6 to 7 (E3 Active): Experiment 3 is active, and its pre-committed decision date lands at the end of Week 7.
  • Dependencies (E2 Blocked): Experiment 2 (team invite mechanic) is blocked by Experiment 1. The dependencies column for Weeks 1 to 5 reads "E2 waiting on E1". E2 building begins in Week 7 after E1 has closed, going active in Week 10.
  • Capacity Sizing: 8 weeks show engineering committed, leaving 4 weeks empty as a visible slack budget.
  • Review Cadence: Alternate Monday biweekly reviews are highlighted as standing meetings.

Pre-Committed vs. Reactive Reviews

The timing of experiment reviews determines whether a team maintains velocity:

AttributeReactive ReviewsPre-Committed Reviews
SchedulingScheduled post-hoc when the team feels the data is ready or ambiguous.Locked on the calendar before the experiment even goes active.
Dates BehaviorDates slip based on team busywork or data questions.Dates do not move, ensuring strict operational discipline.
Under Incomplete DataOpen-ended execution; experiments run indefinitely.Team forces the decision (scale, kill, iterate) using whatever data exists.
OutcomeCapacity remains locked to dead experiments; quarter ends with unfinished tests.Slots are freed on schedule, permitting mid-quarter capacity reallocation.

Roadmap Failure Modes

Growth roadmaps can collapse mid-quarter due to four common failure modes:

Failure ModeSymptomsRoot CauseCure
Over-Planned Month 1Weeks 1 to 4 are 100% committed with zero buffer.No capacity budgeted for instrumentation bugs or cohort surprises.Keep Month 1 lean; preserve empty slack cells in capacity tracking.
No Decision DatesExperiments run open-ended; no reviews occur on schedule.Relying on reactive review scheduling.Pre-commit review dates on the calendar before activating tests.
Hidden DependenciesExperiment 2 must be redesigned in Week 6 because it ran in parallel with blocker Experiment 1.Failure to map blocking relationships during sequencing.Walk the dependency graph before scheduling; schedule blockers first.
No Review CadenceThe plan written in Week 1 remains identical in Week 11 despite changes.Lacking a standing biweekly strategic meeting.Establish a biweekly alternate Monday review to ask "are we working on the right things".

Integrating the Full Growth Engine

An integrated growth engine connects individual experimentation disciplines across the full AARRR funnel, compounding knowledge across quarters.

AARRR Funnel: Checklist vs. System

Viewing the funnel as a system rather than a siloed checklist dictates company-level impact:

Siloed Checklist ViewClosed-Loop System View
Funnel is treated as five separate stages (Acquisition, Activation, Retention, Referral, Revenue).Stages are connected in a self-reinforcing master loop.
Each stage is assigned an isolated tactic (e.g., Acquisition gets ads, Activation gets onboarding).Feedback loops inform other stages (e.g., referral feeds acquisition).
Sub-teams optimize their stages independently, ignoring cross-stage feedback.Cross-stage experiment design compounds data (e.g., activation tests designed on retention data).
Local Optimization: Stage metrics improve, but company-level growth remains unchanged.Compounding Growth: Experiment outputs from one stage become inputs for the next.

The classic Dropbox referral program is an example of a system loop: referral drove acquisition, referred users brought friends, friends were activated by sharing folders, and the sharing loop drove retention. Referral, retention, and acquisition were not separate programs but one single mechanism.

The Six-Discipline Stack Hierarchy

The growth operating system is a vertical stack of six disciplines integrated in a single workbook:

Stack Layer (Top to Bottom)Operational RoleInterdependent Connection
6. 90-Day RoadmapCalendar Layer.Scheduled based on capacity limits.
5. Capacity AllocationSizing Layer.Sized based on validated metrics, preventing over-commitment.
4. Validated MetricsTruth Layer.Verifies metric integrity inside decision rules.
3. Decision RulesDeciding Layer.Converts backlog experiment results into contracts.
2. Backlog SystemStorage Layer.Houses experiment lifecycles, ranked by ICE.
1. ICE PrioritisationScoring Layer.Evaluates ideas to feed the backlog.

Removing any layer collapses the entire stack.

The six-discipline stack and the AARRR loop are the same workbook read in two different ways: AARRR tells the team what to work on (which funnel stage holds the binding constraint), and the stack tells the team how and when to work on it. Every discipline applies to every AARRR stage; if next quarter's constraint shifts from retention to revenue expansion, the stack does not change, only the ideas being scored, ruled, and scheduled do.

A Monday Morning Running the System

A concrete 30-minute walkthrough of a growth lead operating the full system at Clairo, at the start of Week 5:

TimeActionDiscipline Used
9:00Open the quarter tab; see it is the start of Week 5.90-Day Roadmap.
9:05The calendar shows E1 hits its decision date on Friday; the capacity column shows engineering still committed to E1.Roadmap, Capacity Allocation.
9:10Open the decision scorecard: E1's pre-registered rule reads scale at 42%, kill at 36% or lower, iterate in between.Decision Rules, Validated Metrics.
9:15E1 currently sits at 41.3%, inside the iterate band; the Friday review meeting is already on the calendar, so the call is made then.Pre-Committed Reviews.
9:20It is also a biweekly review Monday: quick check confirms E2 build is on schedule and E3 is building with no surprises.Review Cadence.
9:25A new idea surfaces from sales (an enterprise tier add-on); it is scored on ICE and added to the backlog, where it stays because it ranks below the active items.ICE Prioritisation, Backlog System.
9:30Done. Every discipline used in 30 minutes.The system runs the team.

Annual rhythm of Growth System Compounding

Running the integrated stack continuously over four quarters creates a compounding growth flywheel:

Chronological PhaseCore Operational FocusCompound Impact / Deliverable
Quarter 1Instrumentation and Calibration.Learnings focus on the system: identifying metrics, cohort segmentation, and verifying the binding currency.
Quarter 2Pattern Emergence.Q1 learnings feed Q2 hypotheses; first known winners appear, and ICE scoring becomes accurate.
Quarter 3Flywheel Compounding.Genuinely validated mechanisms are shipped at scale; 70% capacity is allocated to verified winners with high confidence.
Quarter 4Flywheel Maturity.Referral loops are measurable and revenue expansion mechanisms are validated, feeding the system continuously.

Teams that restart the system every quarter without a structured backlog fail to compound learnings.

Integration Failure Modes

Four integration failures can occur even when individual disciplines appear mature:

Failure ModeSymptomsCore Root CauseStructural Cure
Siloed DisciplinesThe workbook has six tabs, but they are treated as separate processes; decision rules never reference metrics.Individual ownership of tabs with zero real-time cross-referencing.Conduct a workbook audit to verify that tabs actively reference and update each other.
Stage-Level OptimizationActivation team optimizes activation while retention flatlines.Sub-teams focus only on local metrics without touching stage boundaries.Design experiments that cross stages (e.g., onboarding changes measured against both activation and 90-day retention).
No Quarterly Learning TransferEach quarter starts fresh; Q2 plans have no reference to Q1 learnings.Teams fail to review the closed learning column during roadmap resets.Enforce a mandatory 30-minute review of all Q1 closed experiments as the first item of Q2 planning.
Write-Only WorkbookThe workbook is updated weekly but never read by founders, sales, or product leads.Growth updates are isolated from cross-department communication.Establish transparent, quick-sync communication and alignment on growth workbooks across all teams.

Four Phased Rollout Steps for Scratch Setup

To avoid building six broken disciplines at once, teams must roll out the growth engine in phases:

Rollout StepTimeframePhased FocusKey Deliverable
Step 1Week 1.Focus on one core metric and cohort.Prevents vanity metrics upfront.
Step 2Weeks 2 & 3.Focus on backlog and ICE scoring.Basic 2-column spreadsheet to score and rank ideas.
Step 3Weeks 4 & 5.Focus on decision rules and roadmap.Pre-register rules for top 3 ideas and build the 12-week calendar.
Step 4Weeks 6 to 12.Focus on capacity allocation and metric audits.The full integrated stack is active, preparing the first compounding cycle for Q2.

Ultra-Quick Revision (Exam Essentials)

Key Concepts & Distinctions

The core growth architecture concepts and their distinctions are contrasted below:

Concept PairCore DistinctionCrucial Exam Takeaway
ICE Score vs. Growth ConstraintICE score ranks individual ideas; growth constraint is the systemic bottleneck of the funnel.An experiment with a high ICE score is useless if its impact does not target the primary constraint.
Build Effort vs. Time to Usable SignalBuild effort is the time required to code or ship; time to signal is the window required to collect statistically valid data.Ease scores must reflect both variables: a 2-day build that needs 6 weeks of data collection is capped at a maximum ease score of 6.
Success/Fail vs. Bounded IterationSuccess/fail is a binary output; bounded iteration is a third path for marginal results with strict caps.Excluding the iterate band forces premature scaling of weak ideas or premature killing of valid hypotheses.
Vanity Metric vs. Validated MetricVanity metrics are cumulative absolute counts that only rise; validated metrics are rates that can move in both directions.Real metrics must pass 5 tests: denominator, comparability, two-directional, actionability, and business outcome.
Backlog vs. RoadmapThe backlog is a stable, prioritized list of ideas ranked by ICE; the roadmap is a 12-week execution calendar.The backlog answers "what matters"; the roadmap answers "when it ships".
Pre-Committed vs. Reactive ReviewsPre-committed reviews have locked dates set before launch; reactive reviews are scheduled when data feels ready.Pre-committed reviews force decisions on schedule even when data is incomplete, keeping capacity unblocked.
Checklist Funnel vs. Closed-Loop SystemChecklist treats AARRR stages as siloed, independent projects; system treats stages as a self-reinforcing loop.System engines design activation experiments based on retention data, producing compound growth.

Must-Know Terms

Definitions for core growth terms are structured below:

TermExam DefinitionKey Operational Requirement
ICE ScoreThe average of Impact, Confidence, and Ease scored on a 1 to 10 scale.Used to sort backlog items in descending order.
WIP LimitA work-in-progress constraint, typically set at 3 to 5 active experiments.Prevents attribution noise, attention spread, and learning degradation.
Unidirectional LifecycleThe rule that experiments move only forward (from Idea to Closed) and can never go backward.Failed or successful experiments must close with documented learnings.
Pre-Registered Decision RuleA contract written and locked before an experiment goes active, defining success thresholds.Prevents post-hoc data manipulation and departmental bias.
Scale ThresholdThe numeric metric value that triggers permanent product implementation.Derived directly from the projected impact in the experiment hypothesis.
Kill ThresholdThe numeric metric value below which the cost of maintaining the change exceeds its benefit.Triggers an immediate halt to execution and forces learning documentation.
Iterate BandThe numeric range between the scale and kill thresholds where the signal is positive but weak.Triggers Version 2 with exactly one variable changed and a new pre-registered rule.
Failed KillAn experiment that is stopped without documenting learnings in the backlog.Wastes resources; the team risks re-running the same failed experiment 6 months later.
Successful KillAn experiment stopped with a detailed backlog summary of what failed and why.Compounds organizational knowledge, ensuring future experiments avoid the dead mechanism.
Two-Hop RuleThe requirement that a metric must trace to revenue, retention, or margin in at most 2 steps.Eliminates unvalidated marketing stories from decision rules.
Earned Leading IndicatorA predictive metric validated internally using own cohort data.Must prove statistical correlation to a lagging business outcome before use in a rule.
Binding CurrencyThe capacity resource (Engineering, Decision, or Audience) that runs out first.All quarterly roadmap planning must be sized against this constraint.
Slack BudgetA portion of binding capacity (20% to 35%) kept uncommitted.Crucial for absorbing overruns, customer crises, and mid-quarter reallocations.
Blocking ExperimentAn experiment whose outcome determines the hypothesis of a downstream test.Must be scheduled to run serially first to preserve logical coherence.
Three-Month LogicSlicing the 12-week roadmap into Month 1 (Validate), Month 2 (Run), and Month 3 (Decide/Reallocate).Prevents treating all weeks identically, ensuring learnings are captured.
Six-Discipline StackThe vertical operating system: ICE -> Backlog -> Rules -> Metrics -> Capacity -> Roadmap.Layers are strictly interdependent; pulling one layer out collapses the stack.
Stage-Level OptimizationAn integration failure where sub-teams optimize individual funnel stages in silos.Improves local metrics but fails to move company-level growth.
Flywheel CompoundThe long-term state where today's experiment outputs become tomorrow's inputs.Achieved by running the integrated system continuously across four quarters.