A/B Testing and Startup Experimentation Math
Module 2
This module builds the statistical core of growth experimentation: null hypothesis testing, sample size logic, Type I and Type II errors, and the practical directional and proxy-test adaptations startups use when traffic is limited.
The Foundation of Null Hypothesis Testing
Defining Hypotheses: Opinions vs. Scientific Tests
In growth experimentation, distinguishing between unstructured thoughts and structured hypotheses is fundamental to scientific validation.
The structural differences between hunches, hypotheses, and null hypotheses are compared below:
| Term | Definition | Testability | Business Example |
|---|---|---|---|
| Hunch | An opinion or gut feeling based on previous experience, lacking metrics, numbers, or timeframes. | Not testable. | Believing that a green pricing button will get more clicks. |
| Hypothesis (Alternative / H1) | A specific, measurable, and falsifiable prediction containing a defined metric and timeframe. | Testable. | Changing the CTA button to "record my first meeting" will increase signups by 10% within 30 days. |
| Null Hypothesis (H0) | The starting position assuming there is no effect, no difference, or that nothing will change. | Testable (with the objective to disprove or reject it). | Changing the CTA button will produce no difference in the signup rate. |
Core Principles of Null Hypothesis Testing
- The Assumption of No Effect: Experiments start by assuming that changes will produce no difference. This baseline protects growth teams from false optimism and confirmation bias.
- The Impossibility of Proving the Null: Growth teams can never prove a null hypothesis is true. They can only fail to reject it due to insufficient data or weak signals.
- Rejecting the Null vs. Proving the Alternative: When a test shows a statistically significant result, the null hypothesis is rejected. This supports the alternative hypothesis but does not prove it with absolute certainty due to the probabilistic nature of statistics.
- Falsifiability: To be valid, a hypothesis must be falsifiable. Vague statements cannot be falsified, whereas a quantified statement (such as a 10% relative increase within 30 days) is falsified if the metric is not reached.
Baseline Metrics in Case Studies: Clairo and Zoko
The baseline conversion scenarios for the two core brands are summarized below:
| Brand | Business Type | Target Customer | Baseline Metric | Tested CTA Variations |
|---|---|---|---|---|
| Clairo | B2B (AI Meeting Recorder) | Organizational Buyers, Decision-Makers | 3.7% free to paid conversion rate | Control: "Start Free" <br> Variant: "Record My First Meeting" |
| Zoko | B2C (Plant-Based Cosmetics) | Individual Consumers, Buyers | 38% starter kit attach rate | Control: "Start Your Skin Journey" <br> Variant: "Try the 21-Day Starter Kit" |
The starting position of no difference prevents confirmation bias. Confirmation bias occurs when marketers monitor results daily and stop tests early upon seeing a temporary spike (such as on day 3) without sufficient data.
Sample Size Logic in Marketing Experiments
Why Sample Size Matters
Sample size is the minimum number of observations needed before an experimental result can be trusted. Without a pre-calculated sample size, a team is merely making an observation rather than running a valid experiment. While flipping a coin 3 times and getting 3 heads tells us nothing, flipping it 1,000 times and getting 700 heads statistically demonstrates a non-random bias.
Three Essential Sizing Inputs
Every valid sample size calculation relies on three inputs that must be established before running the test:
- Baseline Conversion Rate: The historical performance of the control variant. B2B conversion baselines are typically lower than B2C baselines due to corporate decision-making structures.
- Minimum Detectable Effect (MDE): The smallest improvement relative to the baseline that would change the business decision.
- Significance Level and Power: Statistical parameters controlling risk.
These statistical parameters are defined below:
| Parameter | Statistical Term | Standard Default | Definition | Historical Origin |
|---|---|---|---|---|
| Significance Level | Alpha () | 0.05 (95% Confidence) | The probability that the test will declare a winner even when there is actually no real difference (false positive rate). | Proposed by statistician Ronald Fisher in 1925 as a pragmatic standard. |
| Power | 1 minus Beta () | 80% (20% Beta Risk) | The probability that the test will detect a real effect when one genuinely exists (false negative protection). | Established as an industry default by Jacob Cohen in 1969. |
Sensitivity of Sample Size to Alpha and Power
Tightening either statistical parameter inflates the required sample size, which is why the defaults should not be changed without a strong documented reason:
- Lowering Alpha: For Clairo's activation test (34% baseline, 10% relative MDE), moving from alpha 0.05 to alpha 0.01 raises the requirement from roughly 1,400 to roughly 2,100 users per variant. Critical industries such as pharma and finance accept this cost and run at 99% confidence.
- Raising Power: Increasing power from 80% to 90% raises the same requirement from roughly 1,400 to roughly 1,870 users per variant. Clinical trials commonly require 90% power, but 80% is the pragmatic marketing default.
- 80% Power in Plain Terms: If the variant truly works at the MDE size, the test detects it 4 out of 5 times. The remaining 1 in 5 cases the test misses the real effect and shows no significant result, which is the accepted cost of the 80% default.
Minimum Detectable Effect Sensitivity
The MDE size dramatically affects the required sample size:
- Tiny MDEs (such as a 1% absolute change) require sample sizes in the hundreds of thousands because very small variations are difficult to distinguish from noise.
- Practical MDEs (such as 15% to 20% relative improvements) reduce the required sample sizes to hundreds or low thousands, which are easily achievable for most startups.
- The Sizing Rule: Teams must set the MDE based on the minimum business improvement required to act, rather than setting a high MDE simply to make the test easier to run.
Sizing Realities and the 7-Day Rule
- Clairo Conversion Test: Testing Clairo's B2B baseline conversion of 3.7% with a 10% relative MDE at standard confidence and power requires approximately 9,500 visitors per variant. At 620 signups per month, this would take 15 months, making a full valid conversion test unrunnable at Clairo's scale.
- Zoko Attach Rate Test: Testing Zoko's B2C baseline conversion of 38% with an 8 percentage points absolute MDE requires approximately 560 visitors per variant.
- The 7-Day Rule: Zoko can reach its required sample size in less than 24 hours. However, the test must run for at least 7 calendar days to capture weekly visitor cycles, as weekday, weekend, morning, evening, and payday shoppers behave very differently.
Sizing Tools and Practical Calculations
- Evan Miller Sample Size Calculator: Outputs the raw sample size required per variation based on pure mathematical formulas before traffic or time are factored in.
- VWO Test Duration Calculator: Outputs the number of days required to reach the sample size given the average daily traffic level.
- Sizing Connection: Duration in days multiplied by daily visitors equals the total sample. Total sample divided by the number of variations equals the sample size per variation. For Clairo: 70 days at roughly 21 daily visitors gives about 1,470 total, or about 735 per variation. Higher traffic always means fewer days, and a lower MDE always means more days.
The two-tool walkthrough run live in the lesson produced these reference numbers:
| Data Point | Evan Miller Output (Sample per Variation) | VWO Output (Duration) |
|---|---|---|
| Clairo activation test (34% baseline, 10% relative MDE) | ~1,400 (lesson reference; the live tool run showed ~3,067) | ~70 days (~2.3 months) at ~21 daily signups (620 per month divided by 30) |
| Clairo with MDE raised to 20% relative | Drops to roughly one quarter (~370 in the lesson reference; ~770 on the live tool) | ~20 days |
| Zoko attach rate test (38% baseline, 8 percentage points absolute MDE) | ~560 | ~1 day at ~1,267 daily visitors (38,000 monthly), then held for the 7-day rule |
| Clairo direct conversion test (3.7% baseline, 10% relative MDE) | ~9,500 | ~15 months at 620 signups per month |
Raising the MDE makes the test cheaper but blinds it to smaller real effects: if Clairo's new onboarding flow only produces a 12% lift, a test sized for a 20% MDE will report no difference even though a real improvement exists.
The three critical errors that break sample size logic are outlined below:
- Peeking: Checking results before the required sample size is reached and stopping the test early. This inflates the false positive rate because random variations are mistaken for true signals.
- Micro MDE Sizing: Setting an MDE too low, which forces an unrealistically large sample size and extended timeframe.
- Miscalculating Variations: Forgetting that the calculator outputs the sample size per variation. Running an A/B test requires multiplying the output by two, and an A/B/C test requires multiplying by three.
Type I and Type II Errors in Marketing Decisions
The Two Failure Modes
A statistical experiment can fail in two distinct directions.
Type I and Type II errors are compared below:
| Detail | Type I Error | Type II Error |
|---|---|---|
| Common Name | False Positive. | False Negative. |
| Statistical Definition | Rejecting the null hypothesis when it is actually true (declaring a winner when there is no real difference). | Failing to reject the null hypothesis when it is actually false (missing a real improvement). |
| Probability Control | Controlled by the Significance Level (Alpha). | Controlled by Power (Beta). |
| Courtroom Analogy | Convicting an innocent defendant. | Acquitting a guilty defendant. |
| Business Cause | Peeking at early results, stopping tests prematurely, or ignoring weekly variations. | Running underpowered tests due to small sample sizes or setting the MDE too high. |
| Business Cost | Scaling a change that does not work, resulting in wasted money, engineering effort, and false beliefs about conversion drivers. | Missing a growth opportunity by killing a genuinely superior variant and keeping a worse baseline. |
Memory hook: Type I: "I declared a winner that was not real" (false positive, controlled by alpha, default 5%). Type II: "II missed a real winner" (false negative, controlled by power, default 80%, beta 20%).
The Decision-Reality Matrix
The 2x2 matrix maps decision outcomes against reality:
| Decision \ Actual Reality | Null Hypothesis (H0) is True (No Real Effect) | Null Hypothesis (H0) is False (Real Effect Exists) |
|---|---|---|
| Reject Null Hypothesis (H0) | Type I Error (False Positive, Probability = Alpha) | Correct Decision (True Positive, Probability = Power) |
| Fail to Reject Null Hypothesis (H0) | Correct Decision (True Negative) | Type II Error (False Negative, Probability = Beta) |
Case Studies: Business Cost of Errors
- Clairo Type I Error: The team stops a CTA test on day 3 because the variant shows a 4.2% conversion rate versus 4.0% for control. This 0.2% change is random daily variation. Clairo deploys the CTA on a false belief, and because the true conversion is lower, performance drops to 3%, costing the brand thousands of dollars.
- Clairo Type II Error: The team tests a new onboarding flow but limits the test to 310 users instead of the required 1,400. The test shows no significant result, leading the team to abandon a genuinely better onboarding flow.
- Zoko Type I Error: The team achieves its sample size in 18 hours on a weekday afternoon, observing a 9% lift. If they had run the test for 7 days, the weekday-specific lift would have normalized to a non-significant 2%.
- Zoko Type II Error: The team sets an MDE of 20% relative for an unboxing insert test. The real effect is a 9% improvement in the 90-day repurchase rate. Because the MDE was too high, the test shows no significant result, and the team discards a valuable retention opportunity.
Context-Based Risk Balancing
The severity of each error type depends on business conditions:
- Worry more about Type I errors when decisions require major resource commitments, high engineering costs, or are hard to reverse (such as changing core product pricing or initiating paid acquisition campaigns). It is also critical in slow-moving markets where mistake correction takes months.
- Worry more about Type II errors when testing opportunities are scarce due to low traffic, or in fast, competitive environments where missing a minor edge allows competitors to capture the market.
Startup Experimentation Math and Proxy Testing
The Rigor-Speed Trade-Off
Startups operate with limited traffic, time, and patience, requiring practical adaptations of ideal testing conditions. The goal is informed trade-offs, not shortcuts: the compromise in sample size or time must be explicit and chosen on purpose.
The spectrum runs from maximum speed to maximum accuracy. At the extreme speed end sits ship and watch: making a change and watching the aggregate metric with no control group, no hypothesis, and no statistical framework. Its risk is that the cause of any movement can never be isolated, which is why it is not treated as a legitimate experiment mode. Choosing a mode is a business decision driven by how big the impact of a wrong call would be, how reversible the decision is, and what the business can afford right now, not by statistics alone.
The three experiment modes available to growth teams are outlined below:
| Experiment Mode | Definition | Ideal Application | Risk Profile |
|---|---|---|---|
| Mode 01: Full Valid Test | Complete statistical control: pre-calculated sample size, run to completion, no peeking. | High-stakes, irreversible, or highly expensive business decisions (such as pricing updates). | Low risk: produces a statistically sound, defensive, and scalable conclusion. |
| Mode 02: Directional Test | Running a test with a defined hypothesis but stopping before reaching the full sample size. | Highly reversible, low-cost marketing changes (such as email headlines, secondary page copy). | High risk: accepts elevated Type I and Type II error rates to gain rapid direction. |
| Mode 03: Proxy Test | Running a full valid test on a rapid leading indicator rather than a slow lagging metric. | Slow-moving lagging metrics with small sample volumes (such as Clairo's conversion rate). | Moderate risk: depends on the strength and documentation of the proxy correlation. |
The required sample size reached dictates the reliability of the signal:
| Percentage of Sample Reached | Effective Power | Type I Error Rate | Statistical Validity |
|---|---|---|---|
| 20% of required sample | ~30% | 15% to 20% | Pure noise: results are completely invalid and should never justify scaling. |
| 40% of required sample | ~50% | 10% to 12% | Weak direction: provides early trend indications but remains highly uncertain. |
| 60% to 80% of required sample | Moderate | Slightly elevated | Emerging signal: reliability increases but remains statistically inconclusive. |
| 100% of required sample | 80% | 5% | Valid result: statistically sound and ready for business decisions. |
Leading vs. Lagging Metrics and the North Star
Global growth teams distinguish metrics by their position in the funnel. In Clairo's illustrative funnel, 1,000 weekly visitors convert at 30% into 300 free users, who convert at 10% into 30 paid users per week.
- North Star Metric: The single most important business metric, here the number of paid users. It cannot be moved directly.
- Leading Metrics: The upstream levers that move the North Star: the visitor volume, the visitor-to-free conversion rate, and the free-to-paid conversion rate. Improving any one of them lifts the paid user count as a lagging effect.
- Lagging Metric: The downstream outcome (paid users, revenue) that responds only after the leading metrics change and aggregates too slowly to test directly at low volume.
Because a leading metric such as free users carries roughly 10 times the volume of the paid user count, changes become visible there with far less traffic. This volume advantage is exactly what makes a leading indicator usable as a proxy.
Proxy Metrics: Validation Requirements
A proxy metric is a leading indicator that reliably predicts a lagging final metric but can be measured faster and with less data. A valid proxy must meet three strict criteria:
- Correlation: The proxy must genuinely predict the end metric, and this relationship must be documented (such as day 14 engagers showing 3 times higher repurchase rates at Zoko).
- Sensitivity: The proxy must respond directly to the specific marketing changes being tested.
- Speed Advantage: The proxy must be significantly faster to measure than the lagging metric. If both take similar time, the proxy offers no advantage.
Once traffic increases sufficiently, teams must replace the proxy and test the lagging final metric directly.
The proxy configurations for Clairo and Zoko are compared below:
| Brand | Lagging Final Metric | Baseline Rate | Target Sample (Lagging) | Selected Proxy Metric | Proxy Baseline | Proxy Target Sample |
|---|---|---|---|---|---|---|
| Clairo | Free to Paid Conversion Rate | 3.7% | 9,500 per variant (15 months) | 7-Day Activation Rate (reaching the "Aha!" moment) | 34% | 1,400 per variant (2.3 months) |
| Zoko | 90-Day Repurchase Rate | 22% | Multi-month observation window | Day 14 Engagement (WhatsApp replies or check-in email opens) | High volume | 480 per variant |
Proxies are selected cross-functionally, not by the growth team alone. Clairo's growth, product, sales, and customer success teams jointly identified that users who activate within the first 7 days are the ones who convert to paid, and Zoko validated its day 14 engagement correlation against its existing base of roughly 6,200 customers before relying on it. Clairo's proxy test runs as a Mode 01 full valid test on the proxy metric itself: the compromise is the switch of metric, not the statistical rigor.
Structuring the Four-Column Experiment Decision Log
Growth teams record all experiments systematically to maintain continuity and accountability:
- Column 1: Hypothesis and Metric: Documents H1, H0, and the specific metric to prevent vague descriptions.
- Column 2: Experiment Mode: Explicitly sets Mode 01, Mode 02, or Mode 03 before the test begins to define acceptable standards of proof.
- Column 3: Sample Target vs. Actual: Sanity check recording the target sample versus the actual achieved sample.
- Column 4: Decision and Risk Acknowledged: Logs the final scaling decision and the exact error risk accepted.
Practical A/B Case Walkthroughs and Results Analysis
Execution of the 6-Step Workflow
Both brands execute an identical 6-step workflow, resulting in different statistical outcomes.
- State alternative (H1) and null (H0) hypotheses before opening any testing tools.
- Calculate the required sample size using baseline rates, MDE, and statistical defaults.
- Select the experiment mode based on decision reversibility and scaling costs, and write the decision rule before any results arrive (for example: if P is below 0.05 at full sample, deploy; if not, do not deploy, retain the hypothesis, and iterate).
- Run the test to completion without peeking and enforce the 7-day rule.
- Read the final statistical results (sample sizes, conversion rates, P-value, and confidence).
- Log the final business decision and risks in the decision sheet.
Case Walkthrough A: Clairo's Onboarding Optimization
- Step 1: Hypothesis: H1 predicts that a new guided first meeting flow increases the 7-day activation rate from 34% to 37.4% or higher (a 10% relative lift). H0 predicts no difference.
- Step 2: Sizing: Baseline of 34%, MDE of 10% relative, alpha of 0.05, and power of 80% requires 1,400 users per variant. This is achievable in 2.3 months given 620 signups/month.
- Step 3: Mode: Mode 01 (Full Valid Test) is selected because activation is a critical quarterly product driver.
- Step 4: Run: The test runs for 70 days, reaching 1,447 users per variant (103% of required sample).
- Step 5: Results: Control activation is 34%, variant activation is 38.4% (absolute lift of 4.4%, relative lift of 12.9%). The P-value is 0.031 (less than 0.05 alpha), with a confidence level of 96.9%.
- Step 6: Decision: P is less than alpha at full sample size. Reject H0 and deploy the guided meeting flow across the platform.
Case Walkthrough B: Zoko's Starter Kit CTA
- Step 1: Hypothesis: H1 predicts that changing the CTA to "Try the 21-Day Starter Kit" increases the starter kit attach rate from 38% to 46% (an 8 percentage points absolute lift). H0 predicts no difference.
- Step 2: Sizing: Baseline of 38%, MDE of 8% absolute, alpha of 0.05, and power of 80% requires 560 users per variant. This is achievable in 1 day given Zoko's traffic, but scheduled for 7 days.
- Step 3: Mode: Mode 01 (Full Valid Test).
- Step 4: Run: The test runs for 7 days, reaching 572 users per variant (102% of required sample).
- Step 5: Results: Control attach rate is 38%, variant attach rate is 40.3% (absolute lift of 2.3%). The P-value is 0.21 (above 0.05 alpha), with a confidence level of 79%.
- Step 6: Decision: P is greater than alpha. Fail to reject H0. Do not deploy the CTA change.
Interpreting Non-Significant Outcomes
A non-significant result is not an experimental failure. Zoko's test provided a valid, complete conclusion: the tested CTA did not produce the targeted 8% absolute lift. It does not prove the CTA has zero effect (since a 2.3% lift occurred) or mean the team should stop testing. The correct next step is to retain the hypothesis, refine the copy or lower the MDE with a larger sample size, and run a new iteration.
Dynamic Experiment Logging and Continuous Lifecycles
The experiment decision log maintains a rigorous audit trail across cycles:
- In Zoko's first run (Row 2), the team lacked traffic, running a directional test (Mode 02) that only reached 220 users per variant (39% of sample). They logged it as inconclusive but retained the hypothesis.
- In Zoko's second run (Row 5), once traffic was sufficient, they ran a completed full valid test. Even though the results were non-significant, logging both runs prevented the team from making the mistake of scaling a premature directional signal.
Ultra-Quick Revision (Exam Essentials)
Key Concepts & Distinctions
| Distinction Area | Concept A | Concept B | Core Difference |
|---|---|---|---|
| Hypotheses | Null Hypothesis (H0) | Alternative Hypothesis (H1) | H0 assumes zero effect or no difference, acting as the scientific starting position to prevent bias. H1 is the specific, measurable, and falsifiable prediction of change. |
| Experimental Errors | Type I Error (False Positive) | Type II Error (False Negative) | Type I occurs when declaring a winner when no real difference exists (controlled by alpha). Type II occurs when failing to detect a real difference that does exist (controlled by power). |
| Testing Risks | Alpha () | Beta () | Alpha is the risk of a false positive (standard default is 5%). Beta is the risk of a false negative (standard default is 20%). |
| Sizing Tools | Evan Miller Calculator | VWO Duration Calculator | Evan Miller calculates the raw sample size per variant using pure formulas. VWO calculates the testing duration in calendar days based on daily traffic. |
| Metrics | Leading Metric (Proxy) | Lagging Metric (End Goal) | Leading metrics (such as activation rate) can be measured quickly and predict lagging metrics. Lagging metrics (such as paid conversion) are slow to aggregate but represent final business outcomes. |
| Mode | Required Sample Reached | Decision Reversibility | Acceptable Risk Level | Exam-Critical Application |
|---|---|---|---|---|
| Mode 01: Full Valid Test | 100% of required sample. | Low (costly or hard to reverse). | Low error risk (standard defaults). | Core pricing changes, major product updates. |
| Mode 02: Directional Test | Less than 100% (such as 40%). | High (cheap or easy to reverse). | Elevated error risk (accepts noise). | Low-impact copy tweaks, secondary CTAs. |
| Mode 03: Proxy Test | 100% of proxy-calculated sample. | Variable (depends on metric). | Moderate error risk (depends on correlation). | Low-volume primary metrics (such as measuring activation instead of conversion). |
Must-Know Terms
| Term | Exam-Focused Definition | Cause-Effect Relationship / Significance |
|---|---|---|
| Falsifiability | The capacity for a hypothesis to be proven wrong. | Essential for a hypothesis to be scientifically testable, requiring specific metrics and timeframes. |
| Confirmation Bias | The tendency to search for, favor, and interpret information that confirms pre-existing beliefs. | Causes growth teams to peek at early results and stop tests prematurely, creating false positives. |
| Minimum Detectable Effect (MDE) | The smallest relative or absolute lift in a metric that is business-critical enough to justify making a decision. | A smaller MDE increases the required sample size exponentially, while a larger MDE reduces the required sample size but makes the test blind to smaller lifts. |
| Power | The probability (standardized at 80%) that an experiment will correctly detect a real effect of the MDE size if one exists. | Higher power requires larger sample sizes. Low-power tests result in elevated Type II error rates. |
| The 7-Day Rule | The mandatory practice of running a test for at least 7 calendar days regardless of when the sample size is reached. | Mitigates weekly traffic variations (such as weekend versus weekday behaviors) to ensure a representative sample. |
| Aha! Moment (Activation) | The specific user milestone where they realize the core value of a product. | Serves as a high-volume, highly correlated proxy metric for B2B free to paid conversion. |
| Experiment Decision Log | A structured, four-column record documenting hypotheses, testing modes, sample ratios, and accepted risks. | Prevents organizational amnesia and stops teams from treating weak directional signals as verified scaling conclusions. |