Module
Experimentation
- For
- Head of ecommerce / growth · also Lean founder-operator, Ecommerce marketer becoming an operator
- Stage
- Early tractionGrowingScaling
- Not for
- Stores without enough eligible traffic or a way to assign visitors consistently between variants
Published · v1.0.0
LockedLessons
2 lessons, in order.
- Lesson 1: Approve the test brief before you build the variant
Lock one hypothesis, the commercial threshold, feasibility inputs, guardrails, and decision rules before an experiment starts.
Locked✓ - Lesson 2: Read the test against the decision you planned
Check integrity, effect size, uncertainty, guardrails, and time patterns before choosing to ship, stop, iterate, or remain inconclusive.
Recommended first: Approve the test brief before you build the variant
Locked✓
An experiment is a decision system, not a dashboard that produces winners.
This module covers the two moments where teams most often create false confidence: before the test starts and after the output arrives. They belong together because the hypothesis and evaluation metrics need to be defined before the test starts.1
First, write a test brief that isolates one change and one commercial question. Name the eligible population, assignment unit, primary metric, guardrails, minimum effect worth acting on, feasibility assumptions, QA owner, and decision rules. Power, sample size, randomization, and evaluation criteria affect whether the test can answer the question.2 If the store cannot run the test cleanly or reach a useful answer in an acceptable time, the correct decision is to use another evidence route.
Second, read the result through that brief. Check assignment and instrumentation before effects. Compare the estimated effect and uncertainty with the commercial threshold. Review guardrails and time patterns. The ASA states that a p-value does not measure effect size or importance and does not provide a complete decision rule.3 An inconclusive result does not establish no effect.
Module diagnostic
Take the latest test result your team discussed. Could a new analyst recover the original hypothesis, planned metrics, exclusions, allocation, stopping method, integrity checks, and decision rule without asking the person who ran it? If not, the result is not yet ready to direct the roadmap.
Assignment
Choose one proposed storefront change. Complete the test brief before implementation begins. Ask the analyst to mark every missing input and assess feasibility. If approved, preserve the brief unchanged beside the final export. Then use the Atlas four-outcome workflow: ship, stop, iterate under a new hypothesis, or remain inconclusive.
The ecommerce lead owns the commercial decision. The analyst owns the statistical method and interpretation. The implementation owner confirms that both variants match the brief and can be rolled back.
Evidence and further reading
- "Patterns of Trustworthy Experimentation: Pre-Experiment Stage," Microsoft Research, https://www.microsoft.com/en-us/research/articles/patterns-of-trustworthy-experimentation-pre-experiment-stage/. Supports defining the hypothesis and evaluation metrics before an experiment. Accessed August 23, 2026.
- "Controlled experiments on the web: survey and practical guide," Data Mining and Knowledge Discovery, https://link.springer.com/article/10.1007/s10618-008-0114-1. Supports evaluation criteria, power, sample size, randomization, and experiment limitations. Accessed August 23, 2026.
- "Statement on Statistical Significance and P-Values," American Statistical Association, https://www.amstat.org/asa/files/pdfs/p-valuestatement.pdf. Supports the limits of p-values as measures of effect size, importance, hypothesis truth, or a complete decision rule. Accessed August 23, 2026.
Was this helpful?