ExperimentLab
An A/B testing workbench that checks the data before it reads the result. Plan the sample, validate uploaded data and interpret the effect alongside its uncertainty, practical significance and guardrails.
3deterministic demo scenarios, including an invalid experiment


Make the evidence visible
A small p-value is not enough to justify shipping. ExperimentLab separates data integrity, statistical evidence, practical significance and guardrails. A broken allocation blocks inference instead of producing a confident recommendation from unreliable data.
Under the hood
- Health checks cover sample ratio mismatch, duplicate users, contamination and missing metrics. Severe integrity failures withhold treatment-effect conclusions.
- Binary outcomes use score tests and Newcombe intervals. Continuous outcomes use Welch tests, with CUPED shown beside the unadjusted analysis.
- Planning and simulation expose sample-size requirements, power and uncertainty. Exploratory segments use false-discovery correction and a heterogeneity test.
- An independent notebook re-derives the engine calculations. The supported design is a two-arm, individually randomised, fixed-horizon experiment.
Source is private. The screenshots use synthetic demo data. There is no hosted public demo.