ExperimentLab

Case study, source private
  • FastAPI
  • React
  • SciPy
  • statsmodels
  • DuckDB

An A/B testing workbench that checks the data before it reads the result. Plan the sample, validate uploaded data and interpret the effect alongside its uncertainty, practical significance and guardrails.

3deterministic demo scenarios, including an invalid experiment
ExperimentLab interface showing synthetic demo results
The Checkout Redesign demo uses synthetic data. Its interval still includes gains smaller than the meaningful threshold, which the interface makes explicit.
Synthetic Signup Flow experiment with results withheld after a sample ratio mismatch
The synthetic Signup Flow demo withholds effects after a failed allocation check.

Make the evidence visible

A small p-value is not enough to justify shipping. ExperimentLab separates data integrity, statistical evidence, practical significance and guardrails. A broken allocation blocks inference instead of producing a confident recommendation from unreliable data.

Under the hood

  1. Health checks cover sample ratio mismatch, duplicate users, contamination and missing metrics. Severe integrity failures withhold treatment-effect conclusions.
  2. Binary outcomes use score tests and Newcombe intervals. Continuous outcomes use Welch tests, with CUPED shown beside the unadjusted analysis.
  3. Planning and simulation expose sample-size requirements, power and uncertainty. Exploratory segments use false-discovery correction and a heterogeneity test.
  4. An independent notebook re-derives the engine calculations. The supported design is a two-arm, individually randomised, fixed-horizon experiment.

Source is private. The screenshots use synthetic demo data. There is no hosted public demo.