Validity gates come before any lift number, not after. A 52/48 split on 200,000 users is a broken test rather than a close one, and a lift computed on a broken assignment is worse than no answer because someone will ship on it.
Steps
-
Discover the variant property before anything else.
posthog.list_properties,mixpanel.list_event_propertiesoramplitude.list_event_properties. The property is rarely calledvariant: look for$feature/<flag>,experiment_variant,ab_groupand similar per-account names. Alsoposthog.list_eventsto find the real exposure event and the real primary-metric event, andposthog.list_cohortsin case the team already defined the exposed population. Declare one primary metric now, in writing, before you have seen any result. -
Gate one, sample ratio mismatch.
posthog.query(HogQL) counting distinct exposed users per variant, ormixpanel.segmentationon the exposure event broken down by the variant property. Count exposures, not sessions. Run a chi-squared test of the observed counts against the intended allocation and use p below 0.001 as the threshold. If it fires, stop: report the broken split and the likely cause, and do not compute a lift. Also check that each distinct user maps to exactly one variant and report how many do not, since assignment that is not sticky across devices or logged-out and logged-in states puts the same person in both arms. -
Gate two, peeking, and gate three, power. Ask how many times the result was already looked at mid-flight, and state in the output that the nominal 95% is not 95% if the dashboard was watched daily. Then compute the minimum detectable effect from the observed per-arm sample and the control rate using
n = 2 * (z_alpha + z_power)^2 * p * (1 - p) / delta^2, which at 80% power and 5% two-sided alpha reduces to about15.7 * p * (1 - p) / delta^2per arm. Invert it and report the smallest lift this test could have detected. -
Gate four, novelty and primacy. Repeat the primary metric bucketed by day with
posthog.queryormixpanel.funnel_reportto produce the daily effect series, then check whether the effect is stable or trending. Novelty is initial intrigue that decays; primacy is initial resistance that fades as users adapt. Never read out a test that ran for fewer than a full weekly cycle, because weekday and weekend users differ. -
Compute the effect on the exposed population only.
posthog.query,mixpanel.funnel_reportoramplitude.funnel_reportfor the primary metric numerator and denominator per variant, restricted to exposed users, never all users and never sessions. Report absolute lift, relative lift as(p_treatment - p_control) / p_control, and a 95% confidence interval, so a flat result reads as "between minus 2% and plus 3%" rather than "not significant". Where the exposure log or the assignment lives in the application database, usepostgres.get_schemathenpostgres.query(orbigquery.query), which is also the only clean place to join assignment to billing identities. -
Check guardrails, then extend to money.
posthog.queryper guardrail metric per variant. If a support connection is available, add support contacts per variant cohort over the test window as a load guardrail. Thenstripe.list_subscriptionsplusstripe.list_invoicesjoined to the exposed user list for paid conversion and ARPU per arm, withstripe.list_refundsas a revenue guardrail. Report the join match rate as matched over attempted. This is where most proxy-metric wins evaporate, which is the point of doing it. -
Report. Table one: gate, result, threshold used, pass or fail, for sample ratio mismatch, peeking, power and minimum detectable effect, and novelty. Table two: metric, control n, control rate, treatment n, treatment rate, absolute lift, relative lift, 95% CI, primary or guardrail, verdict. Then one line: ship, kill, or extend with the additional sample required, framed against the roughly 10% median experiment success rate so a flat result is read as normal rather than as a failed measurement.
Gotchas
- A lift reported before the sample ratio check is the classic error. Run the chi-squared test first and refuse to report an effect if it fires. The imbalance itself is the finding.
- Peeking compounds and cannot be undone retroactively. Evan Miller's finding is a false positive rate near 26% when testing after every observation. If the team has already been watching daily, say the stated confidence is not what it claims, then either fix the remaining sample size or move to a sequential method.
- Analysing the wrong population manufactures lift. The denominator must be users exposed to the experiment. Mixing sessions and users across arms with different session counts produces a difference out of nothing.
- Many metrics means many chances to be wrong. Declare the single primary metric before looking; everything else is a guardrail or exploratory and must be labelled as such in the table.
- The proxy won and revenue did not. Clicks, activations and page views move far more easily than paid conversion. Always extend to the Stripe-side outcome before recommending a ship, and report the identity match rate, because a low match rate biases the revenue read toward whichever arm happened to have better identity coverage.
- Pick one analytics tool and stay in it. PostHog only is the best fit, since HogQL does the sample ratio counts, the daily series and the guardrails in a few queries. Amplitude only works through
amplitude.funnel_reportandamplitude.event_segmentationwith the variant property as the breakdown. Mixpanel only works throughmixpanel.segmentationandmixpanel.funnel_report, withmixpanel.export_eventsas the fallback when you need per-user assignment. Do not query two of them and merge the arms: they count users differently and will not reconcile. - GA4 is disqualified for experiment readouts.
google_analytics.run_reportapplies sampling and thresholding, condenses long-tail dimension values into the(other)row, and offers no user-level export, so neither the sample ratio check nor an exposed-population denominator can be computed. If GA4 is the only connection, say the test cannot be read out validly instead of producing a worse number.