+2
Product Analytics

Read out an A/B test honestly

Is this A/B test result real, or statistically significant only by accident? Runs the validity gates before reporting anything (sample ratio mismatch, peeking, power and minimum detectable effect, novelty), then gives absolute and relative lift with a confidence interval, a ship or kill call, and the winning variant extended through to revenue instead of stopping at the proxy metric.

  • +3
    Read out the checkout experiment: is the lift real?
  • +3
    Was our A/B test even valid? Check the split and the sample size
  • +3
    The variant looks 8% better after four days. Can we ship it?

The playbook

Validity gates come before any lift number, not after. A 52/48 split on 200,000 users is a broken test rather than a close one, and a lift computed on a broken assignment is worse than no answer because someone will ship on it.

Steps

  1. Discover the variant property before anything else. posthog.list_properties, mixpanel.list_event_properties or amplitude.list_event_properties. The property is rarely called variant: look for $feature/<flag>, experiment_variant, ab_group and similar per-account names. Also posthog.list_events to find the real exposure event and the real primary-metric event, and posthog.list_cohorts in case the team already defined the exposed population. Declare one primary metric now, in writing, before you have seen any result.

  2. Gate one, sample ratio mismatch. posthog.query (HogQL) counting distinct exposed users per variant, or mixpanel.segmentation on the exposure event broken down by the variant property. Count exposures, not sessions. Run a chi-squared test of the observed counts against the intended allocation and use p below 0.001 as the threshold. If it fires, stop: report the broken split and the likely cause, and do not compute a lift. Also check that each distinct user maps to exactly one variant and report how many do not, since assignment that is not sticky across devices or logged-out and logged-in states puts the same person in both arms.

  3. Gate two, peeking, and gate three, power. Ask how many times the result was already looked at mid-flight, and state in the output that the nominal 95% is not 95% if the dashboard was watched daily. Then compute the minimum detectable effect from the observed per-arm sample and the control rate using n = 2 * (z_alpha + z_power)^2 * p * (1 - p) / delta^2, which at 80% power and 5% two-sided alpha reduces to about 15.7 * p * (1 - p) / delta^2 per arm. Invert it and report the smallest lift this test could have detected.

  4. Gate four, novelty and primacy. Repeat the primary metric bucketed by day with posthog.query or mixpanel.funnel_report to produce the daily effect series, then check whether the effect is stable or trending. Novelty is initial intrigue that decays; primacy is initial resistance that fades as users adapt. Never read out a test that ran for fewer than a full weekly cycle, because weekday and weekend users differ.

  5. Compute the effect on the exposed population only. posthog.query, mixpanel.funnel_report or amplitude.funnel_report for the primary metric numerator and denominator per variant, restricted to exposed users, never all users and never sessions. Report absolute lift, relative lift as (p_treatment - p_control) / p_control, and a 95% confidence interval, so a flat result reads as "between minus 2% and plus 3%" rather than "not significant". Where the exposure log or the assignment lives in the application database, use postgres.get_schema then postgres.query (or bigquery.query), which is also the only clean place to join assignment to billing identities.

  6. Check guardrails, then extend to money. posthog.query per guardrail metric per variant. If a support connection is available, add support contacts per variant cohort over the test window as a load guardrail. Then stripe.list_subscriptions plus stripe.list_invoices joined to the exposed user list for paid conversion and ARPU per arm, with stripe.list_refunds as a revenue guardrail. Report the join match rate as matched over attempted. This is where most proxy-metric wins evaporate, which is the point of doing it.

  7. Report. Table one: gate, result, threshold used, pass or fail, for sample ratio mismatch, peeking, power and minimum detectable effect, and novelty. Table two: metric, control n, control rate, treatment n, treatment rate, absolute lift, relative lift, 95% CI, primary or guardrail, verdict. Then one line: ship, kill, or extend with the additional sample required, framed against the roughly 10% median experiment success rate so a flat result is read as normal rather than as a failed measurement.

Gotchas

  • A lift reported before the sample ratio check is the classic error. Run the chi-squared test first and refuse to report an effect if it fires. The imbalance itself is the finding.
  • Peeking compounds and cannot be undone retroactively. Evan Miller's finding is a false positive rate near 26% when testing after every observation. If the team has already been watching daily, say the stated confidence is not what it claims, then either fix the remaining sample size or move to a sequential method.
  • Analysing the wrong population manufactures lift. The denominator must be users exposed to the experiment. Mixing sessions and users across arms with different session counts produces a difference out of nothing.
  • Many metrics means many chances to be wrong. Declare the single primary metric before looking; everything else is a guardrail or exploratory and must be labelled as such in the table.
  • The proxy won and revenue did not. Clicks, activations and page views move far more easily than paid conversion. Always extend to the Stripe-side outcome before recommending a ship, and report the identity match rate, because a low match rate biases the revenue read toward whichever arm happened to have better identity coverage.
  • Pick one analytics tool and stay in it. PostHog only is the best fit, since HogQL does the sample ratio counts, the daily series and the guardrails in a few queries. Amplitude only works through amplitude.funnel_report and amplitude.event_segmentation with the variant property as the breakdown. Mixpanel only works through mixpanel.segmentation and mixpanel.funnel_report, with mixpanel.export_events as the fallback when you need per-user assignment. Do not query two of them and merge the arms: they count users differently and will not reconcile.
  • GA4 is disqualified for experiment readouts. google_analytics.run_report applies sampling and thresholding, condenses long-tail dimension values into the (other) row, and offers no user-level export, so neither the sample ratio check nor an exposed-population denominator can be computed. If GA4 is the only connection, say the test cannot be read out validly instead of producing a worse number.

Sequel CLI

Install Sequel skills into your agent

One command connects your agent to Sequel and installs the Sequel skill, so it knows this playbook exists and reads it when a question matches. The CLI signs you in, provisions a scoped API key and writes the config for you.

Already have an MCP client?

https://api.sequel.sh/mcp

Point it at this URL and sign in when prompted, or send an API key from Settings as a Bearer token. Skills come with it; nothing else to install. Manual setup per client

  1. 1

    Install the Sequel CLI

    One line installs the latest CLI with whatever package manager you have.

    curl -fsSL https://sequel.sh/install | sh
  2. 2

    Sign in

    Authenticate in your browser and pick an organization.

    sequel login
  3. 3

    Install into your agent

    Writes the MCP config and installs the Sequel skill file for agents that support skills. Pick an agent from the list, or target one directly by its slug.

    sequel install
    • Claude Code
      sequel install claude-code
    • Claude
      sequel install claude
    • Cursor
      sequel install cursor
    • VS Code
      sequel install vscode
    • Windsurf
      sequel install windsurf
    • Zed
      sequel install zed
    • Codex
      sequel install codex
    • OpenClaw
      sequel install openclaw
    • Hermes
      sequel install hermes

More like this

Other Product Analytics skills

Can we trust our product metrics

Two tools show different numbers for the same metric and nobody knows which to trust. Reconciles the disputed figure against a system of record such as Stripe or the application database, then finds the instrumentation defect behind it: duplicate and near-duplicate event names, events that silently stopped firing, missing properties and system-fired events counted as user actions.

Feature launch impact readout

Measure whether a feature we just shipped actually landed: adoption against an eligible denominator of users who could even reach it, breadth, depth and sustained use after the release, a pre and post impact estimate with or without a control group, and a keep, iterate or sunset call.

Find our activation moment

Discover the aha moment: which early in-product action, taken in the first session or first week, most separates users who stay from users who vanish, how to score candidate first actions by retention lift, and what share of new signups ever reach it. Covers activation metric definition, activation rate and time to value.

Is our retention curve flattening

Do users keep coming back, or do they drift away after the first week or month? Builds the cohort retention curve for your key action at its natural frequency and judges whether it flattens to a plateau or decays to zero, covering N-day versus unbounded retention, repeat usage, leaky bucket diagnosis and how the plateau compares to published benchmarks.

Signup funnel drop-off analysis

Find the step where the most users abandon signup or onboarding, and how that varies by segment.

FAQ

Frequently asked questions

What is sample ratio mismatch and why check it first?
Sample ratio mismatch is a statistically significant difference between the allocation you configured and the allocation you observed. Test the observed counts against the intended split with a chi-squared test at a conservative threshold, commonly p below 0.001, to avoid false alarms. If it fires, the randomisation or the logging is broken and the lift is meaningless, so there is nothing to report until it is fixed.
Does peeking at A/B test results really matter?
Yes, badly. Evan Miller's analysis in How Not To Run an A/B Test shows that testing for significance after every observation pushes the false positive rate to about 26%, more than five times the intended 5%, and that ten peeks turn a nominal 1% into roughly 5%. The fix is to fix the sample size in advance and look once, or to use a sequential method built for repeated looks.
What should I do when the test comes back flat?
Report the minimum detectable effect the test could actually have found given the observed sample, not the effect you hoped for. For two proportions at 80% power and 5% two-sided alpha the per-arm sample is roughly 15.7 times p times (1 minus p) divided by the squared absolute lift, so invert it to say this test could only have detected a lift of X or larger. A flat result on an underpowered test is not evidence of no effect.
How often do A/B tests actually win?
Ronny Kohavi's published figures put the median experiment success rate at about 10%, with a range of roughly 8% to 33%: Booking.com around 10% and Microsoft around 33%, where about a third of ideas were positive and significant, a third flat, and a third negative and significant. Set that prior before interpreting: a flat result is the modal outcome, not a measurement failure.
Can an AI agent read out an experiment?
Yes. With PostHog, Mixpanel or Amplitude plus Stripe connected, the agent discovers the real variant property name, runs the sample ratio, peeking, power and novelty gates before touching the result, reports lift with a confidence interval, and extends the winning arm to paid conversion and ARPU. Sequel provides those connections over MCP and the experiment-readout playbook.

Put this playbook to work

Connect a source, ask the question, and the agent follows these steps. Free to start.