+3
Product Analytics

Feature launch impact readout

Measure whether a feature we just shipped actually landed: adoption against an eligible denominator of users who could even reach it, breadth, depth and sustained use after the release, a pre and post impact estimate with or without a control group, and a keep, iterate or sunset call.

  • +4
    We shipped the Slack integration six weeks ago. Did anyone use it?
  • +4
    Did the new dashboard change retention, or is that just seasonality?
  • +4
    Which features are used by under 5% of paid accounts, and what ARR do those accounts represent?

The playbook

The dishonest version of this readout auto-enrols everyone, divides adopters by all users, and reports launch week. The honest version computes an eligible denominator, splits breadth from depth from sustained use, and states plainly whether there was a control group at all.

Steps

  1. Resolve the feature to real event names and check when tracking shipped. posthog.list_events, mixpanel.list_events or amplitude.list_events, then have the requester confirm which events constitute the feature. A feature is one to several events, sometimes plus a property value, and the PM's name for it rarely matches the event name. Query min(timestamp) on each: if the event first appears days after the release date, the early adoption curve is a tracking artefact, not slow uptake. Then posthog.list_properties or mixpanel.list_event_properties to find the plan, flag, platform and account-id properties the denominator needs.

  2. Compute the eligible denominator explicitly. posthog.query for two numbers in one query: unique users or accounts who fired the feature events in the period, and unique active users or accounts in the same period filtered to eligible plan, flag cohort and platform. Call posthog.list_cohorts first: if a rollout or feature-flag cohort already exists, this readout becomes a holdout comparison instead of a time series estimate, which is a much stronger result. Report numerator and denominator, never only the ratio.

  3. Separate user-initiated from system-fired usage. If the feature is on by default, or an automated or backend event fires on the user's behalf, your adopter set includes people who never chose it. Check for suspiciously uniform inter-arrival times and near-perfect one-to-one ratios with another event using posthog.query, and report user-initiated and system-fired adoption as two numbers. Auto-enrolling users and then reporting adoption is the single most common way this readout lies.

  4. Report breadth, depth and sustained use as three separate numbers. Breadth is the share of eligible users who used it at least once. Depth is median and p90 uses per adopting user, from posthog.query or amplitude.event_segmentation with uniques per period. Sustained use is the share of first-week adopters still using it at day 30, 60 and 90, from amplitude.retention_report on the adopter cohort. Also plot weekly adopters from launch: a spike that decays to near zero within three weeks is announcement traffic. Judge the level after the decay, and report the decay rate.

  5. Locate the setup friction and the time to first use. mixpanel.funnel_report or amplitude.funnel_report for discovery to first use, with the conversion window set from the observed time-to-first-use distribution rather than the tool default, which otherwise truncates any feature whose real setup takes longer and manufactures a cliff. Report the distribution of first feature event minus per-user exposure date, plus the share of eligible users who never adopted.

  6. Run the impact read at the highest credibility the data supports, and name which one you used. In descending order: a feature-flag holdout, where posthog.query compares flagged against unflagged on the primary metric and PostHog retention can break down by flag directly; difference in differences against an untouched plan tier, region or platform, valid only if the pre-period trends are parallel; interrupted time series on the downstream outcome metric with segmented regression, y_t = b0 + b1*t + b2*D_t + b3*(t - T0)*D_t, where b2 is the level change and b3 the slope change, using at least 8 periods either side; or matched adopter versus non-adopter comparison, labelled correlational. Close on money with stripe.list_subscriptions plus stripe.list_invoices joined to the adopter list, and hubspot.search_companies or postgres.query for plan tier and the account-to-user mapping. Report the join match rate as matched over attempted.

  7. Report. One table: segment, eligible active users or accounts, adopters, breadth %, uses per adopter (median and p90), sustained use at day 30, time to first use (median), pre-period primary metric, post-period primary metric, estimated level change, estimated slope change, paid conversion for adopters versus non-adopters, and verdict. Then one sentence naming the impact method used, whether a control group existed, the concurrent events in the window that could explain the change, and a keep or sunset call with the ARR of the adopting segment attached.

Gotchas

  • Launch date is not exposure date. Staged rollouts, app store review lag and cached clients make real exposure a distribution. Use the per-user first exposure timestamp, from the flag evaluation event or an app version property, as t equals zero, not the release note date.
  • Week one is novelty or primacy. Novelty is initial intrigue that decays; primacy is initial resistance that fades as users adapt. Both make the first week unrepresentative. Report week 1 separately and base the verdict on weeks 2 to 4.
  • A pre and post jump is not an effect. Without a control the estimate absorbs seasonality, a marketing push, a pricing change and every other release in the same window. List the known concurrent events in the output and call the number an estimate.
  • Denominator drift flatters breadth. If eligible active users shrink while adopters stay flat, breadth rises for the wrong reason. Show both sides of the fraction in every period.
  • Events are not users. One power user firing the event 400 times makes total-event adoption look healthy at near-zero breadth. Lead with unique users or accounts. Equally, account-level and user-level adoption answer different questions: a feature used by one person in a 50-seat account is not 2% adoption of that account's value, so compute and label both.
  • Low adoption is not automatically failure. Compliance, procurement-checklist and once-a-year features are legitimately low-breadth. Check revenue concentration of the adopting segment before recommending a sunset, and recommend deprecation paths rather than deletion.
  • Pick one tool and respect its limits. PostHog only is best for the causal read, because flags and retention live together. Amplitude only handles sustained use and adopter cohorts cleanly through amplitude.retention_report but the time series must be assembled across calls. Mixpanel only gives breadth, depth and the setup funnel natively, with the eligible denominator as the weak point since plan must come from whatever is on the user profile. Do not query several analytics tools and merge: they count users differently and will not reconcile. GA4 is limited to page or screen-level usage and the marketing side of the launch, because the (other) row condenses less common dimension values, any dimension above 500 values is high cardinality, and google_analytics.run_report applies thresholding, so a long-tail feature can vanish from the report entirely; use google_analytics.run_report for changelog page traffic as an exposure proxy only.

Sequel CLI

Install Sequel skills into your agent

One command connects your agent to Sequel and installs the Sequel skill, so it knows this playbook exists and reads it when a question matches. The CLI signs you in, provisions a scoped API key and writes the config for you.

Already have an MCP client?

https://api.sequel.sh/mcp

Point it at this URL and sign in when prompted, or send an API key from Settings as a Bearer token. Skills come with it; nothing else to install. Manual setup per client

  1. 1

    Install the Sequel CLI

    One line installs the latest CLI with whatever package manager you have.

    curl -fsSL https://sequel.sh/install | sh
  2. 2

    Sign in

    Authenticate in your browser and pick an organization.

    sequel login
  3. 3

    Install into your agent

    Writes the MCP config and installs the Sequel skill file for agents that support skills. Pick an agent from the list, or target one directly by its slug.

    sequel install
    • Claude Code
      sequel install claude-code
    • Claude
      sequel install claude
    • Cursor
      sequel install cursor
    • VS Code
      sequel install vscode
    • Windsurf
      sequel install windsurf
    • Zed
      sequel install zed
    • Codex
      sequel install codex
    • OpenClaw
      sequel install openclaw
    • Hermes
      sequel install hermes

More like this

Other Product Analytics skills

Can we trust our product metrics

Two tools show different numbers for the same metric and nobody knows which to trust. Reconciles the disputed figure against a system of record such as Stripe or the application database, then finds the instrumentation defect behind it: duplicate and near-duplicate event names, events that silently stopped firing, missing properties and system-fired events counted as user actions.

Find our activation moment

Discover the aha moment: which early in-product action, taken in the first session or first week, most separates users who stay from users who vanish, how to score candidate first actions by retention lift, and what share of new signups ever reach it. Covers activation metric definition, activation rate and time to value.

Is our retention curve flattening

Do users keep coming back, or do they drift away after the first week or month? Builds the cohort retention curve for your key action at its natural frequency and judges whether it flattens to a plateau or decays to zero, covering N-day versus unbounded retention, repeat usage, leaky bucket diagnosis and how the plateau compares to published benchmarks.

Read out an A/B test honestly

Is this A/B test result real, or statistically significant only by accident? Runs the validity gates before reporting anything (sample ratio mismatch, peeking, power and minimum detectable effect, novelty), then gives absolute and relative lift with a confidence interval, a ship or kill call, and the winning variant extended through to revenue instead of stopping at the proxy metric.

Signup funnel drop-off analysis

Find the step where the most users abandon signup or onboarding, and how that varies by segment.

FAQ

Frequently asked questions

How do you calculate feature adoption rate correctly?
Divide unique users or accounts who used the feature in the period by the unique active users or accounts who could have used it in that period. Could have means on a plan that includes it, inside the rolled-out flag cohort, on a supported platform, and active at all in the period. Dividing by all users understates a gated feature; dividing by eligible users while the flag sat at 10% rollout overstates it.
Is 20% feature adoption good or bad?
On its own it is neither. Pendo's 2019 Feature Adoption Report, built from 615 subscriptions with more than a year of data, found that about 80% of features in the average software product are rarely or never used, which makes a modest breadth number the normal case rather than a failure. The report is vendor research against customer-created Pendo tags, so use it as a prior, not a pass or fail bar.
How do I measure feature impact without a control group?
Use interrupted time series on the downstream outcome metric: model the pre-launch trend and compare the post-launch level and slope against that counterfactual, with roughly 8 to 12 periods either side of a clean intervention date. If an untouched comparable segment exists, difference in differences is stronger, but only if the pre-period trends are visibly parallel. Either way the number is an estimate, not a measurement, and the output must say so.
Why do adopters always look better than non-adopters?
Because your most engaged users adopt everything. Adopter versus non-adopter retention and revenue gaps are self-selection as much as feature value, so they overstate impact. Match on signup cohort, plan and pre-period activity level before comparing, label the result correlational, and prefer a flag holdout or the time series estimate for the headline.
Can an AI agent do a feature launch readout?
Yes. With PostHog, Mixpanel or Amplitude plus Stripe and a product database connected, the agent resolves the feature to real event names, computes the eligible denominator, reports breadth, depth and sustained use separately, and runs a holdout or time series impact read with the self-selection caveat attached. Sequel provides those connections over MCP and the feature-launch-impact playbook.

Put this playbook to work

Connect a source, ask the question, and the agent follows these steps. Free to start.