The dishonest version of this readout auto-enrols everyone, divides adopters by all users, and reports launch week. The honest version computes an eligible denominator, splits breadth from depth from sustained use, and states plainly whether there was a control group at all.
Steps
-
Resolve the feature to real event names and check when tracking shipped.
posthog.list_events,mixpanel.list_eventsoramplitude.list_events, then have the requester confirm which events constitute the feature. A feature is one to several events, sometimes plus a property value, and the PM's name for it rarely matches the event name. Querymin(timestamp)on each: if the event first appears days after the release date, the early adoption curve is a tracking artefact, not slow uptake. Thenposthog.list_propertiesormixpanel.list_event_propertiesto find the plan, flag, platform and account-id properties the denominator needs. -
Compute the eligible denominator explicitly.
posthog.queryfor two numbers in one query: unique users or accounts who fired the feature events in the period, and unique active users or accounts in the same period filtered to eligible plan, flag cohort and platform. Callposthog.list_cohortsfirst: if a rollout or feature-flag cohort already exists, this readout becomes a holdout comparison instead of a time series estimate, which is a much stronger result. Report numerator and denominator, never only the ratio. -
Separate user-initiated from system-fired usage. If the feature is on by default, or an automated or backend event fires on the user's behalf, your adopter set includes people who never chose it. Check for suspiciously uniform inter-arrival times and near-perfect one-to-one ratios with another event using
posthog.query, and report user-initiated and system-fired adoption as two numbers. Auto-enrolling users and then reporting adoption is the single most common way this readout lies. -
Report breadth, depth and sustained use as three separate numbers. Breadth is the share of eligible users who used it at least once. Depth is median and p90 uses per adopting user, from
posthog.queryoramplitude.event_segmentationwith uniques per period. Sustained use is the share of first-week adopters still using it at day 30, 60 and 90, fromamplitude.retention_reporton the adopter cohort. Also plot weekly adopters from launch: a spike that decays to near zero within three weeks is announcement traffic. Judge the level after the decay, and report the decay rate. -
Locate the setup friction and the time to first use.
mixpanel.funnel_reportoramplitude.funnel_reportfor discovery to first use, with the conversion window set from the observed time-to-first-use distribution rather than the tool default, which otherwise truncates any feature whose real setup takes longer and manufactures a cliff. Report the distribution of first feature event minus per-user exposure date, plus the share of eligible users who never adopted. -
Run the impact read at the highest credibility the data supports, and name which one you used. In descending order: a feature-flag holdout, where
posthog.querycompares flagged against unflagged on the primary metric and PostHog retention can break down by flag directly; difference in differences against an untouched plan tier, region or platform, valid only if the pre-period trends are parallel; interrupted time series on the downstream outcome metric with segmented regression,y_t = b0 + b1*t + b2*D_t + b3*(t - T0)*D_t, whereb2is the level change andb3the slope change, using at least 8 periods either side; or matched adopter versus non-adopter comparison, labelled correlational. Close on money withstripe.list_subscriptionsplusstripe.list_invoicesjoined to the adopter list, andhubspot.search_companiesorpostgres.queryfor plan tier and the account-to-user mapping. Report the join match rate as matched over attempted. -
Report. One table: segment, eligible active users or accounts, adopters, breadth %, uses per adopter (median and p90), sustained use at day 30, time to first use (median), pre-period primary metric, post-period primary metric, estimated level change, estimated slope change, paid conversion for adopters versus non-adopters, and verdict. Then one sentence naming the impact method used, whether a control group existed, the concurrent events in the window that could explain the change, and a keep or sunset call with the ARR of the adopting segment attached.
Gotchas
- Launch date is not exposure date. Staged rollouts, app store review lag and cached clients make real exposure a distribution. Use the per-user first exposure timestamp, from the flag evaluation event or an app version property, as t equals zero, not the release note date.
- Week one is novelty or primacy. Novelty is initial intrigue that decays; primacy is initial resistance that fades as users adapt. Both make the first week unrepresentative. Report week 1 separately and base the verdict on weeks 2 to 4.
- A pre and post jump is not an effect. Without a control the estimate absorbs seasonality, a marketing push, a pricing change and every other release in the same window. List the known concurrent events in the output and call the number an estimate.
- Denominator drift flatters breadth. If eligible active users shrink while adopters stay flat, breadth rises for the wrong reason. Show both sides of the fraction in every period.
- Events are not users. One power user firing the event 400 times makes total-event adoption look healthy at near-zero breadth. Lead with unique users or accounts. Equally, account-level and user-level adoption answer different questions: a feature used by one person in a 50-seat account is not 2% adoption of that account's value, so compute and label both.
- Low adoption is not automatically failure. Compliance, procurement-checklist and once-a-year features are legitimately low-breadth. Check revenue concentration of the adopting segment before recommending a sunset, and recommend deprecation paths rather than deletion.
- Pick one tool and respect its limits. PostHog only is best for the causal read, because flags and retention live together. Amplitude only handles sustained use and adopter cohorts cleanly through
amplitude.retention_reportbut the time series must be assembled across calls. Mixpanel only gives breadth, depth and the setup funnel natively, with the eligible denominator as the weak point since plan must come from whatever is on the user profile. Do not query several analytics tools and merge: they count users differently and will not reconcile. GA4 is limited to page or screen-level usage and the marketing side of the launch, because the(other)row condenses less common dimension values, any dimension above 500 values is high cardinality, andgoogle_analytics.run_reportapplies thresholding, so a long-tail feature can vanish from the report entirely; usegoogle_analytics.run_reportfor changelog page traffic as an exposure proxy only.