The activation metric is discovered, not declared. Every candidate event you score is a predictor of retention, not a cause of it, so the output of this analysis is a ranked hypothesis list that a holdout still has to confirm.
Steps
-
Enumerate the event namespace and check when each event started firing.
posthog.list_events,mixpanel.list_eventsoramplitude.list_eventsfor the full list. Then querymin(timestamp)andmax(timestamp)per event name withposthog.query(HogQL) ormixpanel.segmentation. Discard pageview-only and system-fired events, and discard any event whose first occurrence is after the cohort start date: its reach is mechanically near zero and its lift is garbage. Never assume an event exists because the product has the feature. -
Pick the retention yardstick before scoring anything. Choose one key action that best represents delivered value, and its natural frequency: weekly for most B2B SaaS, monthly for infrequent transactional products. Use
posthog.list_propertiesormixpanel.list_event_propertiesto find the signup-date property and the segmentation properties (plan, acquisition channel, platform, company size). Callposthog.list_cohortsto see whether the team already has a saved activation or power-user cohort, so the answer uses their definition rather than a new one. -
Build the retained versus churned split for one signup cohort. Take all users whose first signup event falls in one month. Label a user
retainedif they performed the key action during a fixed later window, for example week 4 through week 6 after signup, elsechurned. Both groups must have had the same elapsed time since signup. Build the cohort from the event stream at signup time, not from the current users table, or GDPR deletions and removed accounts bias the retained share upward. -
Score every candidate event by reach and lift inside a fixed early window. One
posthog.querydoes the whole table: a CTE for the cohort, a CTE for events wheretimestamp < signup_ts + INTERVAL N DAY, a CTE for the retained label, thenGROUP BY event. For each event computereach = doers / cohort sizeandlift = P(retained | did it) / P(retained | did not do it). Run N at 1, 7 and 14 days. A 6x lift at 2% reach is a power-user marker, not an activation metric; the candidate you want is the highest-lift event a meaningful share of the cohort can plausibly be driven to. -
Add the count dimension and find the knee. Re-score the top candidates at thresholds of at least 1, 2, 3 and 5 occurrences within the window. Activation metrics are almost always "X actions in Y days", and the right X is the point where additional occurrences stop adding lift. Report that knee explicitly.
-
Validate the winner and test for Simpson's paradox.
amplitude.retention_reportwith the candidate activation cohort versus everyone else: the two curves must visibly separate and the activated curve must flatten. Then re-run the lift holding plan, acquisition channel and platform fixed, and only promote a candidate whose sign is stable across all three. Optionally join tostripe.list_subscriptionson email or customer id to check whether the candidate also predicts paid conversion. -
Report. One table ranked by lift: candidate event, window in days, threshold count, reach %, retained share among doers, retained share among non-doers, lift, implied activation rate, and sign-stable across segments (yes or no). Then one sentence naming the recommended activation definition, the activation rate it implies for the latest cohort, and the explicit statement that this is a correlational finding to be confirmed with a feature flag holdout.
Gotchas
- Exposure bias makes downstream events look magical. A user can only fire
invited_teammateif they got through onboarding, so part of the lift is just "survived long enough to be able to do it". Report lift conditional on reaching the preceding funnel step, or at minimum flag which candidates sit downstream of each other. - Correlation is not causation, and this analysis cannot produce causation. High lift makes an event a good predictor, which is all an activation metric needs to be. It does not license "if we push users to do X, retention rises". State this in the output every time.
- Hold both windows to days since signup, never to calendar dates. Comparing "did X in the first 7 days" against a retained set measured over a variable calendar window silently rewards older users and invents lift.
- Account versus user activation in B2B. In seat-based products the activating actor is often an admin while the retained actor is an end user. If group analytics is configured, run the whole analysis at group level. If not, say the user-level answer may be an artefact of who holds the admin seat.
- Identity stitching splits the first day in two. Pre-signup anonymous events and post-signup identified events are different actors until the identify call fires, and PostHog excludes anonymous events from lifecycle status. A first-day window that straddles identification drops real activity.
- Connector degradation is real and you must declare it. PostHog only is the best case, since HogQL makes the lift table one query. Mixpanel only has no arbitrary SQL: approximate with
mixpanel.funnel_reportplacing the candidate as step 2 and the key action as step 3, ormixpanel.export_eventsfor the cohort and compute locally, warning that export is slow and volume-capped. Amplitude only cannot scan the whole namespace cheaply, so ask the PM to nominate 5 to 10 candidates and useamplitude.event_segmentationfor reach plusamplitude.retention_reportwith a behavioural cohort for lift. - Do not attempt this on GA4 alone.
google_analytics.run_reporthas no user-level event export and cannot express user-scoped behavioural cohorts. Degrade to signup-to-first-key-event conversion rate by channel, say explicitly that it is not an activation analysis, and do not compare the result to any activation benchmark.