There are no published benchmarks for event-namespace health, so the right output is not a percentile. It is "these nine events are broken, these four are duplicates of each other, this metric is at 62% coverage against Stripe, and here are the three dashboards currently lying because of it."
Steps
-
Pull the full namespace, not a sample.
posthog.list_events,mixpanel.list_eventsoramplitude.list_eventsfor every event, thenposthog.list_properties,mixpanel.list_event_propertiesoramplitude.list_event_propertiesfor the property namespace. Event names are per-account and nothing in the product guarantees an event exists, so this list is the only authoritative input. Take all of it: the taxonomy-drift check is worthless on a sample. -
Find out which events are load-bearing.
posthog.list_insightsandposthog.list_cohorts, plusmixpanel.list_funnelswhere Mixpanel is connected. An event with volume that no saved insight uses is a deprecation candidate. An event a live dashboard depends on that has just gone quiet is a P0 incident. This step is what turns an alphabetical dump into a priority list, so do it before any querying. -
Inventory volume, lifespan and week-over-week deltas.
posthog.query(HogQL) for per-event 30-day volume, unique actors,min(timestamp)andmax(timestamp), then a second query for week-over-week volume change, excluding the current in-flight period. Classify each event active, dead or new. Flag any drop above 60% and any spike above 5x, then correlate the drop dates against release dates before calling anything breakage. Where HogQL is unavailable, usemixpanel.segmentationoramplitude.event_segmentationper event, cap to the top N by volume, and state that the long tail was not audited. -
Detect near-duplicates, renames and system-fired events. Normalise names to a key and cluster:
Button Clicked,button_clickedandButtonClickare one defect, and launch-driven sprawl is where they come from. A rename shows up as one event dying on the exact day another is born, so match actor overlap and property signature between a dying and an emerging event and report it as a rename rather than two defects. Then flag events with suspiciously uniform inter-arrival times or a near one-to-one ratio with another event as probably system or backend fired, and escalate loudly if any of them is being used as an activation or adoption metric. -
Measure property completeness per event, and identity health.
posthog.queryfor the null or missing rate of the properties that matter for segmentation (plan, account id, platform, version) on the top 30 events by volume, plus the share of events with no identified actor. Account id can be 100% present on 40 events and 0% on the one event your retention metric uses, which is exactly the failure a global completeness number hides. A high anonymous share systematically breaks lifecycle and new-user retention, since PostHog excludes anonymous events from lifecycle status. -
Reconcile the disputed metric against a system of record, not against another analytics tool.
postgres.get_schemathenpostgres.query(orbigquery.query) to count the same business entity in the application database: rows in signups, orders, documents. For money usestripe.list_payment_intentswithstatus: "succeeded"net ofstripe.list_refunds, orstripe.list_invoiceswithstatus: "paid". Convert both sides to an identical explicit UTC window at an identical grain, then report coverage as tool count divided by system-of-record count with both absolute counts shown. If GA4 is involved, include it only here, withgoogle_analytics.run_reportondimensions: ["date"], and attach the thresholding and(other)row caveats. -
Report. Table one: metric, source A value, source B value, system-of-record value, coverage %, most likely cause, evidence, fix owner. Table two: event name, normalised key, duplicate siblings, first seen, last seen, 28-day volume, week-over-week change, null rate on the key property, user-fired or system-fired, saved insights affected, verdict. Then one sentence naming the single defect that invalidates the most downstream analysis, and whether the client-side loss rate is measurable at all given the connections available.
Gotchas
- Comparing the same calendar date across tools is the most common false alarm. Each tool has its own independently configured time zone, so the same calendar range covers different hours. Convert both sides to an explicit UTC window before comparing anything, and remember that unique users, sessions and total events are three different quantities: state the grain in the output.
- Cross-tool deltas are structural, not necessarily bugs. PostHog, Mixpanel, Amplitude and GA4 resolve identity, sessions and uniqueness differently, so a 10% to 20% delta on unique users for the same event is normal. The job is to explain the delta, not to demand it be zero. Reporting "Mixpanel and GA4 disagree" as a defect destroys the audit's credibility. Reconcile each tool against the system of record instead, and never merge two analytics tools into one number.
- Volume drops that are real business events. A 70% fall in checkout starts after a pricing change is not a tracking bug. Correlate every cliff against release dates and against a warehouse ground-truth count before calling it breakage.
- Late-arriving events make the newest week a false cliff. Mobile SDKs batch and retry, so events can arrive days late, and GA4 is additionally incomplete for the most recent 24 to 48 hours. Exclude the in-flight period and end the window at least two days ago.
- Dead is not always dead. Annual renewal, tax-season and admin-only events legitimately fire a handful of times a year. Check 12 months of history before recommending anything, recommend deprecation rather than deletion, and never recommend removing an event a live insight depends on.
- The application database is a bad system of record unless filtered. Check for
deleted_at,is_test, internal email domains and seed rows first. Stripe's own objects also disagree with each other: a subscription that exists is not revenue collected, and an invoice markedpaidcan later be refunded or disputed, so say which object you treated as cash truth. - Identity resolution silently breaks every cross-source join, and client-side loss is invisible without one. A CRM email maps to a product analytics distinct id only if an identify call carried that email, and anonymous pre-signup activity under a device id will never join. Report matched over attempted for every join. If no database or payment system is connected, say plainly that the client-side loss rate is unmeasurable rather than implying the tool's number is complete. GA4 alone cannot support this audit at all: its event model, thresholding and
(other)row make most checks un-runnable, so degrade to events present versus expected plus a thresholding warning.