+3
Product Analytics

Can we trust our product metrics

Two tools show different numbers for the same metric and nobody knows which to trust. Reconciles the disputed figure against a system of record such as Stripe or the application database, then finds the instrumentation defect behind it: duplicate and near-duplicate event names, events that silently stopped firing, missing properties and system-fired events counted as user actions.

  • +4
    Which of our events stopped firing last month, and which dashboards depend on them?
  • +4
    Do we have duplicate event names we should consolidate?
  • +4
    Our signup count in Mixpanel doesn't match the database. Where is the gap?

The playbook

There are no published benchmarks for event-namespace health, so the right output is not a percentile. It is "these nine events are broken, these four are duplicates of each other, this metric is at 62% coverage against Stripe, and here are the three dashboards currently lying because of it."

Steps

  1. Pull the full namespace, not a sample. posthog.list_events, mixpanel.list_events or amplitude.list_events for every event, then posthog.list_properties, mixpanel.list_event_properties or amplitude.list_event_properties for the property namespace. Event names are per-account and nothing in the product guarantees an event exists, so this list is the only authoritative input. Take all of it: the taxonomy-drift check is worthless on a sample.

  2. Find out which events are load-bearing. posthog.list_insights and posthog.list_cohorts, plus mixpanel.list_funnels where Mixpanel is connected. An event with volume that no saved insight uses is a deprecation candidate. An event a live dashboard depends on that has just gone quiet is a P0 incident. This step is what turns an alphabetical dump into a priority list, so do it before any querying.

  3. Inventory volume, lifespan and week-over-week deltas. posthog.query (HogQL) for per-event 30-day volume, unique actors, min(timestamp) and max(timestamp), then a second query for week-over-week volume change, excluding the current in-flight period. Classify each event active, dead or new. Flag any drop above 60% and any spike above 5x, then correlate the drop dates against release dates before calling anything breakage. Where HogQL is unavailable, use mixpanel.segmentation or amplitude.event_segmentation per event, cap to the top N by volume, and state that the long tail was not audited.

  4. Detect near-duplicates, renames and system-fired events. Normalise names to a key and cluster: Button Clicked, button_clicked and ButtonClick are one defect, and launch-driven sprawl is where they come from. A rename shows up as one event dying on the exact day another is born, so match actor overlap and property signature between a dying and an emerging event and report it as a rename rather than two defects. Then flag events with suspiciously uniform inter-arrival times or a near one-to-one ratio with another event as probably system or backend fired, and escalate loudly if any of them is being used as an activation or adoption metric.

  5. Measure property completeness per event, and identity health. posthog.query for the null or missing rate of the properties that matter for segmentation (plan, account id, platform, version) on the top 30 events by volume, plus the share of events with no identified actor. Account id can be 100% present on 40 events and 0% on the one event your retention metric uses, which is exactly the failure a global completeness number hides. A high anonymous share systematically breaks lifecycle and new-user retention, since PostHog excludes anonymous events from lifecycle status.

  6. Reconcile the disputed metric against a system of record, not against another analytics tool. postgres.get_schema then postgres.query (or bigquery.query) to count the same business entity in the application database: rows in signups, orders, documents. For money use stripe.list_payment_intents with status: "succeeded" net of stripe.list_refunds, or stripe.list_invoices with status: "paid". Convert both sides to an identical explicit UTC window at an identical grain, then report coverage as tool count divided by system-of-record count with both absolute counts shown. If GA4 is involved, include it only here, with google_analytics.run_report on dimensions: ["date"], and attach the thresholding and (other) row caveats.

  7. Report. Table one: metric, source A value, source B value, system-of-record value, coverage %, most likely cause, evidence, fix owner. Table two: event name, normalised key, duplicate siblings, first seen, last seen, 28-day volume, week-over-week change, null rate on the key property, user-fired or system-fired, saved insights affected, verdict. Then one sentence naming the single defect that invalidates the most downstream analysis, and whether the client-side loss rate is measurable at all given the connections available.

Gotchas

  • Comparing the same calendar date across tools is the most common false alarm. Each tool has its own independently configured time zone, so the same calendar range covers different hours. Convert both sides to an explicit UTC window before comparing anything, and remember that unique users, sessions and total events are three different quantities: state the grain in the output.
  • Cross-tool deltas are structural, not necessarily bugs. PostHog, Mixpanel, Amplitude and GA4 resolve identity, sessions and uniqueness differently, so a 10% to 20% delta on unique users for the same event is normal. The job is to explain the delta, not to demand it be zero. Reporting "Mixpanel and GA4 disagree" as a defect destroys the audit's credibility. Reconcile each tool against the system of record instead, and never merge two analytics tools into one number.
  • Volume drops that are real business events. A 70% fall in checkout starts after a pricing change is not a tracking bug. Correlate every cliff against release dates and against a warehouse ground-truth count before calling it breakage.
  • Late-arriving events make the newest week a false cliff. Mobile SDKs batch and retry, so events can arrive days late, and GA4 is additionally incomplete for the most recent 24 to 48 hours. Exclude the in-flight period and end the window at least two days ago.
  • Dead is not always dead. Annual renewal, tax-season and admin-only events legitimately fire a handful of times a year. Check 12 months of history before recommending anything, recommend deprecation rather than deletion, and never recommend removing an event a live insight depends on.
  • The application database is a bad system of record unless filtered. Check for deleted_at, is_test, internal email domains and seed rows first. Stripe's own objects also disagree with each other: a subscription that exists is not revenue collected, and an invoice marked paid can later be refunded or disputed, so say which object you treated as cash truth.
  • Identity resolution silently breaks every cross-source join, and client-side loss is invisible without one. A CRM email maps to a product analytics distinct id only if an identify call carried that email, and anonymous pre-signup activity under a device id will never join. Report matched over attempted for every join. If no database or payment system is connected, say plainly that the client-side loss rate is unmeasurable rather than implying the tool's number is complete. GA4 alone cannot support this audit at all: its event model, thresholding and (other) row make most checks un-runnable, so degrade to events present versus expected plus a thresholding warning.

Sequel CLI

Install Sequel skills into your agent

One command connects your agent to Sequel and installs the Sequel skill, so it knows this playbook exists and reads it when a question matches. The CLI signs you in, provisions a scoped API key and writes the config for you.

Already have an MCP client?

https://api.sequel.sh/mcp

Point it at this URL and sign in when prompted, or send an API key from Settings as a Bearer token. Skills come with it; nothing else to install. Manual setup per client

  1. 1

    Install the Sequel CLI

    One line installs the latest CLI with whatever package manager you have.

    curl -fsSL https://sequel.sh/install | sh
  2. 2

    Sign in

    Authenticate in your browser and pick an organization.

    sequel login
  3. 3

    Install into your agent

    Writes the MCP config and installs the Sequel skill file for agents that support skills. Pick an agent from the list, or target one directly by its slug.

    sequel install
    • Claude Code
      sequel install claude-code
    • Claude
      sequel install claude
    • Cursor
      sequel install cursor
    • VS Code
      sequel install vscode
    • Windsurf
      sequel install windsurf
    • Zed
      sequel install zed
    • Codex
      sequel install codex
    • OpenClaw
      sequel install openclaw
    • Hermes
      sequel install hermes

More like this

Other Product Analytics skills

Feature launch impact readout

Measure whether a feature we just shipped actually landed: adoption against an eligible denominator of users who could even reach it, breadth, depth and sustained use after the release, a pre and post impact estimate with or without a control group, and a keep, iterate or sunset call.

Find our activation moment

Discover the aha moment: which early in-product action, taken in the first session or first week, most separates users who stay from users who vanish, how to score candidate first actions by retention lift, and what share of new signups ever reach it. Covers activation metric definition, activation rate and time to value.

Is our retention curve flattening

Do users keep coming back, or do they drift away after the first week or month? Builds the cohort retention curve for your key action at its natural frequency and judges whether it flattens to a plateau or decays to zero, covering N-day versus unbounded retention, repeat usage, leaky bucket diagnosis and how the plateau compares to published benchmarks.

Read out an A/B test honestly

Is this A/B test result real, or statistically significant only by accident? Runs the validity gates before reporting anything (sample ratio mismatch, peeking, power and minimum detectable effect, novelty), then gives absolute and relative lift with a confidence interval, a ship or kill call, and the winning variant extended through to revenue instead of stopping at the proxy metric.

Signup funnel drop-off analysis

Find the step where the most users abandon signup or onboarding, and how that varies by segment.

FAQ

Frequently asked questions

Why do two analytics tools report different numbers for the same event?
Five structural causes explain most of it: time zone, since each tool has an independently configured project or property time zone; attribution window; client-side loss; metric type mismatch, where sessions, unique users and total events are three different quantities; and model differences in how credit is assigned. Walk them in that order, because each is cheap to test, and expect a residual gap rather than zero.
How much event data do ad blockers actually cost?
Client-side tracking can lose 30% to 50% of events to ad blockers and Do Not Track, per the Mixpanel community guidance on GA4 versus Mixpanel discrepancies. The tool cannot see what it never received, so the loss is invisible from inside it. Only a comparison against a server-side system of record such as the application database or Stripe can measure it.
Which number should I trust when dashboards disagree?
Neither of the two analytics tools. Declare a system of record per domain first: Stripe for money, the application database for behavioural entities, then compute coverage as tool count divided by system-of-record count over an identical UTC window at an identical entity grain. Reconciling two analytics tools against each other produces a debate; reconciling both against the system of record produces an answer.
How do I find duplicate or broken event names?
Normalise every event name to a key by lowercasing and stripping spaces, underscores and hyphens, then flag any key with more than one surface form, any pair within a Levenshtein distance of 1 or 2, and any event whose daily volume dropped to zero on the day a near-named sibling appeared. Segment's object-action convention, for example Cart Item Added, with snake_case properties, is a reasonable target naming standard to drift-check against.
Can an AI agent audit our event tracking?
Yes, and this is the most agent-native analysis in the set, because the whole thing is list events, list properties, list insights, plus a handful of queries. The agent returns a prioritised defect list keyed to the saved insights and playbooks each defect breaks, plus a coverage ratio against the system of record. Sequel provides the product analytics, database and Stripe connections over MCP and the product-metric-trust-audit playbook.

Put this playbook to work

Connect a source, ask the question, and the agent follows these steps. Free to start.