WRITINGGagandeep Bhatia
Chat with GaganChatPortfolio
Sign in with GoogleSign in with Google. Opens in new tab
‹ Index17 min left
17 min read

Building Self-Hosted Product Analytics as an npm Package, Where Counting Was the Hard Part

A capture SDK, an Express collector, a funnel and retention engine, and a React dashboard in one zero-dependency install. Shipping the events was an afternoon; counting them correctly found a funnel that double counted, a retention number wrong by 3x, and an SDK that rewrote history.

Why I Stopped Reaching for a Hosted Tool

I run four sites off one box: a portfolio, an engineering blog, a book reader, and a suite of client-side PDF tools. All four needed the same boring answers. Who visits, which pages hold them, what do they click, do they come back. The usual answer is a script tag from a hosted analytics vendor.

That trade is real and sometimes worth taking. You get mature tooling in five minutes, and in exchange you hand every visitor's behaviour to a third party, add a consent banner, and write a cookie policy. I wanted the tooling without the trade, so I built the tool.

@gagandeep023/event-analyzer is a capture SDK, an Express collector, an analysis engine, and a React dashboard, published as one npm package with zero runtime dependencies. Events land as JSONL on my own disk. All four sites report into a single collector.

The surprise was where the difficulty sat. Getting events over the network is a solved problem and took an afternoon. Counting them correctly took the rest of the time, and every real bug I found was arithmetic.

Deciding What a Product Analytics Tool Actually Is

Before writing anything, I read the public developer documentation of the mature hosted vendors end to end. Between them they expose close to 250 documented endpoints. That is not a weekend of work, and more importantly, most of it is not analysis.

So I applied one filter to every endpoint: does this exist because analytics is hard, or because selling analytics to enterprises is hard? SCIM user provisioning, SSO group sync, audit log export, per-seat permission grants, data residency controls, subscription management, taxonomy governance. All of it is real engineering. None of it answers a question about your users.

Strip that half away and the category collapses to a surprisingly small core. Nine genuine analyses survived:

  • segmentation: counts over time, split by any event or user property
  • funnel: ordered, unordered and sequential step completion, with conversion windows
  • retention: n-day, unbounded and bracket measures
  • cohort: define a user set once, then reuse it as a filter everywhere else
  • sessions: length, depth, and stickiness distributions
  • events: per-event volume and unique actors
  • breakdown: top values of a property, with an explicit long tail
  • growth: new, retained, resurrected and dormant accounting
  • activity: an hour-by-weekday matrix of when people actually show up

The finished collector exposes nine routes. Four ingest (collect, identify, group-identify, alias), one fans out to the nine analyses above, and four are operational: a metadata catalogue, a raw export, a live SSE stream, and a health check. That is the whole product. Almost everything else the hosted tools sell is a business model rather than a feature.

Four Pieces, One Install

 browser / node                your server                        you
 +--------------+              +------------------+          +---------------+
 |  ./sdk       |   batch of   |  ./backend       |          |  ./frontend   |
 |  track()     |   events     |  POST /collect   |          |  dashboard    |
 |  identify()  | -----------> |    validate      |          |  8 pages      |
 |  revenue()   |   (keepalive |    dedupe        |          +-------+-------+
 |              |    on unload)|    store         |                  |
 |  retry queue |              +--------+---------+                  | POST /query/:kind
 |  in storage  |                       |                            v
 +--------------+                       v                  +--------------------+
                              events.jsonl on disk  ----->  |  ./core            |
                                                            |   funnel()         |
                                                            |   retention()      |
                                                            |   pure functions   |
                                                            |   no I/O, no HTTP  |
                                                            +--------------------+
How the pieces fit

The subpath exports matter more than they look. ./core is pure functions over an array of events, with no I/O and no Express anywhere in its import graph, so you can pull in funnel() and run it against rows from any database with no server involved. express, react and react-dom are optional peer dependencies, which means installing this for the analysis engine alone adds exactly one package to your tree.

typescript
// Any of the four pieces, independently.
import { funnel, retention } from '@gagandeep023/event-analyzer/core';
import { createClient }      from '@gagandeep023/event-analyzer/sdk';
import { createEventAnalyzerRouter } from '@gagandeep023/event-analyzer/backend';
import { EventAnalyzerDashboard }    from '@gagandeep023/event-analyzer/frontend';

Sessions Without a Session Table

The first design decision that kept paying off: a session is not a record. It is a session_id that equals the session's own start timestamp.

typescript
/**
 * Sessions are derived, never stored. A session is the group of one user's
 * events sharing a `session_id`, which is itself the session's start timestamp.
 * That choice removes the need for a session table anywhere.
 */

Because the id carries its own start time, session length is a max minus the id, ordering is a numeric sort, and nothing has to be written when a session opens or cleaned up when it ends. There is no half-finished row to clean up when someone kills a tab mid-session, because there was never a row.

The browser SDK stamps it. Server-side SDKs and raw HTTP clients will not have one, so the engine reconstructs sessions for those by splitting each user's sorted event stream wherever the inter-event gap exceeds a timeout, 30 minutes by default. Both paths coexist inside a single query, which matters when one site posts from a browser and another posts from a server and you still want one number.

One Person, Two Identities

Events arrive keyed by device_id before login and user_id after it. Without resolution, one person is two users, and every retention number is wrong in the same direction: too low.

This is a connected components problem, so it gets union-find. Every user_id and every device_id becomes a node, every event carrying both unions them, and the canonical key falls out of the find.

typescript
/** Namespaced node ids, so a device id can never collide with a user id. */
function userNode(id: string): string {
  return `u:${id}`;
}
function deviceNode(id: string): string {
  return `d:${id}`;
}

The namespacing is not paranoia. Ids in the wild are uuids on both sides, drawn from the same alphabet, and an accidental collision would silently merge two unrelated people into one. Two characters of prefix make that impossible by construction.

The graph is built once per query and threaded down into every analysis, so a dashboard page that runs six analyses pays for resolution once rather than six times. It also reports skipped, the count of events carrying neither identifier, because an analysis that quietly drops rows is worse than one that tells you it dropped them.

Retention Is Three Different Questions

This is the part I would have got wrong if I had not read carefully first, and I suspect most hand-rolled implementations do get it wrong.

"Day 7 retention" is ambiguous. It can mean three genuinely different things:

  • n-day: the user came back on exactly day 7
  • unbounded: the user came back on day 7 or any day after it
  • bracket: the user came back inside a window you define, say days 5 through 9

Shipping only n-day is the usual way a retention implementation is quietly wrong, and the gap is not academic. Measured on this deployment's own data, n-day reports roughly a third of what unbounded reports for the same users, about 20 percent against about 58 percent.

The consequence is worth sitting with. If your product has a weekly rhythm rather than a daily one, day-7 n-day retention makes a healthy product look like it is bleeding users, because a person who comes back on day 8 is counted as churned. All three measures are implemented, and the measure is a field on the query rather than a constant in the code.

Young cohorts cannot be measured yet

The second retention trap: a cohort that started yesterday cannot have a day-30 number. It is not zero, it is unknown. Plotting it as zero draws a cliff at the right edge of every retention curve, and that cliff is an artifact of the calendar rather than a fact about the product.

typescript
const observable = isObservable(cohortStart, spec.maxPeriod, q, tz);

cells.push({
  period: spec.period,
  retained,
  rate: ratio(retained, members.length),
  incomplete: !observable,
});

if (observable) {
  curveRetained[i]! += retained;
  curveDenominator[i]! += members.length;
} else {
  // Kept separately, so an incomplete point still reports something
  // rather than a bare zero.
  partialRetained[i]! += retained;
  partialDenominator[i]! += members.length;
}

Every cell carries its own denominator and an incomplete flag. The aggregate curve is built only from cohorts old enough to be measured fairly, with the under-observed counts accumulated separately so an incomplete point still reports a number instead of a zero. The dashboard renders those cells differently, so a thin tail reads as "not known yet", which is exactly what it is.

The Funnel Bug That Only Appeared in Totals

A funnel counts two ways. Uniques asks how many people reached each step. Totals asks how many times each step was reached, which is a different question the moment one person runs the funnel three times.

My first implementation walked each user's events, found their best attempt, and stopped.

typescript
// The original. Correct for `uniques`, quietly wrong for `totals`.
if (attempt.depth === q.steps.length) break;

That early exit is right for uniques. Once somebody has completed the funnel they cannot beat it, so there is nothing left to find. Under totals it silently discards every attempt after the first complete one, so a user who checked out three times was reported as one.

Removing the exit surfaced the opposite bug. Without it, a step-0 match sitting inside an already-counted walk starts a fresh attempt, and the same checkout gets counted twice. Both fixes are needed together.

typescript
let consumedUpto = -1;
for (const startIdx of starts) {
  // Under `totals` each attempt consumes its events, so a completed walk does
  // not get re-counted from a step-0 match that sat inside it.
  if (q.countBy === 'totals' && startIdx <= consumedUpto) continue;

  const attempt =
    q.order === 'unordered'
      ? walkUnordered(userEvents, startIdx, q, opts)
      : walkInOrder(userEvents, startIdx, q, opts);

  attempts.push(attempt);
  if (q.countBy === 'totals') consumedUpto = attempt.lastIndex;

  // A complete funnel cannot be beaten, so under `uniques` there is nothing
  // left to find. Under `totals` every attempt counts, so keep walking.
  if (q.countBy !== 'totals' && attempt.depth === q.steps.length) break;
}

What makes this class of bug dangerous is that nothing announces it. Nothing crashes, no test fails unless you happened to write the test that counts repeat conversions, and the number that comes out the other end is entirely plausible. It is just smaller than the truth.

The Collector Is a Trust Boundary

/collect has to be public, because a browser has to reach it. /query/:kind does not. They get different gates: a write key for collection, a real owner check for analysis. Two trust levels inside one router, which is the only security decision in the package that actually matters.

Validation runs per event, never per batch. A batch of 50 containing one malformed event should cost you one event, not 50.

json
{
  "code": 400,
  "events_ingested": 47,
  "events_with_missing_fields": { "event_type": [12] },
  "events_with_invalid_fields":  { "time": [31] },
  "duplicate_events": [44]
}

Failures are addressed by array index. A client reading that response can drop exactly the poison events and retry the rest without guessing which ones offended.

A duplicate is not an error

Every event carries an insert_id, and the collector remembers recent ones so that a retry after a timeout does not double count. I first treated duplicates as a rejection reason like any other, which produced a bug that appears on exactly the path insert_id exists to protect.

A client sends a batch. The response times out on the way back. The client does the right thing and retries the identical batch. Now every event is a duplicate, nothing is accepted, and the collector answers 400. The client sees a client error, retries, and loops forever on a batch that was stored correctly the first time.

typescript
// A batch rejected purely as duplicates is a successful no-op, not a client
// error. This is exactly the retry-after-timeout case insert_id exists for.
const realIssues = issues.filter((i) => i.kind !== 'duplicate');
if (accepted.length === 0 && realIssues.length > 0) {
  return badRequest(res, issues);
}

A batch that is entirely duplicates means the first attempt worked. That is a 200.

The SDK Bug That Rewrote History

The subtlest bug in the package. enqueue() is async, because plugins in the pipeline can be async. The original version read the current identity after awaiting readiness, which makes this ordinary-looking sequence produce a wrong event.

typescript
ea.track('Pricing Viewed');   // anonymous, there is no user_id yet
await signIn();
ea.setUserId('user_123');     // ...and the awaited enqueue resumes after this

The pricing view genuinely happened while the visitor was anonymous. But the await inside enqueue resumed after setUserId, so the event was stamped user_123 and attributed to a logged-in user retroactively. The same hole ran in the other direction for consent: an event dropped because tracking was off could be resurrected by a later setOptOut(false).

The fix is to snapshot everything mutable synchronously, before any await can interleave.

typescript
/**
 * Everything mutable is snapshotted SYNCHRONOUSLY, before any await. Reading
 * identity or opt-out after the microtask resumes would let a `setUserId`
 * that happened later retroactively attribute an event that was anonymous
 * when it occurred, and would let a `setOptOut(false)` resurrect an event
 * that was dropped when it was tracked.
 */
private enqueue(partial: AnalyticsEvent): Promise<DeliveryResult> {
  const optedOutNow = this.optOut;
  const now = partial.time ?? Date.now();
  // ...stamp the event from the snapshot, then await freely
}

An event describes a moment. Anything you read after an await describes a different one.

Zero Dependencies, Including the Charts

The package installs nothing. That was mostly easy to hold, because the engine is arithmetic over arrays and the collector is a single Express router. The place it bit was the dashboard.

I started with a charting library. It was immediately the heaviest thing in the tree, and for six chart types (a time series, a donut, bars, stacked bars, a heatmap and an activity grid) it was doing very little that a few hundred lines of SVG could not do. So I dropped it and hand-rolled the six.

What that bought, beyond install size, was vocabulary. The charts speak the engine's types directly. A retention cell knows whether it is incomplete, so the grid renders it hatched rather than as a zero. Handing that to a generic charting library means flattening it into an anonymous series first, and the flag is precisely the thing that gets lost in the flattening.

Deploying It Across Four Sites

Three of the sites proxy through their own /api path. The PDF tools site is purely client-side with no backend of its own, so it posts cross-origin to the API subdomain directly. That one case found most of the deployment bugs.

The prerender step logged itself as traffic

The portfolio prerenders its routes at build time with a headless browser. That browser loads the real page, which loads the real SDK, which faithfully reported 39 page views from a build machine. Analytics polluted by its own deploy pipeline, which is a satisfying kind of stupid.

typescript
function isAutomated(): boolean {
  if (typeof navigator === 'undefined') return false;
  if (navigator.webdriver) return true;
  return / HeadlessChrome\/|Puppeteer|Playwright/i.test(navigator.userAgent || '');
}

export const analytics = createClient({
  endpoint: `${API_BASE}/events/collect`,
  defaultContext: { site: SITE },
  optOut: isAutomated(),
});

Next build: 21 routes prerendered, zero events recorded.

Vite loads .env in every mode

The PDF site shipped to production with http://localhost:3001 as its collector endpoint. A .env.production existed and was correct. The catch is that Vite loads .env in every mode and .env.production only layers on top, so a variable set in .env and not restated in .env.production leaks straight into the production bundle.

Nothing fails. The build succeeds, the site works perfectly, and the analytics simply never arrive. I found it by grepping the built bundle for the endpoint string instead of trusting that a green build meant a correct build, which is now the habit.

The version constant nobody updates

Every event is tagged with the library that sent it. At package version 0.5.0, every event in production was tagged event-analyzer-sdk/0.1.0, because SDK_VERSION had been written once at the start and never touched again.

That makes "which version sent this event" unanswerable at exactly the moment you need to ask it, which is when one site's numbers look wrong and you want to know whether it is running old code. Fixed, along with a test asserting that every version constant matches package.json, so the next drift is a failing test rather than a discovery.

What It Looks Like Running

Eight dashboard pages: overview, audience, pages, clicks, events, funnels, retention, and a live feed over SSE. 458 tests across 25 files. Zero runtime dependencies.

The clicks page is the one I use most, and the one I nearly shipped broken. The capture plugin originally recorded each clicked element's tag, selector and text, which is enough to tell you that a link labelled "Read more" was pressed 400 times and nothing whatsoever about where those 400 people went. It now records the href, its host, and whether the destination is external, which turns "what gets clicked" into "where does this site actually send people", and only the second question is worth having a page for.

Capture is off by default, all of it. Nothing is autocaptured until you ask for it by name, click capture is allowlist-only, text is truncated, input values are never read, and password fields are excluded outright with a test that proves it stays that way.

Try It

bash
npm i @gagandeep023/event-analyzer

The full setup guide, an API reference for all nine analyses, and a live dashboard are at analytics.gagandeep023.com. The package is MIT licensed and the source is on GitHub.

Spotted a typo or have a thought on this post?