WRITINGGagandeep Bhatia
Chat with GaganChatPortfolio
Sign in with GoogleSign in with Google. Opens in new tab
‹ Index17 min left
17 min read

The Defaults That Can Never Be Fixed

fetch() still waits forever. A pg pool still queues forever. Nobody can change that without breaking the entire ecosystem, so the gap is permanent and the tooling has to live inside your running app. Here is what I learned building a runtime auditor for it, including the three bugs that only showed up once the whole thing ran at once.

@gagandeep023/production-auditPRODUCTION-AUDIT ON NPMSOURCE ON GITHUB

There is a category of production incident that has nothing to do with the code anyone wrote. The logic is correct, the tests pass, the review found nothing. The service falls over anyway, because a library shipped a default that nobody questioned and nobody could see.

I went looking for a problem worth building an npm package around, and the filter I ended up using was this: does the painful surface live in the user's own code, and can a platform fix it instead? If a registry, a runtime, a bundler or a regulator owns the problem, they will eventually close it, and anything you build is a stopgap with an expiry date. That question disqualified three candidates in a single afternoon. This one survived, and it survived for a reason that is worth the rest of the article.

Eleven of Thirty

An analysis of thirty real production incidents, each costing somewhere between ten thousand and a million dollars, found that eleven of them traced back to a single configuration default. Not a bug. A default value that shipped with a framework, a database driver or an infrastructure tool, and was never revisited.

The companies had no connection to each other. The same handful of root causes kept appearing, independently, in organisations that had never spoken. That is the signature of a problem in the tooling rather than a problem in the teams.

Three of them are still true as I write this:

  • fetch() has no timeout. Not a long one. None. A slow or hung upstream holds the request open until the connection dies on its own, which may be never.
  • A pg pool's connectionTimeoutMillis defaults to 0, which means a caller waiting for a free connection waits forever. Under load, every request queues silently, and the symptom reads as a slow database while the database sits idle and healthy.
  • A Node HTTP server closes idle keep-alive sockets after five seconds. AWS ALB's default idle timeout is sixty. The load balancer reuses a socket Node has just decided to close, and the client gets an intermittent 502 that reproduces nowhere else.

Dangerous Because It Mostly Works

The sharpest description of this class of problem I found came from a practitioner writeup about keep-alive. It is not dangerous because it is broken. It is dangerous because it mostly works. That is what fools teams.

A missing fetch timeout is invisible when your upstream responds in 40ms. It stays invisible through local development, through CI, through staging, through the first six months of production. It becomes visible on the one afternoon a downstream service stops answering instead of returning an error, and by then the failure has propagated through every caller that was waiting.

None of these defaults fail loudly. They fail under conditions your test suite does not create, in a way that looks like a different problem entirely.

The Defaults Are Frozen Forever

Here is the part that made me pick this problem over the others.

Node cannot make fetch() default to a timeout. The moment it did, every application that legitimately holds a long-lived request would break on a patch upgrade. pg cannot change its pool behaviour for the same reason. These values are not unfixed because nobody has got round to it. They are frozen by backwards compatibility, indefinitely, and any fix is a breaking change for the entire ecosystem.

Compare that to the other candidates I looked at. npm supply-chain tooling? npm owns that problem and is actively closing it on a published timeline. Prompt-injection detection? Occupied by several projects, and detection loses that arms race by construction. In both cases the gap has an owner and an expiry date.

Here, nobody can close the gap. That is an unusual property, and it is the whole reason the tool is worth building.

Why Your Linter Cannot See It

The obvious response is to write an ESLint rule. It does not work, and understanding why is the technical crux of the whole design.

A static linter reads your source and sees new Pool(). It cannot know the effective value, because the effective value does not exist until environment variables, config merging, framework defaults and the specific installed version of the library have all had their say. Two applications with byte-identical source can have different effective pool sizes.

It gets worse. The most valuable findings are not single wrong values at all. They are relationships between two numbers, and one of the two is not in your codebase:

your code                      your infrastructure
----------                     -------------------
server.keepAliveTimeout        load balancer idle timeout
     5000ms          <              60000ms

           |
           v
  LB reuses a socket Node is closing
           |
           v
  intermittent 502s, reproduce nowhere
The finding no source reader can produce

Neither number is wrong on its own. 5000 is a perfectly reasonable keep-alive. 60000 is a perfectly reasonable load balancer setting. The defect is the ordering, and half of it lives in Terraform.

So the tool has to run inside the booted application and introspect what the objects actually are.

diagnostics_channel, the Part That Patches Nothing

Putting a library in the boot path of a production service is a big ask, so I wanted the observation half to be provably harmless before anything was allowed to modify a running system.

diagnostics_channel is a Node core module that most people have never used, and undici already publishes to it. Outbound HTTP traffic can be watched with no monkey-patching whatsoever:

typescript
import { subscribe } from 'node:diagnostics_channel';

subscribe('undici:request:create', (message) => {
  // A real request, as it happens, with no interception anywhere.
  const origin = originOf(message);
  state.requests += 1;
  if (origin !== undefined) state.origins.add(origin);
});

There is a limit here that I want to state plainly, because glossing over it would be the exact failure this package exists to find. The channel publishes the request. It does not publish the dispatcher, and the dispatcher is where headersTimeout and bodyTimeout live. So this observer can prove traffic is flowing and to where, and it genuinely cannot read those two values.

The honest thing to do with a value you cannot read is to say so. Every setting in the package carries one of three states, and the third one is the important one:

typescript
export type SettingState = 'explicit' | 'unset' | 'unknown';

// explicit -> the developer chose this, never touch it
// unset    -> nobody chose it, the library default applies
// unknown  -> we could not determine it, and that is not the same as fine

"Could not determine" and "nothing wrong" must never look the same in a report. A tool that quietly reports unknown values as passing is worse than no tool, because it produces confidence it has not earned.

The Hardest Problem in the Package

For libraries with no diagnostics channel, construction has to be intercepted. That part is mechanical. The part that is not mechanical is knowing whether a value was chosen.

keepAliveTimeout defaults to 5000. A developer who deliberately writes server.keepAliveTimeout = 5000, having thought hard about their load balancer and concluded that five seconds is right, is indistinguishable by value from a developer who never touched it.

This matters enormously, because the first rule of the fixing half is that it never overrides an explicit choice. If intent is inferred from the value, that rule is unimplementable, and a tool that overwrites deliberate configuration gets removed within a week.

So the value is never used to infer intent. Assignment is observed instead:

typescript
for (const name of TRACKED) {
  const current = server[name];
  state.values.set(name, current);

  Object.defineProperty(server, name, {
    configurable: true,
    enumerable: true,
    get: () => state.values.get(name),
    set: (value) => {
      state.assigned.add(name);   // <- the entire guarantee
      state.values.set(name, value);
      publish(server);
    },
  });
}

Anything the developer touched is theirs. Presence of a constructor option key counts as explicit too, even when its value equals the library default, because typing the key is the choice. This one mechanism is what the whole never-override promise rests on.

There is a timing consequence. The developer's own assignment runs in the same tick, immediately after createServer returns. So filling in a default has to wait for the end of that tick:

typescript
// The developer's `server.keepAliveTimeout = ...` runs in this same tick.
// Waiting until it ends is what lets their assignment win without us
// having to guess from a value that is identical either way.
if (isPolicyRelevant()) setImmediate(() => fillFromPolicy(server));

The ESM Problem, and the Two Entry Points It Forces

I wanted the pitch to be one line at the top of your entrypoint. It cannot honestly be, and working out why changed the API.

javascript
import { Pool } from 'pg';                            // evaluated first
import { applySafeDefaults } from '@gagandeep023/...'; // and this
applySafeDefaults();                                  // only now do we run

ESM imports are hoisted and evaluated before any of your code. By the time that call executes, pg is loaded and Pool is a binding somebody else is already holding. Patching the module export afterwards does not reach it.

Full coverage needs a preload, not a function call, which is the same problem OpenTelemetry has and the same solution:

bash
node --import @gagandeep023/production-audit/register app.js

So there are two entry points and the README labels them honestly. The flag covers everything. The function covers globals like fetch, plus anything constructed after the call, and misses modules your app already imported. Claiming otherwise would ship a tool that silently protects less than it says, which is precisely the failure class this package exists to catch.

Three Bugs That Only Appeared When Everything Ran at Once

Every unit test passed while each of these was broken. All three were caught by the one test that runs a real child process under the real preload against a real application.

1. A Proxy is not a stack frame

Pools are wrapped with a Proxy construct trap, which preserves instanceof, statics and the prototype chain in a way a subclass would not. To report where a pool was built, the trap captures a stack trace and hands V8 a function to cut the stack above.

typescript
// Wrong. `proxied` is a Proxy object and never appears as a stack frame,
// so captureStackTrace matches nothing, keeps the whole stack, and
// silently reports OUR file as the pool's call site.
observePool(options, version, proxied);

// Right. A named function declaration, passed by identity.
function construct(target, args, newTarget) {
  const site = captureCallSite(construct);
  ...
}
const proxied = new Proxy(Pool, { construct });

The failure mode is the nasty kind: no error, no warning, just a plausible-looking file path pointing at the wrong place.

2. Filtering stack frames by path

The first version identified its own frames by checking whether the path contained the package name. That works right up until somebody's project directory is called production-audit, which is exactly what the repository itself is called. Every call site in the test suite came back undefined, because every frame looked like ours.

The fix was to stop matching paths entirely and let V8 do it, by handing captureStackTrace the observer's own outermost function. The path-based check was not just failing in the tests, it was a latent bug for any user whose directory happened to share a name.

3. The fixer created its own finding

This is the one I am most glad the test caught, because it is the exact failure a reliability tool must never have.

Enforce mode filled in keepAliveTimeout at 65000, correctly, to sit above a typical load balancer idle timeout. It left headersTimeout at Node's default of 60000. Node requires headersTimeout to exceed keepAliveTimeout, so the tool had just introduced a violation of one of its own rules.

before:  keepAliveTimeout 5000   headersTimeout 60000   ok

apply:   keepAliveTimeout 65000  headersTimeout 60000   <- broken by us

fixed:   keepAliveTimeout 65000  headersTimeout 66000   ok
What enforce mode did to itself

The fix was conceptual rather than mechanical. Settings that have to move together belong in one rule, checked with match: all, so the pair is never raised by halves.

A reliability tool that causes an outage is worse than no tool at all. I had written that sentence in the design doc a day earlier, and still shipped the bug. The end-to-end test is the only reason it did not reach npm.

The Four Rules That Keep the Fixer Safe

Given the above, the constraints on the fixing half are not negotiable:

  • Only fill what is unset. A deliberate connectionTimeoutMillis of 0 stands, because the developer meant it. The audit still reports it; the fixer leaves it alone.
  • Log every change at boot, loudly, naming the rule that caused it. Silent behaviour modification is how a helpful library becomes a three-hour debugging session six months later.
  • Opt in per subsystem. All-or-nothing means people either take risks they did not evaluate or skip the feature entirely.
  • Report mode first. The default call computes and logs exactly what it would change, and changes nothing. This is the same warn-then-enforce shape that makes TypeScript strict mode adoptable.
typescript
import { applySafeDefaults } from '@gagandeep023/production-audit';

applySafeDefaults();                                 // report only
applySafeDefaults({ mode: 'enforce', pg: true });    // then one subsystem

Nobody puts an unfamiliar library in the boot path of a production service in enforce mode on day one. A tool that requires them to is a tool they will not adopt.

Why the Rules Are Data

Every check is a JSON record, not a function. Contributing one requires no knowledge of the internals at all:

json
{
  "id": "pg/pool-connection-timeout-unbounded",
  "library": "pg",
  "versionRange": ">=8",
  "settings": ["connectionTimeoutMillis"],
  "dangerousWhen": { "kind": "unsetOr", "value": 0 },
  "severity": "critical",
  "why": "...the symptom reads as a slow database, while the database itself is idle and healthy.",
  "safeDefaults": { "connectionTimeoutMillis": 5000 },
  "profileOverrides": { "worker": { "connectionTimeoutMillis": 30000 } }
}

This is the part that keeps growing without the maintainer doing all the work, which is what lets an open-source project outlive its author's enthusiasm. Every entry is somebody's postmortem turned into a check that everyone else inherits for free.

Two details make the data format work in practice. The profileOverrides field exists because one timeout number is right for a payment path and wrong for a batch worker, and a tool that produces a wall of findings which are correct in general and wrong here gets muted once and never looked at again. A null override disables a rule for a profile entirely, which is how a statement timeout check stays off a CLI.

The other is that rules are validated in CI. A malformed rule does not throw at runtime. It silently never matches, which is the worst failure a check can have, so the schema test is the only thing standing between a bad contribution and a check that quietly does nothing.

Suppressions That Expire

Every static analysis tool eventually accumulates a suppression list nobody revisits. The fix is to make an entry impossible to write without a reason and a date:

typescript
{
  ruleId: 'pg/pool-connection-timeout-unbounded',
  reason: 'migration in flight, tracked in OPS-412',
  expires: '2026-12-01',
}

An expired suppression turns back into a finding. An expired suppression whose rule no longer fires is reported separately at info level, so stale entries get deleted rather than accumulating. An unparseable expiry date is treated as expired, because a suppression that cannot be checked must not silence anything.

Testing for Silence

The failure mode that kills a tool like this is not missing a problem. It is firing on configuration somebody thought about. One false positive on deliberate code and the whole thing gets muted, and muted tools get deleted.

So the most important test in the suite asserts nothing happens:

typescript
it('an explicitly configured app produces zero findings', () => {
  const config = effective([
    configuredPool(),
    configuredServer(),
    fetchSite(explicit('AbortSignal')),
  ]);
  expect(evaluate(config, { rules })).toEqual([]);
});

it('an explicit value that happens to equal the default is still explicit', () => {
  // keepAliveTimeout defaults to 5000. Writing `= 5000` is
  // indistinguishable by value from never touching it.
  expect(evaluate(effective([serverWithExplicitDefaults]), { rules })).toEqual([]);
});

There is also a mutation check, because a check only ever observed passing has not been shown to work. Disable a rule's condition, and assert the failing fixture starts passing. If it does not, the rule was never the thing producing the finding.

What It Will Not Catch

Stating the limits is part of the product, not an apology for it.

  • Native-ESM packages are not intercepted. The module hook wraps CommonJS loads, which covers pg and most drivers, but a package shipping only ESM resolves through a different path and is missed.
  • A version-scoped rule stays inert when the installed version cannot be resolved. Firing would be a guess, and a guessed finding the user cannot verify is how a tool earns a permanent mute.
  • Only what actually ran is checked. A pool constructed lazily on the first query does not exist at boot, which is why the report prints at exit and why there is a line listing what was watched and never happened.
  • Every default nobody has written a rule for yet. The corpus ships with seven. That is most of the cascading-failure incident class and nowhere near all of it.

Try It

Point it at something you already run. It takes one command and changes nothing:

bash
npx production-audit run dist/server.js

The output names the setting, says what the library default actually does, gives you the file and line where the object was built, and shows the lines to add as a diff. If your application is already configured properly, it says so and stops, which is the outcome I spent most of the test suite protecting.

If you have a default that cost you an incident, the most useful thing you can send is a rule. It is a JSON record in one file, and it turns your bad afternoon into a check everyone else gets for free.

Spotted a typo or have a thought on this post?