Revision note (2026-09-15). The earlier version of this article presented an in-memory JavaScript flag evaluator whose "Example test cases" were console.assert calls; a failing console.assert prints a line and exits 0, so that verification block could never fail a build. Its percentage rollout hashed the user id alone, which makes every percentage flag select the same users, a limitation the text did not mention. Its metrics object counted only enabled evaluations, so it could not report what share of users a flag reached, while the text called it "critical in measuring feature impact". This version keeps the same three toggle types and rebuilds the rollout and metrics parts around a lab with twelve tests.
The three toggle types, and which one needs care
A flag evaluator answers one question: is flag X on for this user? The earlier version had three ways to decide, and they are still the right three:
- User allowlist: the user id is in a set. Exact, small, used for internal testers.
- Segment: a predicate over user attributes (
subscription === 'beta'). Exact, but the predicate runs on every evaluation. - Percentage rollout: a deterministic function of the user maps to a bucket, and the flag is on for buckets below the rollout percentage.
The first two are lookups and need no measurement. The third has properties you can get wrong silently, and the rest of this article is about those. Tested with Node v22.22.2 on macOS (arm64), no dependencies, Node's built-in node:test runner. The lab is examples/feature-flag-metrics in the companion repository.
The evaluator
The rollout bucket is computed from the flag key and the user id together:
const crypto = require('node:crypto');
const BUCKETS = 100000; // 0.001 % resolution
function bucketFor(flagKey, userId) {
const digest = crypto.createHash('sha256').update(`${flagKey}:${userId}`).digest();
return digest.readUInt32BE(0) % BUCKETS;
}
The decision per type, with the percentage case using that bucket:
function decide(flag, flagKey, user) {
switch (flag.type) {
case 'user':
return Boolean(user?.id) && flag.enabledUsers.has(user.id) ? 'on' : 'off';
case 'segment':
for (const segment of flag.enabledSegments) {
const predicate = flag.segments[segment];
if (predicate && predicate(user)) return 'on';
}
return 'off';
case 'gradual': {
if (!user?.id) return 'off';
const bucket = bucketFor(flagKey, user.id);
return bucket < flag.rolloutPercentage * (BUCKETS / 100) ? 'on' : 'off';
}
default:
return 'off';
}
}
And the public function, which records every evaluation, including the ones that return false because the flag does not exist or a predicate threw:
function isFlagEnabled(flagKey, user) {
const flag = featureFlags[flagKey];
if (!flag) {
metrics.record(flagKey, 'off', 'unknown_flag');
return false;
}
let result;
try {
result = decide(flag, flagKey, user);
} catch (err) {
metrics.record(flagKey, 'off', 'error'); // fail closed, but count it
return false;
}
metrics.record(flagKey, result, result === 'on' ? 'match' : 'no_match');
return result === 'on';
}
Does a 30% rollout reach 30% of users?
For 100,000 users with ids user0 to user99999 and rolloutPercentage: 30, the evaluator above enabled 29,969 users (test 30% rollout enables close to 30% of 100,000 users).
The earlier version's hash, a 31-bit string hash taken modulo 100, was also fine on this point. Run verbatim from the article, hashUser(id) % 100 < 30 enabled 29,978 of 100,000 sequential ids, 29,997 of 100,000 numeric ids starting at 1,000,000, and 30,032 of 100,000 random UUIDs. The distribution claim in the old text held; the problem was elsewhere.
The flaw: one bucket per user, shared by every flag
If the bucket depends on the user id only, then every percentage flag at 30% enables exactly the same 30% of users, and a flag at 10% enables a subset of them. The lab adds a second gradual flag under a different key to the article's original code and evaluates both for 100,000 users: the decisions were identical for 100,000 of 100,000 users (test the article hash ignores the flag key).
Why this matters in production:
- The same users get every half-finished feature first. If two rollouts each carry a 1% risk of a bad experience, those users take both.
- Two flags rolled out at the same time cannot be told apart in any metric split by flag, because the exposed populations are the same population.
- A team running an experiment behind one flag and a rollout behind another has confounded them without knowing.
Keying the hash by flag makes the cohorts independent. With bucketFor(flagKey, userId), two flags at 30% were both on for 8,870 of 100,000 users and made the same decision for 57,892 (independent 30% cohorts predict 9,000 and 58,000). The 100,000-bucket resolution is borrowed from LaunchDarkly, whose documentation describes its percentage rollout as a hash of the context key and context kind divided into 100,000 buckets; the flag-key input above is this article's addition, and the test is what backs it.
Raising the percentage must keep the users you already have
A rollout that goes 10% to 30% should keep the 10% enabled and add 20% more, not reshuffle. This holds for any scheme that fixes the bucket per (flag, user) and compares it to a threshold: in the lab, raising rolloutPercentage from 10 to 30 kept all 10,057 users enabled at 10% (revised evaluator) and all 9,969 (article code). The property comes from storing the percentage and computing membership, never from storing membership. It breaks the moment someone changes the hash input (a salt, a different id field) mid-rollout, which is why the bucket function should be treated as an interface with a version, not an implementation detail.
A flag metric with a denominator
The earlier version's metric was an object that incremented a per-flag count when the flag evaluated to true. In the lab, 1,000 evaluations of a 30% flag recorded 305 activations, and the 1,000 was stored nowhere (test the activation counter has no denominator). A count of true results cannot distinguish a 30% rollout with little traffic from a 3% rollout with a lot of traffic, and it cannot tell you a flag is being evaluated for users it is not meant for.
The replacement is one counter with labels for the flag, the result, and the reason:
$ node src/exposition-demo.js
# HELP feature_flag_evaluations_total Feature flag evaluations by flag, result and reason.
# TYPE feature_flag_evaluations_total counter
feature_flag_evaluations_total{flag="dark-mode",result="off",reason="no_match"} 688
feature_flag_evaluations_total{flag="dark-mode",result="on",reason="match"} 312
feature_flag_evaluations_total{flag="does-not-exist",result="off",reason="unknown_flag"} 1
feature_flag_evaluations_total{flag="new-dashboard",result="off",reason="no_match"} 1
feature_flag_evaluations_total{flag="new-dashboard",result="on",reason="match"} 1
Exposure for dark-mode is 312 / (312 + 688) = 0.312 for that sample of 1,000 users, and in PromQL it is sum by (flag) (rate(feature_flag_evaluations_total{result="on"}[5m])) / sum by (flag) (rate(feature_flag_evaluations_total[5m])). The unknown_flag reason catches a typo in a flag key, which with the old design was indistinguishable from a flag that is simply off. The error reason (a segment predicate that throws) is counted, not swallowed; test exposure is on / total, unknown flags and throwing predicates are counted as off pins both.
The exposition text was linted with promtool check metrics from Prometheus 3.14.0 (run in Docker); it passed with exit code 0. To make sure the check means something, a deliberately wrong version without the _total suffix and without HELP was fed in too, and promtool rejected it with counter metrics should have "_total" suffix and no help text. The naming follows the Prometheus guidance for counters: a base unit-free name with _total, labels for dimensions you will actually query by, and no user id in a label. A label per user would create one series per user and is the fastest way to take down a Prometheus server.
Tests that can fail
The earlier version's verification section used console.assert. In Node, console.assert(1 === 2, 'this is false') prints Assertion failed: this is false to stderr, the process keeps running, and the exit code is 0. The lab runs exactly that and asserts on the exit code (test the article verification block cannot fail a run). It also runs the old article's assertion block verbatim: exit 0, which proves nothing either way. Use node:test with node:assert/strict; a failing assertion then fails the run.
Full run, 12 tests, 12 passed, in about 0.4 seconds; a second run produced the same counts (the only randomised input is the UUID sample, which moved from 30,032 to 30,072 enabled):
$ npm test
ok 1 - user, segment and gradual toggles behave as the article states
ok 2 - hashUser % 100 is close to uniform for sequential and random ids (30% rollout)
ok 3 - the article hash ignores the flag key: two 30% flags enable exactly the same users
ok 4 - raising the percentage keeps the users already enabled (monotonic ramp)
ok 5 - the activation counter has no denominator: it cannot tell 30% exposure from 3%
ok 6 - the article verification block cannot fail a run: console.assert exits 0
ok 7 - bucket is deterministic and bounded
ok 8 - 30% rollout enables close to 30% of 100,000 users
ok 9 - two 30% flags have independent cohorts: overlap near 9%, not 100%
ok 10 - raising the percentage keeps already-enabled users
ok 11 - exposure is on / total, unknown flags and throwing predicates are counted as off
ok 12 - exposition has one HELP, one TYPE counter line and _total-suffixed samples
What an evaluation costs
The earlier version asked readers to "confirm SDK evaluation runs in near-constant time" and gave no number. One million calls per line, single thread, one machine (Apple Silicon laptop, Node 22.22.2), after a 10,000-call warm-up:
| Evaluation | ns per call |
|---|---|
| Article code: user allowlist | 7 |
| Article code: segment predicate | 12 |
| Article code: percentage (31-bit string hash) | 24 |
| Revised: user allowlist + counter | 125 |
| Revised: segment predicate + counter | 139 |
| Revised: percentage (SHA-256 keyed by flag) + counter | 617 |
The labelled counter costs about 120 ns (building a string key and updating a Map), and SHA-256 through node:crypto costs about 600 ns. Both are far below anything that touches a network, but the 31-bit hash is 25 times cheaper than SHA-256 here. If that matters at your request rate, a non-cryptographic 32-bit hash keyed by flagKey:userId keeps the independence property at a fraction of the cost; the choice of hash is not the point, the input to it is. Not measured: contention, garbage collection under load, or a real request path.
What this does not cover
- Remote flag storage, streaming updates, caching, or an SDK/server split. Everything here is in memory, which is the right place to test the decision logic and the wrong place to run production.
- Scraping the exposition over HTTP. The text was linted by promtool; no Prometheus server scraped it.
- Multivariate flags, segment targeting beyond predicate functions, and bucketing on anything other than a user id (sessions, accounts).
- Automated rollback on metric thresholds. With the counter above you have the exposure series; wiring it to an alert or a rollout controller is a separate topic.
Reproduce it
cd examples/feature-flag-metrics
npm test
node bench.js
node src/exposition-demo.js | docker run --rm -i --entrypoint promtool prom/prometheus:v3.14.0 check metrics
