Writing Backend correctness

Green tests, forty grants, limit ten

A serial suite can bless a quota check that collapses the moment two requests overlap.

Local tests like a single clerk. Production is a crowd at the same counter.

I keep seeing the same shape: a limit of N, a read of “how many used,” a branch, then a write. The suite is sequential. Every case passes. Two overlapping requests in a real process both read “under limit,” both proceed, and you issued more than N.

This note is that pattern only. No product, no vendor scorecard. A small in-process lab you can rerun.

What the suite actually proves

A typical unit test calls take() once, or loops one after another on one thread of control. That proves the happy path and the “already full” path. It does not prove two take() calls that interleave between the read and the write.

I ran node lab/check-then-act.mjs in this repo on 5 August 2026. Setup: limit 10, 40 concurrent workers, an 8 ms delay between check and increment (stand-in for a DB round-trip or a bit of work). Five rounds each.

  • Serial (one at a time): avg used 10, max 10, never over.
  • Naive check-then-act + delay: avg used 40, max 40, over limit every round.
  • Serialized take() (promise chain): avg used 10, max 10, never over.

Serial matches what a green suite usually sees. Naive concurrent granted all 40. The lock path stayed at 10.

Raw report from that run:

{
  "limit": 10,
  "workers": 40,
  "windowMs": 8,
  "serial": { "avgUsed": 10, "maxUsed": 10, "overLimit": false },
  "naive": { "avgUsed": 40, "maxUsed": 40, "overLimit": true },
  "locked": { "avgUsed": 10, "maxUsed": 10, "overLimit": false }
}

The naive body is the bug in miniature:

async function take() {
  if (used >= LIMIT) return false;
  await sleep(WINDOW_MS);
  used += 1;
  return true;
}

Every worker passes the if before anyone increments. The delay is the race window. In production the window is “time to the database,” not setTimeout, but the interleaving is the same class of mistake.

What to change in the check

Fix the shared counter, not the test names.

  • Make take() atomic relative to other take() calls: one transaction, one compare-and-swap, one row lock, or a single-threaded actor that owns the counter.
  • Put the limit in the write: UPDATE … SET used = used + 1 WHERE used < :limit, then check rows affected. A read-then-write in app code is the naive snippet again.
  • If you scale out to multiple processes, an in-memory mutex is not enough. The lock has to live where the counter lives.

The serialized path in the lab is a promise chain around take(). It is enough for one Node process. It is not a distributed lock. I used it only to show the race is in the interleaving, not in “JavaScript can’t count.”

How I would test this next time

Keep the serial cases. Add one test that fires many overlapping take() calls and asserts used <= LIMIT. If the implementation is a database, run that test against the same isolation level you ship, not against a fake that serializes everything.

If that concurrent case is annoying to write, that is a signal: the production shape is annoying too, and the suite has been lying by omission.

Rerun: node lab/check-then-act.mjs. Numbers will jitter with scheduling; overshoot on the naive path should stay obvious.

A public notebook

Notes from real delivery: racey quotas, sync when a device is offline, messaging APIs, and the gap between a clean local demo and production.

Not a product catalog, a tutorial syllabus, or a course funnel. If a post names a tool, I used it. If it describes a failure, it happened.

About Contact