Calin Gabriel Full Stack Developer · Node.js / TypeScript

← Lab Runs in your browser

1,000 charges, one imperfect API

You need to charge 1,000 customers through a payment provider's API. The API has limits, and now and then it fails or answers slowly. Below: the obvious version, what goes wrong with it, and the four changes that fix it.

Live demo Rate limits Retries Idempotency

The provider

It behaves like a real third-party API:

  • At most 10 requests in flight. Above that it answers 503.
  • At most 100 requests in any one second. Above that it answers 429.
  • About 1 request in 20 fails with a 500.
  • About 1 in 33 takes 1.5 seconds. The charge happens even if you stop waiting.
  • Every 97th customer's card is declined (400). Sending it again won't change that.
  • It accepts an idempotency key: a request with a key it has already seen gets the first answer back instead of a new charge.

The obvious way

const results = await Promise.all(
  ids.map((id) => fetch(`${baseUrl}/charge`, {
    method: "POST",
    body: JSON.stringify({ id, amount: 1000 }),
  })),
);

It starts all 1,000 requests at the same moment. Press the first button to see what the provider makes of that.

Run it

Each button adds one change to the one before it. The chart shows every request by the time it started, coloured by what came back.

Press a button to send 1,000 charges.

Pacing: one request every 11 ms

The provider allows 100 a second, so aim for 90 and leave some room. That's one request every 11 ms. The pacer keeps one number, the time the next request may start, and every request does three things in this order: work out its turn, book the next turn, then wait.

function createPacer(ratePerSec) {
  const gapMs = 1000 / ratePerSec;
  let nextSlot = Date.now();

  return async function waitForSlot() {
    const now = Date.now();
    const myTurn = Math.max(now, nextSlot); // 1. when is my turn?
    nextSlot = myTurn + gapMs;              // 2. book the next one, now
    await sleep(myTurn - now);              // 3. only then wait
  };
}

The order is the whole trick. JavaScript runs a function without stopping until the first await. If nextSlot is updated before the wait, each of the 1,000 calls sees the slot the previous one booked. If it's updated after, all 1,000 read the same value before any of them writes it, and they all go at once.

Retrying, but not everything

A 500 or 503 is the provider's problem, and it's usually gone a moment later, so try again: after 100 ms, then 200, 400, 800. A 400 "card declined" is an answer. Sending the same request again gets the same answer, so stop.

const backoff = (attempt) => 100 * 2 ** (attempt - 1); // 100, 200, 400...

async function chargeOne(id) {
  const key = `ck-${id}`; // the same key on every retry of this charge
  for (let i = 0; i < MAX_ATTEMPTS; i++) {
    await waitForSlot(); // a retry is a request too
    let res;
    try {
      res = await fetch(`${baseUrl}/charge`, {
        method: "POST",
        headers: { "content-type": "application/json", "idempotency-key": key },
        body: JSON.stringify({ id, amount: 1000 }),
        signal: AbortSignal.timeout(1000),
      });
    } catch {
      await sleep(backoff(i + 1)); // timed out: try again
      continue;
    }
    if (res.status === 200) return { id, ok: true };
    if (res.status === 400) return { id, ok: false, reason: "declined" }; // won't change
    await sleep(backoff(i + 1)); // 500, 503, 409: try again
  }
  return { id, ok: false, reason: "gave up" };
}

Every retry goes through the pacer too. To the provider a retry is just another request, and it counts against the 100 a second like any other.

The catch: giving up doesn't stop the charge

Without a timeout, a provider that hangs keeps your request waiting forever. So after one second the client gives up and tries again. But the provider doesn't know the client gave up. It finishes the slow charge anyway.

Without a key, the retry looks like a new charge, and the customer pays twice. With the same key on every retry of one charge, the provider recognises it. If the first request is still running, it answers 409 "in progress", and the client backs off and asks again. Once it's done, it returns the first result.

The key belongs to the operation, not the customer. Charging the same customer next month needs a new key.

What this isn't

The demo is a model. It runs in simulated time, so thirteen seconds of traffic take a fraction of a second, and there is no network.

The same client ran against a mock provider in Node, over real HTTP. Started all at once, 900 of the 1,000 requests came back 429. Paced, there were no 429s and 939 customers were charged in 12.2 seconds; the other 51 valid ones hit a 500 and were never retried. With retries, all 990 valid customers were charged with 1,053 requests in 13 seconds. With the timeout but no key, 29 customers were charged twice. With the key, none. The last run ended like this:

provider stats: {
  requests: 1084,
  charges: 990,
  duplicates: 0,
  maxInflight: 9,
  rateLimited: 0,
  serverErrors: 53,
  slow: 29,
  inProgressConflicts: 2,
  declined: 10,
  retriedDeclined: 0,
} elapsed: 13236 ms

What's still missing

Look at the last run: only two 409s. A request that timed out after one second asked the pacer for a new turn, but all 1,000 first attempts had booked their turns already, so the retry went to the back of the line, about 11 seconds later. By then the slow charge had long finished.

The limit of 10 in flight held too, but by luck: at 90 a second and 15 ms per answer, only a few requests overlap. The fix for both is a pool of workers, say six, each taking the next customer when it's free. Then at most six requests are ever in flight, and a retry waits behind five others, not a thousand. That's the next version.

Where this came from

Interview preparation, published. Sending many requests to a rate-limited API comes up often in backend interviews. I wrote the client one step at a time against the mock provider, and the numbers above are from those runs. The job queue lab ended on idempotency from the worker's side; this is the same idea from the client's side.