Check first, then do the work

Every request to a limited route goes through the same three steps:

  1. Ask. Who is this, and do they have any allowance left?
  2. Decide. If yes, the request runs. If no, it stops here and returns an error like 429 Too Many Requests. Your code never starts.
  3. Record. When the work finishes, the usage is counted against the caller.

That's the whole idea. The check happens before the work, so a blocked request costs you nothing.

You add it once

You don't write a check into every endpoint. You put one piece of code in front of your routes, and it runs the three steps above for each request. Which routes are limited, who counts as a caller, and how much they get all live in policy, not in your handlers.

The code for it depends on your stack. The docs have the setup for each language.

What if the limit check is down?

Every design has to answer this, so decide before it happens:

Neither is wrong. It depends on whether an overage or an outage costs you more.

Why start with a request that costs 1

Here every request counts as exactly 1, so nothing has to be guessed. That makes it the easiest kind of usage to enforce, and a good first example. AI is harder, because the cost isn't known until the work is done. That's the next post.

← PreviousYour app needs a usage policy Next →AI usage isn't one request = one unit

FAQ

Does this work for things that aren't AI?

Yes. The engine counts whatever you tell it to: API calls, exports, file downloads, model calls.

What does the caller see when blocked?

Whatever you choose. A common answer is an HTTP 429 with a short error message.

Fail open or fail closed?

Closed protects your limits. Open protects your uptime. Pick one on purpose.