Where we left off

In the last post, identity gave every AI call an owner. Now imagine that owner keeps calling.

At some point, their usage crosses a level that matters to you. Maybe you want to move them to a cheaper model. Maybe you want to stop them outright. You could put that logic in your application:

if usage > 100_000:
    block()
elif usage > 50_000:
    use_cheaper_model()

Or you can define it once as a workflow. Your application stays simple:

usage.chat({
  identity: 'cust_acme',
  workflow: 'support-agent',
  ...
})

The workflow tells Vibe what to do as that identity's usage changes.

A workflow is a ladder of decisions

A workflow is made of tiers: "when usage reaches this, do that." For example:

100,000 tokens → Block
 50,000 tokens → Route to a cheaper model

Tiers are evaluated top to bottom, and the first matching tier wins — so the strictest threshold has to come first, or it never gets reached. Here's that shape in the Console:

Console policy editor showing two tiers: usage reaches 100,000 tokens then Block the request, usage reaches 1 token then Route to a different model

A workflow isn't a pipeline every request steps through — it's a set of decisions evaluated against the identity's usage.

What can a tier do?

ActionWhat happens
RouteSend the call to a different model, potentially cheaper. result.model reports what actually ran.
DegradeReport a lower effort or quality level on the result, for your code to act on.
BlockReject the call before the provider is contacted.

Your application doesn't implement any of that. It can ask for one model while the workflow decides another actually runs — and not just within one provider:

Console model picker for a Route tier, searching 'anthropic · claude-' and showing anthropic · claude-opus-5 as a match

Your app asked for one model. Vibe decided another actually ran — no code change required.

The usage belongs to the identity

This is where identity matters. A workflow isn't watching one global counter — it's evaluating the usage tied to that identity:

cust_acme    → 42,000 tokens
cust_other   →  8,000 tokens

Both can use the same workflow and reach different decisions, because their usage is different.

And it's not only about blocking

A matched tier can also fire an effect alongside its action — notify Slack, call a webhook, or report usage to a Stripe meter:

100,000 tokens
    ↓
Block + Notify

Change the rule, not the code

Your application doesn't need to know the customer's allowance, when to downgrade them, or when to block them. That lives outside the application. Today:

50k → cheaper model
100k → block

Tomorrow:

75k → cheaper model
150k → block

The call stays usage.chat({ identity, workflow: 'support-agent', ... }). Only the workflow changes.

That's what a workflow gives you: usage-driven behavior attached to the identity, without putting the rules in your application code.

← PreviousGet started with Vibe in 2 minutes Next →Stop AI costs before they happen

FAQ

Do I need a workflow for my first call?

No. Leave workflow off and the call is recorded with no policy applied.

What happens if tiers are ordered wrong?

Nothing errors. The first matching tier wins, so a low threshold listed before a high one catches every call and the high one never fires.