Know what AI costs before the invoice arrives.
Every call is priced the moment it happens and attributed to a product and a person. Caps warn at a threshold and then block at the ceiling. The company report runs live off that data instead of being reconstructed from provider invoices at the end of the quarter.
Metering happens at the call, not at the invoice
Each call through the gateway writes a usage event carrying the provider, the model, token counts, credits and cost, plus the product and member it belongs to.
The distinction matters because provider invoices arrive monthly, in aggregate, per API key. If three departments share one key, the invoice cannot tell you which of them spent the money. Pricing the call as it happens is what makes attribution possible at all.
Caps that warn before they block
A spend limit is scoped either to the whole workspace or to an individual user, it runs over a period that defaults to monthly, and it carries a credit ceiling.
- Warn
- The default threshold is 80% of the ceiling. Consumption past it shows as a warning state rather than stopping anyone working.
- Block
- A limit set to block refuses calls once the ceiling is reached, and the caller receives a 402 rather than a silent failure or a surprise overage.
Warn is the default action rather than block, which is deliberate: a cap that silently stops a department mid-quarter causes more damage than the overspend it prevented.
Three billing modes
| Mode | Whose keys | How you pay |
|---|---|---|
| BYOK | Yours | The provider bills you directly. This is the default. |
| Prepaid | Ours | You top up a wallet and each call draws it down. An empty wallet returns a 402. |
| Postpaid | Ours | Invoiced monthly. Off by default and enabled per workspace. |
Bring-your-own-keys is the default because it is the mode that requires the least trust from a company evaluating us. Your keys, your provider relationship, your invoice, with the platform adding identity, metering and control on top.
The questions the report is built to answer
Reporting exists to answer the questions a board actually asks, which are rarely about tokens.
- Which department gained the most hours back?
- Attribution by product and member makes this answerable rather than anecdotal.
- What did AI cost us last quarter, by product?
- Every event carries a product, so the breakdown is a query rather than a reconstruction.
- Who is about to hit their spend cap?
- Consumption is tracked against each limit continuously, not at period end.
- Which model are we actually paying for?
- Model and provider are recorded per call, including when auto-routing chose the provider.
Straight answers.
Gateway keys can carry a product scope, and every usage event records the product and the member alongside the provider, model, tokens and cost. Attribution therefore comes from the call itself rather than from anyone remembering to tag their requests.
It depends on the action set on the limit. A warn limit raises a warning state at the threshold, which defaults to 80% of the ceiling. A block limit refuses further calls at the ceiling and returns a 402.
Yes. A spend limit is scoped either to the workspace or to an individual user, over a period that defaults to monthly.
No. Bring-your-own-keys is the default mode: you keep your provider relationship and your provider invoice, and the platform supplies identity, metering, attribution and caps on top. Prepaid and postpaid exist for companies that would rather have one bill.
Live in production
Build your company’s AI System.
Bring one real workflow to the call. We run it through the gateway while you watch, metered and attributed, before you decide anything.
30 minutes · no deck · reply within 24h