How usage is measured
This page describes the mechanism. For what you have used, see Usage. For overage settings, see Overage. Plan details are on Plans.
What is recorded
Every call from the assistant to a model is recorded as one entry in an append-only ledger. An entry holds:
- your account and the model used
- input, output and cached token counts
- the cost of the call, kept as a whole number
- the response time, the status and any error code
- the time of the call
An entry does not contain your prompt or the model's answer.
How cost is worked out
Cost is the token counts multiplied by the rate for that model. A call that fails costs nothing.
One request from you can make several calls. Planning, checking and summarising steps are calls too, and all are counted.
The check before each call
Before forwarding a call, the service checks your allowance.
Overage off or on
- Off (the default). When your allowance is used up, calls are refused. A call that was already running can take you slightly below zero. The difference is taken from the next month's allowance.
- On. Only the part beyond your allowance is counted as overage. It is billed later.
The sidebar shows a refused call as a usage-exhausted message. See quota_exhausted.
Concurrency limit
Each account has a limit on how many calls can run at once. A call over the limit is refused with a rate-limit error. The assistant tries the call up to three times, waiting a few seconds between tries. A used-up allowance is not retried. If the tries all fail, the sidebar asks you to try again shortly.