How billing works

Charged on actual token usage; how failed requests are handled.

Charged on real usage

Chat models are billed on input + output tokens at the prices shown on the Models page. Cache hits, reasoning tokens, audio and images are priced separately and itemised on your bill.

Hold and settle

We place a conservative hold when a request starts, then settle against actual usage when it finishes and release the excess immediately. A briefly lower balance mid-request is expected.

Failed requests

The test is whether upstream actually spent compute, not whether you got a 200.

Related articles

Blog