Every token is billed.
Output is billed 4–5×.
LLM APIs charge for what you send and what the model says back — at very different rates. Watch one call get billed, then run the audit that tells you which half of your bill to attack first.
Rates shown are illustrative shapes, not quotes — see sources below. Tap any element in the sections that follow for detail.
One call, two prices
Everything you send bills at the base rate. Everything the model says back bills at the premium. Click any layer.
Same million tokens, different bill
Providers price the compute difference between reading and generating. The ratio holds across tiers.
Why long chats get expensive
The whole input stack is resent — and re-billed — on every turn. Click any bar.
The Asymmetry Audit
One division tells you which half of the bill to attack first. Drag your output share of spend.
Compress what you send
- Prompt caching — stable prefix first
- Compress the system prompt
- Summarize older history
- Surgical RAG — rerank, inject less
Cap what comes back
- Cap max_tokens per call
- Terse, structured output formats
- Cut filler from response templates
- Then recompute S — the priority flips
The expensive token type is not automatically the dominant cost — volume skew can beat price skew. That's what S catches.