Billing and Pricing
Actual cost depends on the model, input and output usage, cache behavior, group multiplier, and media or task parameters. Use the live Model Catalog and usage logs as the source of truth.
| Dimension | Meaning |
|---|---|
| Input | Prompts, history, documents, and tool definitions sent to the model |
| Output | Generated text, reasoning, or structured content |
| Cache | Eligible cached-input reads or writes |
| Group multiplier | Adjustment applied by the key's selected group |
| Per-operation media | Some image or task endpoints charge by count, duration, or size |
Costs change as chat history grows, agents make multiple calls, model prices differ, groups change, and cache hits vary. Test a small but realistic workload rather than estimating an agent from one short chat.
Control spending
- Use separate keys for development, staging, production, and each desktop client.
- Start every new key with a small quota.
- Restrict expensive or unneeded models.
- Limit conversation history and tool-result size.
- Monitor usage by key and model.
- Add retry limits and circuit breakers for loops.
The account balance funds all keys. A key quota is an extra per-key cap. An unlimited key cannot spend past the account balance, and an expired or disabled key cannot call the API.
Prices and groups can change. Do not hard-code documentation prices into customer-facing billing logic without a clear update mechanism and timestamp.
