Pricing
Per-token, pay as you go. No subscriptions and no minimums; you pay only for the tokens you use.
| model | input↑ | output | cached | ctx |
|---|---|---|---|---|
| deepseek-v4-flash-0731ZDR | $0.045 | $0.09 | $0.009 | 1M |
| qwen3.6-35b-a3b | $0.045 | $0.30 | $0.015 | 208K |
| qwen3.8-27bZDR | $0.045 | $0.32 | $0.015 | 1M |
| glm-5.3-flash(Ox Alpha)newZDR | $0.135 | $0.45 | $0.027 | 262K |
| glm-5.2ZDR | $0.68 | $1.50 | $0.18 | 262K |
| glm-5.3newZDR | $0.98 | $3.08 | $0.18 | 328K |
| kimi-k3ZDR | $1.95 | $9.75 | $0.195 | 1M |
| $ per 1M tokens | ||||
Prompt-cache hits bill at the cached rate automatically, with no config and no cache_control markers. Agentic workloads (coding assistants, multi-turn tools) typically hit 90%+ cache on repeated prefixes, so effective input cost is usually far below the headline rate.
ZDR marks models under zero data retention: served only on hardware engy operates, prompts and outputs never stored or trained on. qwen3.6-35b-a3b carries no such mark because it routes across independent operators; see terms and privacy.
Prices are live from the billing engine. The API reports the same numbers at https://api.engy.ai/v1/models.