engyMiniMax H3 is now live!

Pricing

Per-token, pay as you go. No subscriptions and no minimums; you pay only for the tokens you use.

modelinput↑outputcachedvs officialctx
MiniMax-H3new
$0.03/second—32K
deepseek-v4.1-flashnewZDR
$0.30$0.04$1.20$0.08$0.006$0.008−90%328K
deepseek-v4-flash-0731ZDR
$0.045$0.09$0.009—1M
qwen3.6-35b-a3b
$0.375$0.045$2.25$0.30$0.015−87%208K
qwen3.8-27bZDR
$0.50$0.045$3.00$0.32$0.015−90%1M
glm-5.3-flash(Ox Alpha)newZDR
$0.15$0.135$0.50$0.45$0.03$0.027−10%262K
glm-5.2ZDR
$1.40$0.68$4.40$1.50$0.26$0.18−59%262K
glm-5.3newZDR
$1.40$0.98$4.40$3.08$0.26$0.18−30%328K
kimi-k3ZDR
$3.00$1.95$15.00$9.75$0.30$0.195−35%1M
$ per 1M tokens (/second = $ per second of output) · $1.40 = the model developer's own API price, checked 2026-09-29 · vs official uses a 3:1 input:output blend

Prompt-cache hits bill at the cached rate automatically, with no config and no cache_control markers. Agentic workloads (coding assistants, multi-turn tools) typically hit 90%+ cache on repeated prefixes, so effective input cost is usually far below the headline rate.

ZDR marks models under zero data retention: served only on hardware engy operates, prompts and outputs never stored or trained on. MiniMax-H3 (not covered by engy's zero-data-retention configuration) and qwen3.6-35b-a3b (routed across permissionless subnet miners, on hardware without confidential computing) carry no such mark; see terms and privacy.

Prices are live from the billing engine. The API reports the same numbers at https://api.engy.ai/v1/models.