Model record / Z.ai
Access glm-5.3-flash from Z.ai through one OpenAI-compatible API with live market-linked pricing and usage-based billing.
Observed reliability
100
Output speed
—
p95 latency
28,515
Evidence strength
—
See how much this model saves and when its recorded price has been lowest.
Savings are visible now, but a best-time recommendation needs at least 12 observations across 6 different hours.
Average saved
30.0%
Cheapest time
Collecting data
Hourly coverage
2/24
Signal quality
Early
Hover a bar for savings per token and observation counts.
Source buckets are UTC; labels are converted to UTC.
Last 24 hours
#61
of 61 ranked models
Last 7 days
#61
of 61 ranked models
Last 30 days
#61
of 61 ranked models
Ready when you are
Marketplace pricing, capped at list, optimized further.
curl https://api.cheaperinference.com/v1/chat/completions \ -H "Authorization: Bearer ci_live_YOUR_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "glm-5.3-flash", "messages": [{"role": "user", "content": "Hello!"}] }'