GLM 4.6
z-ai/glm-4.6
Compare with GLM 4.6V, GLM 4.5 Air, GLM 4.5V
CompareGLM 4.6
| State | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Healthy | 100% | 99.92% | 99.16% | 99.16% | 524ms | 10.3s | 42 tps | 203K | $0.500 / $2.00 | |
| Healthy | 100% | 99.34% | 98.74% | 98.74% | 1.12s | 19.6s | 27 tps | 198K | $0.430 / $1.75 | |
| Healthy | 96.00% | 97.26% | 96.76% | 96.76% | 2.02s | 27.7s | 27 tps | 205K | $0.550 / $2.20 | |
| No data | — | 92.69% | 97.11% | 97.11% | 905ms | 11.9s | 47 tps | 203K | $0.600 / $2.20 | |
| No data | — | 95.76% | 91.74% | 91.74% | 1.01s | 23.5s | 28 tps | 203K | $0.600 / $2.20 |
- DeepInfra100%Healthy24h 99.92%
- 7d
- 99.16%
- 30d
- 99.16%
- p50
- 524ms
- p99
- 10.3s
- Throughput
- 42 tps
- Context
- 203K
- $/Mtok
- $0.500 / $2.00
deepinfra/fp4fp4 - Venice100%Healthy24h 99.34%
- 7d
- 98.74%
- 30d
- 98.74%
- p50
- 1.12s
- p99
- 19.6s
- Throughput
- 27 tps
- Context
- 198K
- $/Mtok
- $0.430 / $1.75
venice/fp4fp4 - Novita96.00%Healthy24h 97.26%
- 7d
- 96.76%
- 30d
- 96.76%
- p50
- 2.02s
- p99
- 27.7s
- Throughput
- 27 tps
- Context
- 205K
- $/Mtok
- $0.550 / $2.20
novita/bf16bf16 - No data24h 92.69%
- 7d
- 97.11%
- 30d
- 97.11%
- p50
- 905ms
- p99
- 11.9s
- Throughput
- 47 tps
- Context
- 203K
- $/Mtok
- $0.600 / $2.20
atlas-cloud/fp8fp8 - Z.AI—No data24h 95.76%
- 7d
- 91.74%
- 30d
- 91.74%
- p50
- 1.01s
- p99
- 23.5s
- Throughput
- 28 tps
- Context
- 203K
- $/Mtok
- $0.600 / $2.20
z-ai/fp4fp4
Measured 5m ago. Latency and throughput are OpenRouter's rolling 30-minute windows.
History
2 endpoints reporting no data are omitted from the charts.
Availability
Axis starts at 90% so small differences stay visible; it expands to 0 when an endpoint falls below.
Time to first token
Logarithmic axis — the fleet spans two orders of magnitude. P50 is OpenRouter's rolling 30-minute window.
Throughput
Median tokens per second.