GLM 5
z-ai/glm-5
Compare with GLM 5 Turbo, GLM 5V Turbo, GLM 5.1
CompareGLM 5
| Provider | State | Uptime 5m | 24h | p50 latency | p99 | Throughput | Context | $/Mtok |
|---|---|---|---|---|---|---|---|---|
AtlasCloudatlas-cloud/fp8fp8 | Healthy | 100% | 99.56% | 1.96s | 5.35s | 39 tps | 203K | $0.950 / $3.15 |
Baidubaidu/fp8fp8 | Healthy | 100% | 99.90% | 1.06s | 2.30s | 51 tps | 203K | $0.700 / $2.24 |
DeepInfradeepinfra/fp4fp4 | Healthy | 100% | 98.88% | 842ms | 10.1s | 29 tps | 203K | $0.600 / $2.08 |
Novitanovita/fp8fp8 | Healthy | 100% | 99.94% | 1.52s | 8.18s | 57 tps | 203K | $1.00 / $3.20 |
Z.AIz-ai/fp8fp8 | Healthy | 100% | 99.35% | 8.32s | 11.6s | 52 tps | 203K | $1.00 / $3.20 |
StreamLakestreamlake/fp8fp8 | Healthy | 99.79% | 99.92% | 2.20s | 5.59s | 60 tps | 198K | $0.600 / $1.92 |
GMICloudgmicloud/fp8fp8 | Healthy | 98.48% | 99.52% | 1.59s | 8.68s | 48 tps | 203K | $0.600 / $1.92 |
Amazon Bedrockamazon-bedrock | Healthy | 98.25% | 92.59% | 1.19s | 3.73s | 104 tps | 203K | $1.00 / $3.20 |
DigitalOceandigitalocean | Down | 89.19% | 93.60% | 1.52s | 16.6s | 7.0 tps | 64K | $0.750 / $2.40 |
SiliconFlowsiliconflow/fp8fp8 | No data | — | 99.75% | 2.32s | 11.7s | 46 tps | 205K | $0.950 / $2.55 |
Venicevenice/fp8fp8 | No data | — | 93.61% | 1.22s | 7.38s | 29 tps | 198K | $1.00 / $3.20 |
- AtlasCloud100%Healthy24h 99.56%
- p50
- 1.96s
- p99
- 5.35s
- Throughput
- 39 tps
- Context
- 203K
- $/Mtok
- $0.950 / $3.15
atlas-cloud/fp8fp8 - Baidu100%Healthy24h 99.90%
- p50
- 1.06s
- p99
- 2.30s
- Throughput
- 51 tps
- Context
- 203K
- $/Mtok
- $0.700 / $2.24
baidu/fp8fp8 - DeepInfra100%Healthy24h 98.88%
- p50
- 842ms
- p99
- 10.1s
- Throughput
- 29 tps
- Context
- 203K
- $/Mtok
- $0.600 / $2.08
deepinfra/fp4fp4 - Novita100%Healthy24h 99.94%
- p50
- 1.52s
- p99
- 8.18s
- Throughput
- 57 tps
- Context
- 203K
- $/Mtok
- $1.00 / $3.20
novita/fp8fp8 - Z.AI100%Healthy24h 99.35%
- p50
- 8.32s
- p99
- 11.6s
- Throughput
- 52 tps
- Context
- 203K
- $/Mtok
- $1.00 / $3.20
z-ai/fp8fp8 - StreamLake99.79%Healthy24h 99.92%
- p50
- 2.20s
- p99
- 5.59s
- Throughput
- 60 tps
- Context
- 198K
- $/Mtok
- $0.600 / $1.92
streamlake/fp8fp8 - GMICloud98.48%Healthy24h 99.52%
- p50
- 1.59s
- p99
- 8.68s
- Throughput
- 48 tps
- Context
- 203K
- $/Mtok
- $0.600 / $1.92
gmicloud/fp8fp8 - Amazon Bedrock98.25%Healthy24h 92.59%
- p50
- 1.19s
- p99
- 3.73s
- Throughput
- 104 tps
- Context
- 203K
- $/Mtok
- $1.00 / $3.20
amazon-bedrock - DigitalOcean89.19%Down24h 93.60%
- p50
- 1.52s
- p99
- 16.6s
- Throughput
- 7.0 tps
- Context
- 64K
- $/Mtok
- $0.750 / $2.40
digitalocean - SiliconFlow—No data24h 99.75%
- p50
- 2.32s
- p99
- 11.7s
- Throughput
- 46 tps
- Context
- 205K
- $/Mtok
- $0.950 / $2.55
siliconflow/fp8fp8 - Venice—No data24h 93.61%
- p50
- 1.22s
- p99
- 7.38s
- Throughput
- 29 tps
- Context
- 198K
- $/Mtok
- $1.00 / $3.20
venice/fp8fp8
Measured 1m ago. Latency and throughput are OpenRouter's rolling 30-minute windows.
History
2 endpoints reporting no data are omitted from the charts.
Availability
Axis starts at 90% so small differences stay visible; it expands to 0 when an endpoint falls below.
Time to first token
Logarithmic axis — the fleet spans two orders of magnitude. p50 and p99 are OpenRouter's rolling 30-minute windows.
Throughput
Median tokens per second.