GLM 5
z-ai/glm-5
Compare with GLM 5 Turbo, GLM 5V Turbo, GLM 5.1
CompareGLM 5
| Provider | State | Uptime 5m | 24h | p50 latency | p99 | Throughput | Context | $/Mtok |
|---|---|---|---|---|---|---|---|---|
Amazon Bedrockamazon-bedrock | Healthy | 100% | 92.69% | 1.30s | 4.89s | 98 tps | 203K | $1.00 / $3.20 |
AtlasCloudatlas-cloud/fp8fp8 | Healthy | 100% | 99.55% | 2.41s | 10.3s | 38 tps | 203K | $0.950 / $3.15 |
Novitanovita/fp8fp8 | Healthy | 100% | 99.94% | 1.64s | 10.3s | 56 tps | 203K | $1.00 / $3.20 |
SiliconFlowsiliconflow/fp8fp8 | Healthy | 100% | 99.76% | 2.58s | 13.2s | 40 tps | 205K | $0.950 / $2.55 |
StreamLakestreamlake/fp8fp8 | Healthy | 100% | 99.92% | 2.51s | 7.49s | 37 tps | 198K | $0.600 / $1.92 |
Z.AIz-ai/fp8fp8 | Healthy | 100% | 99.37% | 8.22s | 12.3s | 53 tps | 203K | $1.00 / $3.20 |
DigitalOceandigitalocean | Healthy | 98.00% | 93.53% | 1.19s | 16.0s | 7.0 tps | 64K | $0.750 / $2.40 |
Baidubaidu/fp8fp8 | Healthy | 97.67% | 99.90% | 1.06s | 11.9s | 36 tps | 203K | $0.700 / $2.24 |
GMICloudgmicloud/fp8fp8 | Healthy | 97.67% | 99.51% | 1.28s | 6.88s | 64 tps | 203K | $0.600 / $1.92 |
DeepInfradeepinfra/fp4fp4 | No data | — | 98.85% | 816ms | 6.66s | 40 tps | 203K | $0.600 / $2.08 |
Venicevenice/fp8fp8 | No data | — | 93.53% | 1.03s | 15.4s | 42 tps | 198K | $1.00 / $3.20 |
- Amazon Bedrock100%Healthy24h 92.69%
- p50
- 1.30s
- p99
- 4.89s
- Throughput
- 98 tps
- Context
- 203K
- $/Mtok
- $1.00 / $3.20
amazon-bedrock - AtlasCloud100%Healthy24h 99.55%
- p50
- 2.41s
- p99
- 10.3s
- Throughput
- 38 tps
- Context
- 203K
- $/Mtok
- $0.950 / $3.15
atlas-cloud/fp8fp8 - Novita100%Healthy24h 99.94%
- p50
- 1.64s
- p99
- 10.3s
- Throughput
- 56 tps
- Context
- 203K
- $/Mtok
- $1.00 / $3.20
novita/fp8fp8 - SiliconFlow100%Healthy24h 99.76%
- p50
- 2.58s
- p99
- 13.2s
- Throughput
- 40 tps
- Context
- 205K
- $/Mtok
- $0.950 / $2.55
siliconflow/fp8fp8 - StreamLake100%Healthy24h 99.92%
- p50
- 2.51s
- p99
- 7.49s
- Throughput
- 37 tps
- Context
- 198K
- $/Mtok
- $0.600 / $1.92
streamlake/fp8fp8 - Z.AI100%Healthy24h 99.37%
- p50
- 8.22s
- p99
- 12.3s
- Throughput
- 53 tps
- Context
- 203K
- $/Mtok
- $1.00 / $3.20
z-ai/fp8fp8 - DigitalOcean98.00%Healthy24h 93.53%
- p50
- 1.19s
- p99
- 16.0s
- Throughput
- 7.0 tps
- Context
- 64K
- $/Mtok
- $0.750 / $2.40
digitalocean - Baidu97.67%Healthy24h 99.90%
- p50
- 1.06s
- p99
- 11.9s
- Throughput
- 36 tps
- Context
- 203K
- $/Mtok
- $0.700 / $2.24
baidu/fp8fp8 - GMICloud97.67%Healthy24h 99.51%
- p50
- 1.28s
- p99
- 6.88s
- Throughput
- 64 tps
- Context
- 203K
- $/Mtok
- $0.600 / $1.92
gmicloud/fp8fp8 - DeepInfra—No data24h 98.85%
- p50
- 816ms
- p99
- 6.66s
- Throughput
- 40 tps
- Context
- 203K
- $/Mtok
- $0.600 / $2.08
deepinfra/fp4fp4 - Venice—No data24h 93.53%
- p50
- 1.03s
- p99
- 15.4s
- Throughput
- 42 tps
- Context
- 198K
- $/Mtok
- $1.00 / $3.20
venice/fp8fp8
Measured 3m ago. Latency and throughput are OpenRouter's rolling 30-minute windows.
History
2 endpoints reporting no data are omitted from the charts.
Availability
Axis starts at 90% so small differences stay visible; it expands to 0 when an endpoint falls below.
Time to first token
Logarithmic axis — the fleet spans two orders of magnitude. p50 and p99 are OpenRouter's rolling 30-minute windows, averaged over the hour.
Throughput
Median tokens per second.