GLM 5
z-ai/glm-5
Compare with GLM 5 Turbo, GLM 5V Turbo, GLM 5.1
CompareGLM 5
| Provider | State | Uptime 5m | 24h | p50 latency | p99 | Throughput | Context | $/Mtok |
|---|---|---|---|---|---|---|---|---|
AtlasCloudatlas-cloud/fp8fp8 | Healthy | 100% | 99.56% | 1.46s | 5.48s | 49 tps | 203K | $0.950 / $3.15 |
Baidubaidu/fp8fp8 | Healthy | 100% | 99.91% | 1.43s | 4.48s | 45 tps | 203K | $0.700 / $2.24 |
DeepInfradeepinfra/fp4fp4 | Healthy | 100% | 98.83% | 768ms | 12.6s | 45 tps | 203K | $0.600 / $2.08 |
GMICloudgmicloud/fp8fp8 | Healthy | 100% | 99.51% | 1.58s | 7.87s | 57 tps | 203K | $0.600 / $1.92 |
Novitanovita/fp8fp8 | Healthy | 100% | 99.95% | 1.48s | 10.2s | 59 tps | 203K | $1.00 / $3.20 |
StreamLakestreamlake/fp8fp8 | Healthy | 100% | 99.92% | 1.91s | 5.29s | 54 tps | 198K | $0.600 / $1.92 |
Venicevenice/fp8fp8 | Degraded | 100% | 93.16% | 6.29s | 20.1s | 25 tps | 198K | $1.00 / $3.20 |
Z.AIz-ai/fp8fp8 | Healthy | 99.09% | 99.40% | 7.91s | 12.1s | 55 tps | 203K | $1.00 / $3.20 |
Amazon Bedrockamazon-bedrock | Healthy | 98.82% | 93.06% | 945ms | 3.40s | 101 tps | 203K | $1.00 / $3.20 |
SiliconFlowsiliconflow/fp8fp8 | Down | 87.18% | 99.75% | 1.88s | 15.4s | 47 tps | 205K | $0.950 / $2.55 |
DigitalOceandigitalocean | Down | — | 93.15% | 5.29s | 24.0s | 4.0 tps | 64K | $0.750 / $2.40 |
- AtlasCloud100%Healthy24h 99.56%
- p50
- 1.46s
- p99
- 5.48s
- Throughput
- 49 tps
- Context
- 203K
- $/Mtok
- $0.950 / $3.15
atlas-cloud/fp8fp8 - Baidu100%Healthy24h 99.91%
- p50
- 1.43s
- p99
- 4.48s
- Throughput
- 45 tps
- Context
- 203K
- $/Mtok
- $0.700 / $2.24
baidu/fp8fp8 - DeepInfra100%Healthy24h 98.83%
- p50
- 768ms
- p99
- 12.6s
- Throughput
- 45 tps
- Context
- 203K
- $/Mtok
- $0.600 / $2.08
deepinfra/fp4fp4 - GMICloud100%Healthy24h 99.51%
- p50
- 1.58s
- p99
- 7.87s
- Throughput
- 57 tps
- Context
- 203K
- $/Mtok
- $0.600 / $1.92
gmicloud/fp8fp8 - Novita100%Healthy24h 99.95%
- p50
- 1.48s
- p99
- 10.2s
- Throughput
- 59 tps
- Context
- 203K
- $/Mtok
- $1.00 / $3.20
novita/fp8fp8 - StreamLake100%Healthy24h 99.92%
- p50
- 1.91s
- p99
- 5.29s
- Throughput
- 54 tps
- Context
- 198K
- $/Mtok
- $0.600 / $1.92
streamlake/fp8fp8 - Venice100%Degraded24h 93.16%
- p50
- 6.29s
- p99
- 20.1s
- Throughput
- 25 tps
- Context
- 198K
- $/Mtok
- $1.00 / $3.20
venice/fp8fp8 - Z.AI99.09%Healthy24h 99.40%
- p50
- 7.91s
- p99
- 12.1s
- Throughput
- 55 tps
- Context
- 203K
- $/Mtok
- $1.00 / $3.20
z-ai/fp8fp8 - Amazon Bedrock98.82%Healthy24h 93.06%
- p50
- 945ms
- p99
- 3.40s
- Throughput
- 101 tps
- Context
- 203K
- $/Mtok
- $1.00 / $3.20
amazon-bedrock - SiliconFlow87.18%Down24h 99.75%
- p50
- 1.88s
- p99
- 15.4s
- Throughput
- 47 tps
- Context
- 205K
- $/Mtok
- $0.950 / $2.55
siliconflow/fp8fp8 - DigitalOcean—Down24h 93.15%
- p50
- 5.29s
- p99
- 24.0s
- Throughput
- 4.0 tps
- Context
- 64K
- $/Mtok
- $0.750 / $2.40
digitalocean
Measured 2m ago. Latency and throughput are OpenRouter's rolling 30-minute windows.
History
Availability
Axis starts at 90% so small differences stay visible; it expands to 0 when an endpoint falls below.
Time to first token
Logarithmic axis — the fleet spans two orders of magnitude. p50 and p99 are OpenRouter's rolling 30-minute windows.
Throughput
Median tokens per second.