GLM 5.1
z-ai/glm-5.1
Compare with GLM 5, GLM 5 Turbo, GLM 5V Turbo
CompareGLM 5.1
| Provider | State | Uptime 5m | 24h | p50 latency | p99 | Throughput | Context | $/Mtok |
|---|---|---|---|---|---|---|---|---|
Alibabaalibaba/fp8fp8 | Down | 100% | 99.12% | 2.02s | 8.24s | 45 tps | 203K | $1.33 / $4.18 |
Baidubaidu/fp8fp8 | Degraded | 100% | 99.37% | 1.05s | 3.55s | 64 tps | 203K | $0.896 / $2.82 |
Crusoecrusoe/fp8fp8 | Degraded | 100% | 99.53% | 795ms | 8.85s | 69 tps | 203K | $1.20 / $4.40 |
DeepInfradeepinfra/fp4fp4 | Healthy | 100% | 97.95% | 1.25s | 581s | 32 tps | 203K | $1.05 / $3.50 |
Friendlifriendli | Healthy | 100% | 99.75% | 375ms | 7.52s | 42 tps | 203K | $1.40 / $4.40 |
GMICloudgmicloud/fp8fp8 | Healthy | 100% | 98.79% | 2.97s | 7.23s | 78 tps | 203K | $0.910 / $2.86 |
Parasailparasail/fp8fp8 | Healthy | 100% | 99.62% | 572ms | 2.92s | 147 tps | 203K | $1.40 / $4.40 |
SiliconFlowsiliconflow/fp8fp8 | Degraded | 100% | 99.46% | 1.91s | 9.21s | 40 tps | 205K | $1.19 / $3.74 |
Z.AIz-ai/fp8fp8 | Degraded | 100% | 98.46% | 8.25s | 12.4s | 22 tps | 203K | $1.40 / $4.40 |
AtlasCloudatlas-cloud/fp8fp8 | Degraded | 99.39% | 99.32% | 1.43s | 11.4s | 54 tps | 203K | $1.26 / $3.96 |
StreamLakestreamlake/fp8fp8 | Degraded | 98.82% | 98.10% | 3.90s | 19.9s | 23 tps | 200K | $0.966 / $3.04 |
Chuteschutes/fp8fp8 | Down | 77.08% | 87.35% | 3.88s | 114s | 26 tps | 203K | $0.980 / $3.08 |
DigitalOceandigitalocean | Down | — | 90.26% | 1.53s | 39.4s | 11 tps | 164K | $0.975 / $4.30 |
Nebiusnebius/fp8fp8 | Degraded | — | 94.73% | 934ms | 18.8s | 32 tps | 203K | $1.40 / $4.40 |
Novitanovita/fp8fp8 | No data | — | 98.95% | 2.67s | 34.7s | 37 tps | 205K | $1.38 / $4.40 |
Phalaphala | Down | — | 86.82% | 6.28s | 847s | 24 tps | 203K | $1.21 / $4.20 |
Venicevenice/fp8fp8 | Down | — | 90.81% | 1.75s | 43.8s | 33 tps | 200K | $1.54 / $4.84 |
- Alibaba100%Down24h 99.12%
- p50
- 2.02s
- p99
- 8.24s
- Throughput
- 45 tps
- Context
- 203K
- $/Mtok
- $1.33 / $4.18
alibaba/fp8fp8 - Baidu100%Degraded24h 99.37%
- p50
- 1.05s
- p99
- 3.55s
- Throughput
- 64 tps
- Context
- 203K
- $/Mtok
- $0.896 / $2.82
baidu/fp8fp8 - Crusoe100%Degraded24h 99.53%
- p50
- 795ms
- p99
- 8.85s
- Throughput
- 69 tps
- Context
- 203K
- $/Mtok
- $1.20 / $4.40
crusoe/fp8fp8 - DeepInfra100%Healthy24h 97.95%
- p50
- 1.25s
- p99
- 581s
- Throughput
- 32 tps
- Context
- 203K
- $/Mtok
- $1.05 / $3.50
deepinfra/fp4fp4 - Friendli100%Healthy24h 99.75%
- p50
- 375ms
- p99
- 7.52s
- Throughput
- 42 tps
- Context
- 203K
- $/Mtok
- $1.40 / $4.40
friendli - GMICloud100%Healthy24h 98.79%
- p50
- 2.97s
- p99
- 7.23s
- Throughput
- 78 tps
- Context
- 203K
- $/Mtok
- $0.910 / $2.86
gmicloud/fp8fp8 - Parasail100%Healthy24h 99.62%
- p50
- 572ms
- p99
- 2.92s
- Throughput
- 147 tps
- Context
- 203K
- $/Mtok
- $1.40 / $4.40
parasail/fp8fp8 - SiliconFlow100%Degraded24h 99.46%
- p50
- 1.91s
- p99
- 9.21s
- Throughput
- 40 tps
- Context
- 205K
- $/Mtok
- $1.19 / $3.74
siliconflow/fp8fp8 - Z.AI100%Degraded24h 98.46%
- p50
- 8.25s
- p99
- 12.4s
- Throughput
- 22 tps
- Context
- 203K
- $/Mtok
- $1.40 / $4.40
z-ai/fp8fp8 - AtlasCloud99.39%Degraded24h 99.32%
- p50
- 1.43s
- p99
- 11.4s
- Throughput
- 54 tps
- Context
- 203K
- $/Mtok
- $1.26 / $3.96
atlas-cloud/fp8fp8 - StreamLake98.82%Degraded24h 98.10%
- p50
- 3.90s
- p99
- 19.9s
- Throughput
- 23 tps
- Context
- 200K
- $/Mtok
- $0.966 / $3.04
streamlake/fp8fp8 - Chutes77.08%Down24h 87.35%
- p50
- 3.88s
- p99
- 114s
- Throughput
- 26 tps
- Context
- 203K
- $/Mtok
- $0.980 / $3.08
chutes/fp8fp8 - DigitalOcean—Down24h 90.26%
- p50
- 1.53s
- p99
- 39.4s
- Throughput
- 11 tps
- Context
- 164K
- $/Mtok
- $0.975 / $4.30
digitalocean - Nebius—Degraded24h 94.73%
- p50
- 934ms
- p99
- 18.8s
- Throughput
- 32 tps
- Context
- 203K
- $/Mtok
- $1.40 / $4.40
nebius/fp8fp8 - Novita—No data24h 98.95%
- p50
- 2.67s
- p99
- 34.7s
- Throughput
- 37 tps
- Context
- 205K
- $/Mtok
- $1.38 / $4.40
novita/fp8fp8 - Phala—Down24h 86.82%
- p50
- 6.28s
- p99
- 847s
- Throughput
- 24 tps
- Context
- 203K
- $/Mtok
- $1.21 / $4.20
phala - Venice—Down24h 90.81%
- p50
- 1.75s
- p99
- 43.8s
- Throughput
- 33 tps
- Context
- 200K
- $/Mtok
- $1.54 / $4.84
venice/fp8fp8
Measured 4m ago. Latency and throughput are OpenRouter's rolling 30-minute windows.
History
1 endpoint reporting no data is omitted from the charts.
Availability
Axis starts at 90% so small differences stay visible; it expands to 0 when an endpoint falls below.
Time to first token
Logarithmic axis — the fleet spans two orders of magnitude. p50 and p99 are OpenRouter's rolling 30-minute windows.
Throughput
Median tokens per second.