GLM 5.1
z-ai/glm-5.1
Compare with GLM 5, GLM 5 Turbo, GLM 5V Turbo
CompareGLM 5.1
| Provider | State | Uptime 5m | 24h | p50 latency | p99 | Throughput | Context | $/Mtok |
|---|---|---|---|---|---|---|---|---|
Baidubaidu/fp8fp8 | Degraded | 100% | 99.40% | 1.04s | 3.57s | 62 tps | 203K | $0.896 / $2.82 |
Crusoecrusoe/fp8fp8 | Degraded | 100% | 99.54% | 802ms | 8.76s | 69 tps | 203K | $1.20 / $4.40 |
Friendlifriendli | Healthy | 100% | 99.76% | 260ms | 8.42s | 49 tps | 203K | $1.40 / $4.40 |
GMICloudgmicloud/fp8fp8 | Healthy | 100% | 98.79% | 2.98s | 6.92s | 79 tps | 203K | $0.910 / $2.86 |
Nebiusnebius/fp8fp8 | Degraded | 100% | 94.79% | 905ms | 20.3s | 31 tps | 203K | $1.40 / $4.40 |
SiliconFlowsiliconflow/fp8fp8 | Degraded | 100% | 99.47% | 1.87s | 9.21s | 40 tps | 205K | $1.19 / $3.74 |
StreamLakestreamlake/fp8fp8 | Degraded | 100% | 98.08% | 3.86s | 19.8s | 22 tps | 200K | $0.966 / $3.04 |
Z.AIz-ai/fp8fp8 | Degraded | 100% | 98.47% | 8.52s | 12.8s | 22 tps | 203K | $1.40 / $4.40 |
AtlasCloudatlas-cloud/fp8fp8 | Healthy | 99.31% | 99.33% | 1.41s | 12.6s | 54 tps | 203K | $1.26 / $3.96 |
DeepInfradeepinfra/fp4fp4 | Healthy | 97.46% | 97.92% | 1.24s | 460s | 30 tps | 203K | $1.05 / $3.50 |
Chuteschutes/fp8fp8 | Down | 72.06% | 87.40% | 5.18s | 119s | 26 tps | 203K | $0.980 / $3.08 |
Alibabaalibaba/fp8fp8 | Degraded | — | 99.15% | 2.02s | 10.0s | 46 tps | 203K | $1.33 / $4.18 |
DigitalOceandigitalocean | Down | — | 90.34% | 1.48s | 38.6s | 13 tps | 164K | $0.975 / $4.30 |
Novitanovita/fp8fp8 | Down | — | 98.97% | 2.70s | 35.5s | 35 tps | 205K | $1.38 / $4.40 |
Parasailparasail/fp8fp8 | No data | — | 99.62% | 581ms | 3.00s | 157 tps | 203K | $1.40 / $4.40 |
Phalaphala | Down | — | 86.83% | 6.15s | 687s | 22 tps | 203K | $1.21 / $4.20 |
Venicevenice/fp8fp8 | Down | — | 90.77% | 1.57s | 39.0s | 35 tps | 200K | $1.54 / $4.84 |
- Baidu100%Degraded24h 99.40%
- p50
- 1.04s
- p99
- 3.57s
- Throughput
- 62 tps
- Context
- 203K
- $/Mtok
- $0.896 / $2.82
baidu/fp8fp8 - Crusoe100%Degraded24h 99.54%
- p50
- 802ms
- p99
- 8.76s
- Throughput
- 69 tps
- Context
- 203K
- $/Mtok
- $1.20 / $4.40
crusoe/fp8fp8 - Friendli100%Healthy24h 99.76%
- p50
- 260ms
- p99
- 8.42s
- Throughput
- 49 tps
- Context
- 203K
- $/Mtok
- $1.40 / $4.40
friendli - GMICloud100%Healthy24h 98.79%
- p50
- 2.98s
- p99
- 6.92s
- Throughput
- 79 tps
- Context
- 203K
- $/Mtok
- $0.910 / $2.86
gmicloud/fp8fp8 - Nebius100%Degraded24h 94.79%
- p50
- 905ms
- p99
- 20.3s
- Throughput
- 31 tps
- Context
- 203K
- $/Mtok
- $1.40 / $4.40
nebius/fp8fp8 - SiliconFlow100%Degraded24h 99.47%
- p50
- 1.87s
- p99
- 9.21s
- Throughput
- 40 tps
- Context
- 205K
- $/Mtok
- $1.19 / $3.74
siliconflow/fp8fp8 - StreamLake100%Degraded24h 98.08%
- p50
- 3.86s
- p99
- 19.8s
- Throughput
- 22 tps
- Context
- 200K
- $/Mtok
- $0.966 / $3.04
streamlake/fp8fp8 - Z.AI100%Degraded24h 98.47%
- p50
- 8.52s
- p99
- 12.8s
- Throughput
- 22 tps
- Context
- 203K
- $/Mtok
- $1.40 / $4.40
z-ai/fp8fp8 - AtlasCloud99.31%Healthy24h 99.33%
- p50
- 1.41s
- p99
- 12.6s
- Throughput
- 54 tps
- Context
- 203K
- $/Mtok
- $1.26 / $3.96
atlas-cloud/fp8fp8 - DeepInfra97.46%Healthy24h 97.92%
- p50
- 1.24s
- p99
- 460s
- Throughput
- 30 tps
- Context
- 203K
- $/Mtok
- $1.05 / $3.50
deepinfra/fp4fp4 - Chutes72.06%Down24h 87.40%
- p50
- 5.18s
- p99
- 119s
- Throughput
- 26 tps
- Context
- 203K
- $/Mtok
- $0.980 / $3.08
chutes/fp8fp8 - Alibaba—Degraded24h 99.15%
- p50
- 2.02s
- p99
- 10.0s
- Throughput
- 46 tps
- Context
- 203K
- $/Mtok
- $1.33 / $4.18
alibaba/fp8fp8 - DigitalOcean—Down24h 90.34%
- p50
- 1.48s
- p99
- 38.6s
- Throughput
- 13 tps
- Context
- 164K
- $/Mtok
- $0.975 / $4.30
digitalocean - Novita—Down24h 98.97%
- p50
- 2.70s
- p99
- 35.5s
- Throughput
- 35 tps
- Context
- 205K
- $/Mtok
- $1.38 / $4.40
novita/fp8fp8 - Parasail—No data24h 99.62%
- p50
- 581ms
- p99
- 3.00s
- Throughput
- 157 tps
- Context
- 203K
- $/Mtok
- $1.40 / $4.40
parasail/fp8fp8 - Phala—Down24h 86.83%
- p50
- 6.15s
- p99
- 687s
- Throughput
- 22 tps
- Context
- 203K
- $/Mtok
- $1.21 / $4.20
phala - Venice—Down24h 90.77%
- p50
- 1.57s
- p99
- 39.0s
- Throughput
- 35 tps
- Context
- 200K
- $/Mtok
- $1.54 / $4.84
venice/fp8fp8
Measured 4m ago. Latency and throughput are OpenRouter's rolling 30-minute windows.
History
1 endpoint reporting no data is omitted from the charts.
Availability
Axis starts at 90% so small differences stay visible; it expands to 0 when an endpoint falls below.
Time to first token
Logarithmic axis — the fleet spans two orders of magnitude. p50 and p99 are OpenRouter's rolling 30-minute windows, averaged over the hour.
Throughput
Median tokens per second.