modelstatus.dev

GLM 5

z-ai/glm-5

Compare with GLM 5 Turbo, GLM 5V Turbo, GLM 5.1

CompareGLM 5
  • AtlasCloud100%
    Healthy24h 99.56%
    p50
    1.46s
    p99
    5.48s
    Throughput
    49 tps
    Context
    203K
    $/Mtok
    $0.950 / $3.15
    atlas-cloud/fp8fp8
  • Baidu100%
    Healthy24h 99.91%
    p50
    1.43s
    p99
    4.48s
    Throughput
    45 tps
    Context
    203K
    $/Mtok
    $0.700 / $2.24
    baidu/fp8fp8
  • DeepInfra100%
    Healthy24h 98.83%
    p50
    768ms
    p99
    12.6s
    Throughput
    45 tps
    Context
    203K
    $/Mtok
    $0.600 / $2.08
    deepinfra/fp4fp4
  • GMICloud100%
    Healthy24h 99.51%
    p50
    1.58s
    p99
    7.87s
    Throughput
    57 tps
    Context
    203K
    $/Mtok
    $0.600 / $1.92
    gmicloud/fp8fp8
  • Novita100%
    Healthy24h 99.95%
    p50
    1.48s
    p99
    10.2s
    Throughput
    59 tps
    Context
    203K
    $/Mtok
    $1.00 / $3.20
    novita/fp8fp8
  • StreamLake100%
    Healthy24h 99.92%
    p50
    1.91s
    p99
    5.29s
    Throughput
    54 tps
    Context
    198K
    $/Mtok
    $0.600 / $1.92
    streamlake/fp8fp8
  • Venice100%
    Degraded24h 93.16%
    p50
    6.29s
    p99
    20.1s
    Throughput
    25 tps
    Context
    198K
    $/Mtok
    $1.00 / $3.20
    venice/fp8fp8
  • Z.AI99.09%
    Healthy24h 99.40%
    p50
    7.91s
    p99
    12.1s
    Throughput
    55 tps
    Context
    203K
    $/Mtok
    $1.00 / $3.20
    z-ai/fp8fp8
  • Amazon Bedrock98.82%
    Healthy24h 93.06%
    p50
    945ms
    p99
    3.40s
    Throughput
    101 tps
    Context
    203K
    $/Mtok
    $1.00 / $3.20
    amazon-bedrock
  • SiliconFlow87.18%
    Down24h 99.75%
    p50
    1.88s
    p99
    15.4s
    Throughput
    47 tps
    Context
    205K
    $/Mtok
    $0.950 / $2.55
    siliconflow/fp8fp8
  • DigitalOcean
    Down24h 93.15%
    p50
    5.29s
    p99
    24.0s
    Throughput
    4.0 tps
    Context
    64K
    $/Mtok
    $0.750 / $2.40
    digitalocean

Measured 2m ago. Latency and throughput are OpenRouter's rolling 30-minute windows.

Availability

Axis starts at 90% so small differences stay visible; it expands to 0 when an endpoint falls below.

Time to first token

Logarithmic axis — the fleet spans two orders of magnitude. p50 and p99 are OpenRouter's rolling 30-minute windows.

Throughput

Median tokens per second.