modelstatus.dev

GLM 4.7

z-ai/glm-4.7

Compare with GLM 5.2, GLM 5.1, GLM 5

  • DeepInfra100%
    Healthy24h 98.90%
    p50
    1.06s
    p99
    8.03s
    Throughput
    23 tps
    Context
    203K
    $/Mtok
    $0.400 / $1.75
    deepinfra/fp4fp4
  • Google100%
    Healthy24h 99.90%
    p50
    494ms
    p99
    9.77s
    Throughput
    97 tps
    Context
    200K
    $/Mtok
    $0.600 / $2.20
    google-vertex
  • Mancer 2100%
    Healthy24h 96.45%
    p50
    1.28s
    p99
    9.08s
    Throughput
    17 tps
    Context
    131K
    $/Mtok
    $0.600 / $2.50
    mancer/fp4fp4
  • Venice100%
    Healthy24h 98.01%
    p50
    1.44s
    p99
    5.88s
    Throughput
    23 tps
    Context
    198K
    $/Mtok
    $0.550 / $2.65
    venice/fp4fp4
  • AtlasCloud97.83%
    Degraded24h 96.92%
    p50
    1.94s
    p99
    24.6s
    Throughput
    27 tps
    Context
    203K
    $/Mtok
    $0.520 / $1.85
    atlas-cloud/fp8fp8
  • Novita93.52%
    Degraded24h 93.56%
    p50
    2.39s
    p99
    35.5s
    Throughput
    24 tps
    Context
    205K
    $/Mtok
    $0.540 / $1.98
    novita/fp8fp8
  • StreamLake93.15%
    Degraded24h 92.38%
    p50
    1.96s
    p99
    25.1s
    Throughput
    30 tps
    Context
    200K
    $/Mtok
    $0.480 / $1.76
    streamlake/fp8fp8
  • Z.AI88.66%
    Down24h 82.45%
    p50
    16.0s
    p99
    39.4s
    Throughput
    26 tps
    Context
    203K
    $/Mtok
    $0.600 / $2.20
    z-ai/fp4fp4

Measured 2m ago. Latency and throughput are OpenRouter's rolling 30-minute windows.

Availability

Axis starts at 90% so small differences stay visible; it expands to 0 when an endpoint falls below.

Time to first token

Logarithmic axis — the fleet spans two orders of magnitude. p50 and p99 are OpenRouter's rolling 30-minute windows.

Throughput

Median tokens per second.