modelstatus.dev

GLM 5.1

GLM 5.1: providers compared

Every endpoint serving this model, side by side. ◆ marks the best value in each column — per model, never fleet-wide.

  • AtlasCloud100% (best)
    Healthy24h 99.68%
    7d
    99.72%
    30d
    99.72%
    p50
    1.49s
    p99
    5.49s
    Throughput
    54 tps
    Context
    203K
    $/Mtok
    $1.26 / $3.96
    atlas-cloud/fp8fp8
  • Crusoe100% (best)
    Healthy24h 99.74%
    7d
    99.45%
    30d
    99.45%
    p50
    656ms
    p99
    12.9s
    Throughput
    76 tps
    Context
    203K
    $/Mtok
    $1.20 / $4.40
    crusoe/fp8fp8
  • Friendli100% (best)
    Healthy24h 99.98%
    7d
    99.96% (best)
    30d
    99.96% (best)
    p50
    248ms (best)
    p99
    3.99s
    Throughput
    63 tps
    Context
    203K
    $/Mtok
    $1.40 / $4.40
    friendli
  • GMICloud100% (best)
    Healthy24h 99.62%
    7d
    99.59%
    30d
    99.59%
    p50
    1.81s
    p99
    15.3s
    Throughput
    62 tps
    Context
    203K
    $/Mtok
    $0.910 / $2.86
    gmicloud/fp8fp8
  • Nebius100% (best)
    Healthy24h 98.20%
    7d
    97.56%
    30d
    97.56%
    p50
    783ms
    p99
    12.0s
    Throughput
    28 tps
    Context
    203K
    $/Mtok
    $1.40 / $4.40
    nebius/fp8fp8
  • Parasail100% (best)
    Healthy24h 98.92%
    7d
    99.43%
    30d
    99.43%
    p50
    707ms
    p99
    11.0s
    Throughput
    96 tps (best)
    Context
    203K
    $/Mtok
    $1.40 / $4.40
    parasail/fp8fp8
  • SiliconFlow100% (best)
    Healthy24h 99.94%
    7d
    99.89%
    30d
    99.89%
    p50
    1.70s
    p99
    6.96s
    Throughput
    48 tps
    Context
    205K
    $/Mtok
    $1.19 / $3.74
    siliconflow/fp8fp8
  • StreamLake100% (best)
    Healthy24h 99.51%
    7d
    98.85%
    30d
    98.85%
    p50
    1.92s
    p99
    6.82s
    Throughput
    29 tps
    Context
    200K
    $/Mtok
    $0.966 / $3.04
    streamlake/fp8fp8
  • Z.AI100% (best)
    Healthy24h 99.92%
    7d
    99.11%
    30d
    99.11%
    p50
    4.06s
    p99
    6.69s
    Throughput
    23 tps
    Context
    203K
    $/Mtok
    $1.40 / $4.40
    z-ai/fp8fp8
  • Baidu99.65%
    Healthy24h 99.97%
    7d
    99.90%
    30d
    99.90%
    p50
    1.55s
    p99
    7.48s
    Throughput
    52 tps
    Context
    203K
    $/Mtok
    $0.896 / $2.82 (best)
    baidu/fp8fp8
  • Chutes75.44%
    Down24h 76.15%
    7d
    83.07%
    30d
    83.07%
    p50
    3.34s
    p99
    74.5s
    Throughput
    26 tps
    Context
    203K
    $/Mtok
    $0.980 / $3.08
    chutes/fp8fp8
  • No data24h 99.90%
    7d
    99.78%
    30d
    99.78%
    p50
    1.78s
    p99
    3.91s (best)
    Throughput
    56 tps
    Context
    203K
    $/Mtok
    $1.33 / $4.18
    alibaba/fp8fp8
  • No data24h 99.70%
    7d
    99.41%
    30d
    99.41%
    p50
    5.00s
    p99
    7.38s
    Throughput
    8.0 tps
    Context
    203K
    $/Mtok
    $1.05 / $3.50
    deepinfra/fp4fp4
  • No data24h 97.37%
    7d
    94.17%
    30d
    94.17%
    p50
    1.82s
    p99
    17.2s
    Throughput
    11 tps
    Context
    164K
    $/Mtok
    $1.30 / $4.30
    digitalocean
  • No data24h 99.87%
    7d
    99.36%
    30d
    99.36%
    p50
    1.76s
    p99
    6.34s
    Throughput
    59 tps
    Context
    205K
    $/Mtok
    $1.38 / $4.40
    novita/fp8fp8
  • No data24h 76.83%
    7d
    83.99%
    30d
    83.99%
    p50
    2.35s
    p99
    27.0s
    Throughput
    28 tps
    Context
    203K
    $/Mtok
    $1.21 / $4.20
    phala
  • No data24h 97.86%
    7d
    96.55%
    30d
    96.55%
    p50
    6.07s
    p99
    8.21s
    Throughput
    10 tps
    Context
    200K
    $/Mtok
    $1.54 / $4.84
    venice/fp8fp8

Measured 4m ago. Latency and throughput are OpenRouter's rolling 30-minute windows. ◆ marks the best value in each column among these endpoints.

Price vs latency

One dot per endpoint at the latest round · further left = faster, lower = cheaper · hover a dot for its tag

History

Availability

Axis starts at 90% so small differences stay visible; it expands to 0 when an endpoint falls below.

Time to first token

Logarithmic axis — the fleet spans two orders of magnitude. P50 is OpenRouter's rolling 30-minute window, averaged over the hour.

Throughput

Median tokens per second.