modelstatus.dev

GLM 5.3

GLM 5.3: providers compared

Every endpoint serving this model, side by side. ◆ marks the best value in each column — per model, never fleet-wide.

  • Alibaba100% (best)
    Healthy24h 99.97%
    7d
    98.33%
    30d
    98.33%
    p50
    1.05s
    p99
    9.70s
    Throughput
    73 tps
    Context
    1.0M
    $/Mtok
    $1.40 / $4.40
    alibaba
  • AtlasCloud100% (best)
    Healthy24h 99.90%
    7d
    99.91%
    30d
    99.62%
    p50
    4.89s
    p99
    10.4s
    Throughput
    35 tps
    Context
    1.0M
    $/Mtok
    $1.40 / $4.40
    atlas-cloud/fp8fp8
  • BaseTen100% (best)
    Healthy24h 99.74%
    7d
    97.92%
    30d
    98.53%
    p50
    1.33s
    p99
    17.6s
    Throughput
    74 tps
    Context
    1.0M
    $/Mtok
    $1.40 / $4.40
    baseten/fp4fp4
  • BaseTen100% (best)
    Healthy24h 95.84%
    7d
    99.20%
    30d
    99.59%
    p50
    1.67s
    p99
    48.0s
    Throughput
    61 tps
    Context
    1.0M
    $/Mtok
    $2.10 / $6.60
    baseten/fp8fp8
  • Crusoe100% (best)
    Healthy24h 97.32%
    7d
    94.05%
    30d
    95.67%
    p50
    1.02s
    p99
    54.9s
    Throughput
    102 tps
    Context
    1.0M
    $/Mtok
    $1.40 / $4.40
    crusoe/fp4fp4
  • Decart100% (best)
    Degraded24h 98.32%
    7d
    99.32%
    30d
    97.96%
    p50
    1.27s
    p99
    14.0s
    Throughput
    115 tps
    Context
    1.0M
    $/Mtok
    $1.19 / $3.74
    decart/fp4fp4
  • DigitalOcean100% (best)
    Healthy24h 98.95%
    7d
    95.80%
    30d
    97.52%
    p50
    983ms
    p99
    5.90s
    Throughput
    46 tps
    Context
    1.0M
    $/Mtok
    $1.08 / $3.39
    digitalocean
  • Fireworks100% (best)
    Healthy24h 99.59%
    7d
    99.73%
    30d
    99.56%
    p50
    1.40s
    p99
    38.4s
    Throughput
    46 tps
    Context
    1.0M
    $/Mtok
    $1.40 / $4.40
    fireworks
  • Friendli100% (best)
    Healthy24h 99.93%
    7d
    99.98%
    30d
    99.92%
    p50
    665ms
    p99
    6.97s
    Throughput
    111 tps
    Context
    1.0M
    $/Mtok
    $1.26 / $3.96
    friendli
  • Inceptron100% (best)
    Healthy24h 98.51%
    7d
    98.51%
    30d
    98.16%
    p50
    706ms
    p99
    9.23s
    Throughput
    69 tps
    Context
    1.0M
    $/Mtok
    $1.07 / $4.09
    inceptron/fp4fp4
  • Io Net100% (best)
    Healthy24h 99.53%
    7d
    99.63%
    30d
    98.98%
    p50
    1.83s
    p99
    7.55s
    Throughput
    65 tps
    Context
    262K
    $/Mtok
    $0.910 / $3.08
    io-net/fp8fp8
  • Makora100% (best)
    Healthy24h 95.92%
    7d
    96.82%
    30d
    94.22%
    p50
    1.37s
    p99
    42.6s
    Throughput
    63 tps
    Context
    980K
    $/Mtok
    $1.05 / $4.20
    makora/fp4fp4
  • Morph100% (best)
    Healthy24h 99.38%
    7d
    98.63%
    30d
    98.63%
    p50
    3.10s
    p99
    37.4s
    Throughput
    20 tps
    Context
    1.0M
    $/Mtok
    $0.952 / $2.99
    morph/fp8fp8
  • Parasail100% (best)
    Healthy24h 99.21%
    7d
    99.03%
    30d
    97.09%
    p50
    1.08s
    p99
    94.3s
    Throughput
    84 tps
    Context
    1.0M
    $/Mtok
    $1.40 / $4.40
    parasail/fp8fp8
  • Phala100% (best)
    Healthy24h 99.71%
    7d
    99.02%
    30d
    98.95%
    p50
    1.29s
    p99
    15.4s
    Throughput
    79 tps
    Context
    1.0M
    $/Mtok
    $1.05 / $3.30
    phala
  • Sail Research100% (best)
    Healthy24h 99.92%
    7d
    99.93%
    30d
    99.94%
    p50
    1.41s
    p99
    7.64s
    Throughput
    54 tps
    Context
    1.0M
    $/Mtok
    $0.966 / $3.04
    sail-research/fp8fp8
  • Sail Research100% (best)
    Healthy24h 99.89%
    7d
    99.91%
    30d
    99.91%
    p50
    1.08s
    p99
    6.73s
    Throughput
    86 tps
    Context
    1.0M
    $/Mtok
    $0.966 / $3.04
    sail-research/usfp8
  • Z.AI99.86%
    Healthy24h 99.94%
    7d
    99.78%
    30d
    99.91%
    p50
    2.96s
    p99
    33.2s
    Throughput
    47 tps
    Context
    1.0M
    $/Mtok
    $1.40 / $4.40
    z-ai/fp8fp8
  • Reka99.51%
    Healthy24h 99.03%
    7d
    99.27%
    30d
    97.14%
    p50
    2.82s
    p99
    45.0s
    Throughput
    19 tps
    Context
    262K
    $/Mtok
    $0.877 / $2.97
    reka/fp8fp8
  • Novita99.44%
    Healthy24h 97.50%
    7d
    99.43%
    30d
    99.62%
    p50
    3.78s
    p99
    52.4s
    Throughput
    36 tps
    Context
    1.0M
    $/Mtok
    $0.910 / $2.86
    novita/fp8fp8
  • Modal99.40%
    Healthy24h 98.10%
    7d
    99.28%
    30d
    99.55%
    p50
    750ms
    p99
    33.3s
    Throughput
    73 tps
    Context
    1.0M
    $/Mtok
    $1.40 / $4.40
    modal
  • DeepInfra98.80%
    Healthy24h 95.47%
    7d
    97.04%
    30d
    96.24%
    p50
    3.51s
    p99
    111s
    Throughput
    20 tps
    Context
    1.0M
    $/Mtok
    $0.900 / $3.00
    deepinfra/fp4fp4
  • Wafer97.97%
    Healthy24h 99.46%
    7d
    98.31%
    30d
    97.44%
    p50
    557ms
    p99
    17.2s
    Throughput
    113 tps
    Context
    1.0M
    $/Mtok
    $1.40 / $4.40
    wafer
  • Together96.42%
    Healthy24h 94.95%
    7d
    96.79%
    30d
    96.76%
    p50
    349ms (best)
    p99
    52.9s
    Throughput
    120 tps (best)
    Context
    1.0M
    $/Mtok
    $1.40 / $4.40
    together
  • Degraded24h 84.34%
    7d
    38.70%
    30d
    38.70%
    p50
    2.85s
    p99
    94.7s
    Throughput
    26 tps
    Context
    1.0M
    $/Mtok
    $0.900 / $3.00
    inference-net/fp4fp4
  • No data24h 99.77%
    7d
    99.36%
    30d
    97.76%
    p50
    614ms
    p99
    2.23s
    Throughput
    105 tps
    Context
    1.0M
    $/Mtok
    $1.30 / $4.40
    akashml/fp8fp8
  • No data24h 100%
    7d
    100% (best)
    30d
    100% (best)
    p50
    624ms
    p99
    1.48s (best)
    Throughput
    84 tps
    Context
    1.0M
    $/Mtok
    $1.40 / $4.40
    baidu/fp8fp8
  • No data24h 99.95%
    7d
    99.78%
    30d
    99.72%
    p50
    4.18s
    p99
    69.0s
    Throughput
    38 tps
    Context
    1.3M
    $/Mtok
    $1.40 / $4.40
    cloudflare
  • No data24h
    7d
    30d
    99.41%
    p50
    p99
    Throughput
    Context
    1.0M
    $/Mtok
    $1.40 / $4.40
    crusoe/fp8fp8
  • No data24h
    7d
    30d
    88.65%
    p50
    p99
    Throughput
    Context
    1.0M
    $/Mtok
    $1.20 / $4.00
    deepinfra/bf16bf16
  • No data24h
    7d
    30d
    95.08%
    p50
    p99
    Throughput
    Context
    1.0M
    $/Mtok
    $1.20 / $4.00
    deepinfra/fp8fp8
  • No data24h
    7d
    30d
    p50
    p99
    Throughput
    Context
    1.0M
    $/Mtok
    $2.10 / $6.60
    fireworks/us
  • No data24h 98.54%
    7d
    99.26%
    30d
    98.74%
    p50
    3.71s
    p99
    8.62s
    Throughput
    21 tps
    Context
    1.0M
    $/Mtok
    $1.05 / $3.30
    gmicloud/fp8fp8
  • No data24h
    7d
    100% (best)
    30d
    99.72%
    p50
    p99
    Throughput
    Context
    262K
    $/Mtok
    $1.13 / $4.32
    io-net/fp4fp4
  • No data24h 98.85%
    7d
    99.34%
    30d
    99.34%
    p50
    4.75s
    p99
    145s
    Throughput
    34 tps
    Context
    1.0M
    $/Mtok
    $1.40 / $4.40
    mistral/nvfp4nvfp4
  • No data24h
    7d
    100% (best)
    30d
    95.37%
    p50
    p99
    Throughput
    Context
    1.0M
    $/Mtok
    $1.00 / $3.41
    morph/fp4fp4
  • No data24h 99.60%
    7d
    99.21%
    30d
    99.28%
    p50
    2.44s
    p99
    39.6s
    Throughput
    46 tps
    Context
    1.0M
    $/Mtok
    $1.12 / $3.52
    siliconflow/fp8fp8
  • No data24h
    7d
    30d
    100% (best)
    p50
    p99
    Throughput
    Context
    1.0M
    $/Mtok
    $1.11 / $3.50
    streamlake
  • No data24h
    7d
    99.65%
    30d
    98.85%
    p50
    p99
    Throughput
    Context
    1.0M
    $/Mtok
    $0.700 / $2.20 (best)
    streamlake/fp8fp8
  • No data24h 97.44%
    7d
    98.51%
    30d
    96.21%
    p50
    1.33s
    p99
    86.4s
    Throughput
    44 tps
    Context
    1.0M
    $/Mtok
    $1.40 / $4.40
    venice

Measured 1m ago. Latency and throughput are OpenRouter's rolling 30-minute windows. ◆ marks the best value in each column among these endpoints.

Price vs latency

One dot per endpoint at the latest round · further left = faster, lower = cheaper · hover a dot for its tag

History

Availability

Axis starts at 90% so small differences stay visible; it expands to 0 when an endpoint falls below.

Time to first token

Logarithmic axis — the fleet spans two orders of magnitude. P50 is OpenRouter's rolling 30-minute window.

Throughput

Median tokens per second.