modelstatus.dev

GLM 5.3

GLM 5.3: providers compared

Every endpoint serving this model, side by side. ◆ marks the best value in each column — per model, never fleet-wide.

  • AkashML100% (best)
    Healthy24h 99.77%
    7d
    99.36%
    30d
    97.76%
    p50
    773ms
    p99
    3.16s
    Throughput
    78 tps
    Context
    1.0M
    $/Mtok
    $1.30 / $4.40
    akashml/fp8fp8
  • Alibaba100% (best)
    Healthy24h 99.97%
    7d
    98.35%
    30d
    98.35%
    p50
    917ms
    p99
    5.08s
    Throughput
    79 tps
    Context
    1.0M
    $/Mtok
    $1.40 / $4.40
    alibaba
  • AtlasCloud100% (best)
    Healthy24h 99.90%
    7d
    99.92%
    30d
    99.62%
    p50
    3.32s
    p99
    8.41s
    Throughput
    59 tps
    Context
    1.0M
    $/Mtok
    $1.40 / $4.40
    atlas-cloud/fp8fp8
  • BaseTen100% (best)
    Healthy24h 99.75%
    7d
    97.93%
    30d
    98.53%
    p50
    1.46s
    p99
    11.3s
    Throughput
    62 tps
    Context
    1.0M
    $/Mtok
    $1.40 / $4.40
    baseten/fp4fp4
  • Crusoe100% (best)
    Healthy24h 97.38%
    7d
    94.09%
    30d
    95.69%
    p50
    633ms
    p99
    12.6s
    Throughput
    103 tps
    Context
    1.0M
    $/Mtok
    $1.40 / $4.40
    crusoe/fp4fp4
  • DigitalOcean100% (best)
    Healthy24h 98.96%
    7d
    95.83%
    30d
    97.53%
    p50
    940ms
    p99
    4.49s
    Throughput
    46 tps
    Context
    1.0M
    $/Mtok
    $1.08 / $3.39
    digitalocean
  • Fireworks100% (best)
    Healthy24h 99.60%
    7d
    99.73%
    30d
    99.56%
    p50
    1.13s
    p99
    39.9s
    Throughput
    46 tps
    Context
    1.0M
    $/Mtok
    $1.40 / $4.40
    fireworks
  • Friendli100% (best)
    Healthy24h 99.93%
    7d
    99.98%
    30d
    99.92%
    p50
    714ms
    p99
    7.69s
    Throughput
    106 tps
    Context
    1.0M
    $/Mtok
    $1.26 / $3.96
    friendli
  • Inceptron100% (best)
    Healthy24h 98.49%
    7d
    98.52%
    30d
    98.16%
    p50
    1.16s
    p99
    14.1s
    Throughput
    77 tps
    Context
    1.0M
    $/Mtok
    $1.07 / $4.09
    inceptron/fp4fp4
  • InferenceNet100% (best)
    Healthy24h 91.94%
    7d
    52.67%
    30d
    52.67%
    p50
    871ms
    p99
    48.2s
    Throughput
    58 tps
    Context
    1.0M
    $/Mtok
    $0.900 / $3.00
    inference-net/fp4fp4
  • Io Net100% (best)
    Healthy24h 99.53%
    7d
    99.64%
    30d
    98.98%
    p50
    1.82s
    p99
    19.3s
    Throughput
    55 tps
    Context
    262K
    $/Mtok
    $0.910 / $3.08
    io-net/fp8fp8
  • Morph100% (best)
    Healthy24h 99.40%
    7d
    98.63%
    30d
    98.63%
    p50
    1.54s
    p99
    11.7s
    Throughput
    67 tps
    Context
    1.0M
    $/Mtok
    $1.01 / $3.18
    morph/fp8fp8
  • Parasail100% (best)
    Healthy24h 99.19%
    7d
    99.03%
    30d
    97.09%
    p50
    910ms
    p99
    57.2s
    Throughput
    95 tps
    Context
    1.0M
    $/Mtok
    $1.40 / $4.40
    parasail/fp8fp8
  • Phala100% (best)
    Healthy24h 99.71%
    7d
    99.03%
    30d
    98.95%
    p50
    1.08s
    p99
    16.1s
    Throughput
    53 tps
    Context
    1.0M
    $/Mtok
    $1.05 / $3.30
    phala
  • Sail Research100% (best)
    Healthy24h 99.89%
    7d
    99.91%
    30d
    99.91%
    p50
    1.12s
    p99
    9.40s
    Throughput
    65 tps
    Context
    1.0M
    $/Mtok
    $1.00 / $3.16
    sail-research/usfp8
  • Z.AI100% (best)
    Healthy24h 99.94%
    7d
    99.78%
    30d
    99.91%
    p50
    3.10s
    p99
    30.6s
    Throughput
    48 tps
    Context
    1.0M
    $/Mtok
    $1.40 / $4.40
    z-ai/fp8fp8
  • Modal99.94%
    Healthy24h 98.13%
    7d
    99.28%
    30d
    99.55%
    p50
    865ms
    p99
    34.5s
    Throughput
    71 tps
    Context
    1.0M
    $/Mtok
    $1.40 / $4.40
    modal
  • Novita99.63%
    Healthy24h 97.52%
    7d
    99.43%
    30d
    99.62%
    p50
    3.82s
    p99
    60.6s
    Throughput
    32 tps
    Context
    1.0M
    $/Mtok
    $0.910 / $2.86
    novita/fp8fp8
  • Wafer99.60%
    Healthy24h 99.46%
    7d
    98.31%
    30d
    97.45%
    p50
    517ms
    p99
    14.5s
    Throughput
    107 tps
    Context
    1.0M
    $/Mtok
    $1.40 / $4.40
    wafer
  • Reka99.30%
    Healthy24h 98.95%
    7d
    99.26%
    30d
    97.14%
    p50
    2.21s
    p99
    41.4s
    Throughput
    19 tps
    Context
    262K
    $/Mtok
    $0.877 / $2.97
    reka/fp8fp8
  • DeepInfra99.13%
    Degraded24h 95.49%
    7d
    97.04%
    30d
    96.24%
    p50
    3.78s
    p99
    67.5s
    Throughput
    21 tps
    Context
    1.0M
    $/Mtok
    $0.900 / $3.00
    deepinfra/fp4fp4
  • Healthy24h 99.92%
    7d
    99.93%
    30d
    99.94%
    p50
    1.25s
    p99
    4.24s
    Throughput
    69 tps
    Context
    1.0M
    $/Mtok
    $1.00 / $3.16
    sail-research/fp8fp8
  • Makora98.17%
    Healthy24h 95.87%
    7d
    96.83%
    30d
    94.23%
    p50
    933ms
    p99
    54.1s
    Throughput
    83 tps
    Context
    980K
    $/Mtok
    $1.05 / $4.20
    makora/fp4fp4
  • Together84.94%
    Down24h 94.67%
    7d
    96.78%
    30d
    96.76%
    p50
    329ms (best)
    p99
    40.2s
    Throughput
    112 tps (best)
    Context
    1.0M
    $/Mtok
    $1.40 / $4.40
    together
  • Decart2.13%
    Down24h 98.19%
    7d
    99.30%
    30d
    97.95%
    p50
    1.22s
    p99
    21.2s
    Throughput
    95 tps
    Context
    1.0M
    $/Mtok
    $1.19 / $3.74
    decart/fp4fp4
  • No data24h 100%
    7d
    100% (best)
    30d
    100% (best)
    p50
    783ms
    p99
    1.30s (best)
    Throughput
    77 tps
    Context
    1.0M
    $/Mtok
    $1.40 / $4.40
    baidu/fp8fp8
  • No data24h 95.88%
    7d
    99.21%
    30d
    99.59%
    p50
    1.23s
    p99
    34.8s
    Throughput
    62 tps
    Context
    1.0M
    $/Mtok
    $2.10 / $6.60
    baseten/fp8fp8
  • No data24h 99.95%
    7d
    99.78%
    30d
    99.72%
    p50
    2.77s
    p99
    68.0s
    Throughput
    35 tps
    Context
    1.3M
    $/Mtok
    $1.40 / $4.40
    cloudflare
  • No data24h
    7d
    30d
    99.41%
    p50
    p99
    Throughput
    Context
    1.0M
    $/Mtok
    $1.40 / $4.40
    crusoe/fp8fp8
  • No data24h
    7d
    30d
    88.65%
    p50
    p99
    Throughput
    Context
    1.0M
    $/Mtok
    $1.20 / $4.00
    deepinfra/bf16bf16
  • No data24h
    7d
    30d
    95.08%
    p50
    p99
    Throughput
    Context
    1.0M
    $/Mtok
    $1.20 / $4.00
    deepinfra/fp8fp8
  • No data24h
    7d
    30d
    p50
    p99
    Throughput
    Context
    1.0M
    $/Mtok
    $2.10 / $6.60
    fireworks/us
  • No data24h 98.53%
    7d
    99.26%
    30d
    98.74%
    p50
    2.62s
    p99
    14.4s
    Throughput
    43 tps
    Context
    1.0M
    $/Mtok
    $1.05 / $3.30
    gmicloud/fp8fp8
  • No data24h
    7d
    100% (best)
    30d
    99.72%
    p50
    p99
    Throughput
    Context
    262K
    $/Mtok
    $1.13 / $4.32
    io-net/fp4fp4
  • No data24h 98.84%
    7d
    99.34%
    30d
    99.34%
    p50
    4.03s
    p99
    209s
    Throughput
    46 tps
    Context
    1.0M
    $/Mtok
    $1.40 / $4.40
    mistral/nvfp4nvfp4
  • No data24h
    7d
    100% (best)
    30d
    95.37%
    p50
    p99
    Throughput
    Context
    1.0M
    $/Mtok
    $1.00 / $3.41
    morph/fp4fp4
  • No data24h 99.60%
    7d
    99.22%
    30d
    99.29%
    p50
    2.00s
    p99
    14.7s
    Throughput
    47 tps
    Context
    1.0M
    $/Mtok
    $1.12 / $3.52
    siliconflow/fp8fp8
  • No data24h
    7d
    30d
    100% (best)
    p50
    p99
    Throughput
    Context
    1.0M
    $/Mtok
    $1.11 / $3.50
    streamlake
  • No data24h
    7d
    99.65%
    30d
    98.85%
    p50
    p99
    Throughput
    Context
    1.0M
    $/Mtok
    $0.700 / $2.20 (best)
    streamlake/fp8fp8
  • No data24h 97.42%
    7d
    98.51%
    30d
    96.21%
    p50
    695ms
    p99
    4.40s
    Throughput
    48 tps
    Context
    1.0M
    $/Mtok
    $1.40 / $4.40
    venice

Measured 3m ago. Latency and throughput are OpenRouter's rolling 30-minute windows. ◆ marks the best value in each column among these endpoints.

Price vs latency

One dot per endpoint at the latest round · further left = faster, lower = cheaper · hover a dot for its tag

History

Availability

Axis starts at 90% so small differences stay visible; it expands to 0 when an endpoint falls below.

Time to first token

Logarithmic axis — the fleet spans two orders of magnitude. P50 is OpenRouter's rolling 30-minute window.

Throughput

Median tokens per second.