modelstatus.dev

← GLM 5.3 Flash

GLM 5.3 Flash: providers compared

Every endpoint serving this model, side by side. ◆ marks the best value in each column — per model, never fleet-wide.

  • AtlasCloud100%◆ (best)
    Healthy24h 62.63%
    7d
    87.93%
    30d
    94.94%
    p50
    3.28s
    p99
    18.5s
    Throughput
    50 tps
    Context
    1.0M
    $/Mtok
    $0.150 / $0.500
    atlas-cloud/fp8fp8
  • BaseTen100%◆ (best)
    Healthy24h 100%
    7d
    99.64%
    30d
    99.47%
    p50
    351ms
    p99
    6.28s
    Throughput
    129 tps
    Context
    1.0M
    $/Mtok
    $0.150 / $0.500
    baseten/fp8fp8
  • CoreWeave100%◆ (best)
    Healthy24h 100%
    7d
    97.99%
    30d
    98.53%
    p50
    647ms
    p99
    6.43s
    Throughput
    262 tps◆ (best)
    Context
    1.0M
    $/Mtok
    $0.150 / $0.500
    coreweave/nvfp4nvfp4
  • Crusoe100%◆ (best)
    Healthy24h 100%
    7d
    99.69%
    30d
    96.80%
    p50
    885ms
    p99
    3.00s
    Throughput
    159 tps
    Context
    1.0M
    $/Mtok
    $0.150 / $0.500
    crusoe/fp4fp4
  • Decart100%◆ (best)
    Healthy24h 100%
    7d
    99.81%
    30d
    98.86%
    p50
    664ms
    p99
    5.54s
    Throughput
    51 tps
    Context
    1.0M
    $/Mtok
    $0.093 / $0.310
    decart/fp4fp4
  • DekaLLM100%◆ (best)
    Healthy24h 99.58%
    7d
    98.37%
    30d
    98.18%
    p50
    792ms
    p99
    3.93s
    Throughput
    124 tps
    Context
    1.0M
    $/Mtok
    $0.100 / $1.00
    dekallm
  • DigitalOcean100%◆ (best)
    Healthy24h 99.93%
    7d
    99.21%
    30d
    97.76%
    p50
    702ms
    p99
    11.3s
    Throughput
    47 tps
    Context
    1.0M
    $/Mtok
    $0.150 / $0.500
    digitalocean
  • Fireworks100%◆ (best)
    Healthy24h 100%
    7d
    98.27%
    30d
    98.98%
    p50
    729ms
    p99
    4.58s
    Throughput
    117 tps
    Context
    1.0M
    $/Mtok
    $0.150 / $0.500
    fireworks
  • Fireworks100%◆ (best)
    Healthy24h 99.83%
    7d
    99.40%
    30d
    99.47%
    p50
    586ms
    p99
    2.22s
    Throughput
    101 tps
    Context
    1.0M
    $/Mtok
    $0.225 / $0.750
    fireworks/us
  • Friendli100%◆ (best)
    Healthy24h 100%
    7d
    98.14%
    30d
    98.00%
    p50
    698ms
    p99
    10.0s
    Throughput
    159 tps
    Context
    1.0M
    $/Mtok
    $0.150 / $0.500
    friendli
  • GMICloud100%◆ (best)
    Healthy24h 99.79%
    7d
    97.24%
    30d
    98.60%
    p50
    2.88s
    p99
    13.2s
    Throughput
    43 tps
    Context
    1.0M
    $/Mtok
    $0.090 / $0.300
    gmicloud/fp8fp8
  • Inceptron100%◆ (best)
    Healthy24h 99.79%
    7d
    95.06%
    30d
    96.54%
    p50
    412ms
    p99
    6.08s
    Throughput
    67 tps
    Context
    1.0M
    $/Mtok
    $0.120 / $0.550
    inceptron/fp8fp8
  • InferenceNet100%◆ (best)
    Healthy24h 100%
    7d
    99.81%
    30d
    98.57%
    p50
    462ms
    p99
    4.33s
    Throughput
    85 tps
    Context
    1.0M
    $/Mtok
    $0.070 / $0.050
    inference-net/fp4fp4
  • Modal100%◆ (best)
    Healthy24h 100%
    7d
    99.82%
    30d
    99.68%
    p50
    257ms◆ (best)
    p99
    1.51s◆ (best)
    Throughput
    148 tps
    Context
    1.0M
    $/Mtok
    $0.150 / $0.500
    modal/nvfp4nvfp4
  • Morph100%◆ (best)
    Healthy24h 85.54%
    7d
    94.57%
    30d
    90.67%
    p50
    1.42s
    p99
    20.8s
    Throughput
    27 tps
    Context
    1.0M
    $/Mtok
    $0.060 / $0.750
    morph/fp8fp8
  • Near AI100%◆ (best)
    Healthy24h 100%
    7d
    99.44%
    30d
    98.41%
    p50
    1.45s
    p99
    14.9s
    Throughput
    31 tps
    Context
    1.0M
    $/Mtok
    $0.105 / $0.350
    near-ai/fp8fp8
  • Novita100%◆ (best)
    Healthy24h 99.86%
    7d
    96.38%
    30d
    98.04%
    p50
    3.20s
    p99
    27.9s
    Throughput
    39 tps
    Context
    1.0M
    $/Mtok
    $0.084 / $0.280
    novita/fp8fp8
  • OpenInference100%◆ (best)
    Healthy24h 100%
    7d
    99.64%
    30d
    96.18%
    p50
    1.77s
    p99
    13.6s
    Throughput
    28 tps
    Context
    1.0M
    $/Mtok
    $0.044 / $0.450
    open-inference/fp4fp4
  • Parasail100%◆ (best)
    Healthy24h 99.86%
    7d
    99.91%
    30d
    99.91%◆ (best)
    p50
    941ms
    p99
    4.76s
    Throughput
    181 tps
    Context
    1.0M
    $/Mtok
    $0.188 / $0.625
    parasail/fastfp4
  • Parasail100%◆ (best)
    Healthy24h 100%
    7d
    99.78%
    30d
    99.69%
    p50
    961ms
    p99
    9.18s
    Throughput
    87 tps
    Context
    1.0M
    $/Mtok
    $0.150 / $0.500
    parasail/fp4fp4
  • Phala100%◆ (best)
    Healthy24h 100%
    7d
    99.74%
    30d
    99.74%
    p50
    4.17s
    p99
    11.2s
    Throughput
    26 tps
    Context
    1.0M
    $/Mtok
    $0.112 / $0.375
    phala/nvfp4nvfp4
  • Reka100%◆ (best)
    Healthy24h 100%
    7d
    99.14%
    30d
    99.44%
    p50
    1.16s
    p99
    9.52s
    Throughput
    177 tps
    Context
    262K
    $/Mtok
    $0.060 / $1.60
    reka
  • Relace100%◆ (best)
    Healthy24h 100%
    7d
    99.83%
    30d
    99.51%
    p50
    1.53s
    p99
    15.5s
    Throughput
    47 tps
    Context
    1.0M
    $/Mtok
    $0.040 / $0.500
    relace
  • Sail Research100%◆ (best)
    Healthy24h 99.30%
    7d
    99.26%
    30d
    99.26%
    p50
    1.43s
    p99
    17.5s
    Throughput
    77 tps
    Context
    1.0M
    $/Mtok
    $0.045 / $0.600
    sail-research/fp4fp4
  • Sail Research100%◆ (best)
    Healthy24h 99.44%
    7d
    99.45%
    30d
    99.52%
    p50
    1.42s
    p99
    24.5s
    Throughput
    59 tps
    Context
    1.0M
    $/Mtok
    $0.045 / $0.600
    sail-research/usfp4
  • SiliconFlow100%◆ (best)
    Healthy24h 68.83%
    7d
    94.61%
    30d
    97.26%
    p50
    2.43s
    p99
    21.6s
    Throughput
    41 tps
    Context
    1.0M
    $/Mtok
    $0.150 / $0.500
    siliconflow/fp8fp8
  • StreamLake100%◆ (best)
    Healthy24h 100%
    7d
    99.07%
    30d
    98.91%
    p50
    1.13s
    p99
    10.6s
    Throughput
    42 tps
    Context
    1.0M
    $/Mtok
    $0.069 / $0.230
    streamlake/fp8fp8
  • Together100%◆ (best)
    Healthy24h 99.93%
    7d
    99.88%
    30d
    99.08%
    p50
    719ms
    p99
    16.0s
    Throughput
    131 tps
    Context
    1.0M
    $/Mtok
    $0.150 / $0.500
    together
  • Venice100%◆ (best)
    Healthy24h 100%
    7d
    98.90%
    30d
    97.79%
    p50
    2.17s
    p99
    9.10s
    Throughput
    90 tps
    Context
    1.0M
    $/Mtok
    $0.150 / $0.500
    venice
  • Wafer100%◆ (best)
    Healthy24h 100%
    7d
    99.93%
    30d
    99.58%
    p50
    644ms
    p99
    8.01s
    Throughput
    55 tps
    Context
    1.0M
    $/Mtok
    $0.090 / $0.500
    wafer
  • Z.AI100%◆ (best)
    Healthy24h 100%
    7d
    99.53%
    30d
    99.41%
    p50
    3.17s
    p99
    20.6s
    Throughput
    43 tps
    Context
    1.0M
    $/Mtok
    $0.150 / $0.500
    z-ai/fp8fp8
  • No data24h —
    7d
    99.19%
    30d
    99.08%
    p50
    —
    p99
    —
    Throughput
    —
    Context
    1.0M
    $/Mtok
    $0.300 / $1.00
    cloudflare
  • No data24h —
    7d
    —
    30d
    98.99%
    p50
    —
    p99
    —
    Throughput
    —
    Context
    1.0M
    $/Mtok
    $0.150 / $0.500
    coreweave/fp8fp8
  • No data24h 100%
    7d
    98.74%
    30d
    98.88%
    p50
    1.85s
    p99
    7.73s
    Throughput
    14 tps
    Context
    1.0M
    $/Mtok
    $0.075 / $0.250
    deepinfra/fp4fp4
  • No data24h —
    7d
    —
    30d
    —
    p50
    —
    p99
    —
    Throughput
    —
    Context
    1.0M
    $/Mtok
    $0.075 / $0.250
    deepinfra/fp8fp8
  • No data24h —
    7d
    98.07%
    30d
    97.04%
    p50
    —
    p99
    —
    Throughput
    —
    Context
    262K
    $/Mtok
    $0.112 / $0.375
    io-net/fp8fp8
  • No data24h —
    7d
    —
    30d
    94.66%
    p50
    —
    p99
    —
    Throughput
    —
    Context
    1.0M
    $/Mtok
    $0.140 / $0.470
    makora
  • No data24h —
    7d
    —
    30d
    99.53%
    p50
    —
    p99
    —
    Throughput
    —
    Context
    1.0M
    $/Mtok
    $0.450 / $1.50
    modal/fp8fp8
  • No data24h —
    7d
    —
    30d
    95.71%
    p50
    —
    p99
    —
    Throughput
    —
    Context
    1.0M
    $/Mtok
    $0.200 / $0.700
    morph
  • No data24h —
    7d
    —
    30d
    —
    p50
    —
    p99
    —
    Throughput
    —
    Context
    1.0M
    $/Mtok
    $0.105 / $0.350
    near-ai
  • No data24h —
    7d
    99.94%◆ (best)
    30d
    99.44%
    p50
    —
    p99
    —
    Throughput
    —
    Context
    1.0M
    $/Mtok
    $0.165 / $0.550
    nextbit/fp8fp8
  • No data24h —
    7d
    —
    30d
    98.84%
    p50
    —
    p99
    —
    Throughput
    —
    Context
    1.0M
    $/Mtok
    $0.150 / $0.500
    parasail/fp8fp8
  • No data24h —
    7d
    98.42%
    30d
    98.26%
    p50
    —
    p99
    —
    Throughput
    —
    Context
    1.0M
    $/Mtok
    $0.112 / $0.375
    phala/fp8fp8
  • No data24h —
    7d
    98.79%
    30d
    98.86%
    p50
    —
    p99
    —
    Throughput
    —
    Context
    1.0M
    $/Mtok
    $0.150 / $0.500
    primeintellect
  • Reka—
    No data24h —
    7d
    —
    30d
    98.89%
    p50
    —
    p99
    —
    Throughput
    —
    Context
    262K
    $/Mtok
    $0.150 / $0.500
    reka/fp8fp8
  • No data24h —
    7d
    —
    30d
    99.71%
    p50
    —
    p99
    —
    Throughput
    —
    Context
    1.0M
    $/Mtok
    $0.035 / $0.500◆ (best)
    relace/fp4fp4
  • No data24h —
    7d
    —
    30d
    —
    p50
    —
    p99
    —
    Throughput
    —
    Context
    1.0M
    $/Mtok
    $0.075 / $0.250
    relace/fp8fp8
  • No data24h —
    7d
    —
    30d
    99.15%
    p50
    —
    p99
    —
    Throughput
    —
    Context
    1.0M
    $/Mtok
    $0.045 / $0.600
    sail-research/fp8fp8
  • No data24h —
    7d
    —
    30d
    —
    p50
    —
    p99
    —
    Throughput
    —
    Context
    1.0M
    $/Mtok
    $0.150 / $0.500
    streamlake

Measured 1m ago. Latency and throughput are OpenRouter's rolling 30-minute windows. ◆ marks the best value in each column among these endpoints.

Price vs latency

One dot per endpoint at the latest round · further left = faster, lower = cheaper · hover a dot for its tag

History

Availability

Axis starts at 90% so small differences stay visible; it expands to 0 when an endpoint falls below.

Time to first token

Logarithmic axis — the fleet spans two orders of magnitude. P50 is OpenRouter's rolling 30-minute window.

Throughput

Median tokens per second.