modelstatus.dev

← gpt-oss-120b

gpt-oss-120b: providers compared

Every endpoint serving this model, side by side. ◆ marks the best value in each column — per model, never fleet-wide.

  • AkashML100%◆ (best)
    Healthy24h 99.95%
    7d
    99.97%◆ (best)
    30d
    99.91%
    p50
    3.15s
    p99
    26.7s
    Throughput
    60 tps
    Context
    131K
    $/Mtok
    $0.037 / $0.187
    akashml/bf16bf16
  • Amazon Bedrock100%◆ (best)
    Healthy24h 99.40%
    7d
    99.91%
    30d
    98.18%
    p50
    566ms
    p99
    17.5s
    Throughput
    171 tps
    Context
    131K
    $/Mtok
    $0.150 / $0.600
    amazon-bedrock
  • BaseTen100%◆ (best)
    Healthy24h 99.97%
    7d
    99.85%
    30d
    99.93%
    p50
    266ms
    p99
    796ms◆ (best)
    Throughput
    118 tps
    Context
    128K
    $/Mtok
    $0.100 / $0.500
    baseten/fp4fp4
  • Cerebras100%◆ (best)
    Healthy24h 99.98%
    7d
    99.95%
    30d
    99.95%
    p50
    220ms◆ (best)
    p99
    2.93s
    Throughput
    710 tps◆ (best)
    Context
    131K
    $/Mtok
    $0.350 / $0.750
    cerebras/fp16fp16
  • Crusoe100%◆ (best)
    Healthy24h 99.99%
    7d
    99.96%
    30d
    99.64%
    p50
    282ms
    p99
    1.49s
    Throughput
    116 tps
    Context
    131K
    $/Mtok
    $0.050 / $0.250
    crusoe/bf16bf16
  • DeepInfra100%◆ (best)
    Healthy24h 99.98%
    7d
    99.96%
    30d
    99.97%◆ (best)
    p50
    478ms
    p99
    1.79s
    Throughput
    127 tps
    Context
    131K
    $/Mtok
    $0.150 / $0.600
    deepinfra/turbobf16
  • DigitalOcean100%◆ (best)
    Healthy24h 99.98%
    7d
    99.94%
    30d
    99.94%
    p50
    359ms
    p99
    1.75s
    Throughput
    39 tps
    Context
    128K
    $/Mtok
    $0.060 / $0.420
    digitalocean
  • Mancer 2100%◆ (best)
    Healthy24h 97.43%
    7d
    98.96%
    30d
    99.06%
    p50
    675ms
    p99
    7.32s
    Throughput
    34 tps
    Context
    131K
    $/Mtok
    $0.045 / $0.250
    mancer/fp8fp8
  • Nebius100%◆ (best)
    Healthy24h 97.66%
    7d
    96.46%
    30d
    97.77%
    p50
    306ms
    p99
    5.93s
    Throughput
    182 tps
    Context
    131K
    $/Mtok
    $0.150 / $0.600
    nebius/fp4fp4
  • Novita100%◆ (best)
    Healthy24h 98.44%
    7d
    99.07%
    30d
    92.15%
    p50
    924ms
    p99
    7.70s
    Throughput
    128 tps
    Context
    131K
    $/Mtok
    $0.050 / $0.250
    novita/fp4fp4
  • Parasail100%◆ (best)
    Healthy24h 99.97%
    7d
    99.89%
    30d
    99.35%
    p50
    379ms
    p99
    1.34s
    Throughput
    109 tps
    Context
    131K
    $/Mtok
    $0.100 / $0.750
    parasail/fp4fp4
  • Phala100%◆ (best)
    Healthy24h 98.91%
    7d
    99.26%
    30d
    92.30%
    p50
    905ms
    p99
    4.99s
    Throughput
    98 tps
    Context
    131K
    $/Mtok
    $0.150 / $0.600
    phala
  • SambaNova100%◆ (best)
    Healthy24h 99.76%
    7d
    99.41%
    30d
    96.63%
    p50
    477ms
    p99
    4.86s
    Throughput
    198 tps
    Context
    131K
    $/Mtok
    $0.140 / $0.950
    sambanova
  • DeepInfra99.99%
    Healthy24h 97.03%
    7d
    98.61%
    30d
    99.02%
    p50
    364ms
    p99
    10.9s
    Throughput
    42 tps
    Context
    131K
    $/Mtok
    $0.037 / $0.170
    deepinfra/bf16bf16
  • Groq99.73%
    Healthy24h 99.39%
    7d
    99.09%
    30d
    99.64%
    p50
    319ms
    p99
    2.45s
    Throughput
    269 tps
    Context
    131K
    $/Mtok
    $0.150 / $0.600
    groq
  • DekaLLM99.72%
    Healthy24h 99.52%
    7d
    99.47%
    30d
    99.43%
    p50
    690ms
    p99
    38.2s
    Throughput
    24 tps
    Context
    131K
    $/Mtok
    $0.030 / $0.180◆ (best)
    dekallm/bf16bf16
  • CoreWeave99.68%
    Healthy24h 98.94%
    7d
    99.00%
    30d
    99.68%
    p50
    416ms
    p99
    24.1s
    Throughput
    35 tps
    Context
    131K
    $/Mtok
    $0.030 / $0.170◆ (best)
    coreweave/fp4fp4
  • Mara95.24%
    Degraded24h 95.71%
    7d
    93.50%
    30d
    85.83%
    p50
    810ms
    p99
    13.7s
    Throughput
    121 tps
    Context
    131K
    $/Mtok
    $0.150 / $0.750
    mara
  • Together78.59%
    Down24h 88.71%
    7d
    92.66%
    30d
    92.64%
    p50
    248ms
    p99
    1.94s
    Throughput
    78 tps
    Context
    131K
    $/Mtok
    $0.150 / $0.600
    together
  • Google78.19%
    Down24h 67.30%
    7d
    54.90%
    30d
    78.14%
    p50
    447ms
    p99
    6.85s
    Throughput
    152 tps
    Context
    131K
    $/Mtok
    $0.090 / $0.360
    google-vertex/global
  • No data24h 99.97%
    7d
    98.34%
    30d
    99.10%
    p50
    429ms
    p99
    10.2s
    Throughput
    120 tps
    Context
    131K
    $/Mtok
    $0.150 / $0.600
    amazon-bedrock/eu-west-1
  • No data24h —
    7d
    69.26%
    30d
    85.97%
    p50
    —
    p99
    —
    Throughput
    —
    Context
    131K
    $/Mtok
    $0.200 / $0.950
    deepinfra/fp8fp8
  • Down24h 90.61%
    7d
    86.28%
    30d
    88.15%
    p50
    —
    p99
    —
    Throughput
    —
    Context
    131K
    $/Mtok
    $0.150 / $0.600
    siliconflow/fp8fp8

Measured 1m ago. Latency and throughput are OpenRouter's rolling 30-minute windows. ◆ marks the best value in each column among these endpoints.

Price vs latency

One dot per endpoint at the latest round · further left = faster, lower = cheaper · hover a dot for its tag

History

Availability

Axis starts at 90% so small differences stay visible; it expands to 0 when an endpoint falls below.

Time to first token

Logarithmic axis — the fleet spans two orders of magnitude. P50 is OpenRouter's rolling 30-minute window.

Throughput

Median tokens per second.