modelstatus.dev

← gpt-oss-120b

gpt-oss-120b: providers compared

Every endpoint serving this model, side by side. ◆ marks the best value in each column — per model, never fleet-wide.

  • Amazon Bedrock100%◆ (best)
    Healthy24h 99.40%
    7d
    99.91%
    30d
    98.18%
    p50
    603ms
    p99
    17.0s
    Throughput
    143 tps
    Context
    131K
    $/Mtok
    $0.150 / $0.600
    amazon-bedrock
  • BaseTen100%◆ (best)
    Healthy24h 99.97%
    7d
    99.85%
    30d
    99.93%
    p50
    265ms
    p99
    5.76s
    Throughput
    124 tps
    Context
    128K
    $/Mtok
    $0.100 / $0.500
    baseten/fp4fp4
  • Cerebras100%◆ (best)
    Healthy24h 99.98%
    7d
    99.95%
    30d
    99.95%
    p50
    243ms
    p99
    3.68s
    Throughput
    745 tps◆ (best)
    Context
    131K
    $/Mtok
    $0.350 / $0.750
    cerebras/fp16fp16
  • Crusoe100%◆ (best)
    Healthy24h 99.99%
    7d
    99.96%
    30d
    99.64%
    p50
    208ms◆ (best)
    p99
    5.45s
    Throughput
    131 tps
    Context
    131K
    $/Mtok
    $0.050 / $0.250
    crusoe/bf16bf16
  • DeepInfra100%◆ (best)
    Healthy24h 99.98%
    7d
    99.96%
    30d
    99.97%◆ (best)
    p50
    561ms
    p99
    2.05s
    Throughput
    98 tps
    Context
    131K
    $/Mtok
    $0.150 / $0.600
    deepinfra/turbobf16
  • DigitalOcean100%◆ (best)
    Healthy24h 99.98%
    7d
    99.94%
    30d
    99.94%
    p50
    483ms
    p99
    1.94s◆ (best)
    Throughput
    25 tps
    Context
    128K
    $/Mtok
    $0.060 / $0.420
    digitalocean
  • Parasail100%◆ (best)
    Healthy24h 99.97%
    7d
    99.89%
    30d
    99.35%
    p50
    406ms
    p99
    1.98s
    Throughput
    86 tps
    Context
    131K
    $/Mtok
    $0.100 / $0.750
    parasail/fp4fp4
  • AkashML99.99%
    Healthy24h 99.95%
    7d
    99.97%◆ (best)
    30d
    99.91%
    p50
    947ms
    p99
    33.0s
    Throughput
    32 tps
    Context
    131K
    $/Mtok
    $0.037 / $0.187
    akashml/bf16bf16
  • DekaLLM99.93%
    Healthy24h 99.53%
    7d
    99.47%
    30d
    99.43%
    p50
    901ms
    p99
    49.8s
    Throughput
    19 tps
    Context
    131K
    $/Mtok
    $0.030 / $0.180◆ (best)
    dekallm/bf16bf16
  • Groq99.90%
    Healthy24h 99.40%
    7d
    99.08%
    30d
    99.64%
    p50
    389ms
    p99
    2.49s
    Throughput
    274 tps
    Context
    131K
    $/Mtok
    $0.150 / $0.600
    groq
  • CoreWeave99.80%
    Healthy24h 98.92%
    7d
    98.99%
    30d
    99.68%
    p50
    475ms
    p99
    25.5s
    Throughput
    32 tps
    Context
    131K
    $/Mtok
    $0.030 / $0.170◆ (best)
    coreweave/fp4fp4
  • SambaNova99.71%
    Healthy24h 99.76%
    7d
    99.39%
    30d
    96.63%
    p50
    800ms
    p99
    7.58s
    Throughput
    214 tps
    Context
    131K
    $/Mtok
    $0.140 / $0.950
    sambanova
  • Together98.10%
    Healthy24h 88.62%
    7d
    92.68%
    30d
    92.63%
    p50
    261ms
    p99
    4.95s
    Throughput
    79 tps
    Context
    131K
    $/Mtok
    $0.150 / $0.600
    together
  • Mancer 293.27%
    Degraded24h 97.50%
    7d
    99.00%
    30d
    99.07%
    p50
    1.50s
    p99
    47.5s
    Throughput
    25 tps
    Context
    131K
    $/Mtok
    $0.045 / $0.250
    mancer/fp8fp8
  • Phala91.82%
    Degraded24h 98.96%
    7d
    99.31%
    30d
    92.31%
    p50
    962ms
    p99
    9.81s
    Throughput
    93 tps
    Context
    131K
    $/Mtok
    $0.150 / $0.600
    phala
  • Novita90.75%
    Degraded24h 98.46%
    7d
    99.10%
    30d
    92.16%
    p50
    845ms
    p99
    19.4s
    Throughput
    104 tps
    Context
    131K
    $/Mtok
    $0.050 / $0.250
    novita/fp4fp4
  • Nebius88.95%
    Down24h 97.73%
    7d
    96.46%
    30d
    97.77%
    p50
    441ms
    p99
    11.5s
    Throughput
    132 tps
    Context
    131K
    $/Mtok
    $0.150 / $0.600
    nebius/fp4fp4
  • Google78.17%
    Down24h 66.73%
    7d
    54.88%
    30d
    78.16%
    p50
    534ms
    p99
    7.77s
    Throughput
    101 tps
    Context
    131K
    $/Mtok
    $0.090 / $0.360
    google-vertex/global
  • Mara74.60%
    Down24h 95.80%
    7d
    93.43%
    30d
    85.85%
    p50
    1.02s
    p99
    10.9s
    Throughput
    118 tps
    Context
    131K
    $/Mtok
    $0.150 / $0.750
    mara
  • DeepInfra64.61%
    Down24h 97.18%
    7d
    98.76%
    30d
    99.05%
    p50
    13.1s
    p99
    41.6s
    Throughput
    13 tps
    Context
    131K
    $/Mtok
    $0.037 / $0.170
    deepinfra/bf16bf16
  • No data24h 99.97%
    7d
    98.34%
    30d
    99.10%
    p50
    469ms
    p99
    10.5s
    Throughput
    127 tps
    Context
    131K
    $/Mtok
    $0.150 / $0.600
    amazon-bedrock/eu-west-1
  • No data24h —
    7d
    70.95%
    30d
    85.99%
    p50
    —
    p99
    —
    Throughput
    —
    Context
    131K
    $/Mtok
    $0.200 / $0.950
    deepinfra/fp8fp8
  • No data24h 91.27%
    7d
    86.45%
    30d
    88.19%
    p50
    —
    p99
    —
    Throughput
    —
    Context
    131K
    $/Mtok
    $0.150 / $0.600
    siliconflow/fp8fp8

Measured 1m ago. Latency and throughput are OpenRouter's rolling 30-minute windows. ◆ marks the best value in each column among these endpoints.

Price vs latency

One dot per endpoint at the latest round · further left = faster, lower = cheaper · hover a dot for its tag

History

Availability

Axis starts at 90% so small differences stay visible; it expands to 0 when an endpoint falls below.

Time to first token

Logarithmic axis — the fleet spans two orders of magnitude. P50 is OpenRouter's rolling 30-minute window.

Throughput

Median tokens per second.