modelstatus.dev

gpt-oss-20b

gpt-oss-20b: providers compared

Every endpoint serving this model, side by side. ◆ marks the best value in each column — per model, never fleet-wide.

  • Amazon Bedrock100% (best)
    Healthy24h 99.66%
    7d
    99.13%
    30d
    99.13%
    p50
    357ms
    p99
    1.51s
    Throughput
    328 tps
    Context
    131K
    $/Mtok
    $0.070 / $0.150
    amazon-bedrock
  • Google100% (best)
    Degraded24h 82.52%
    7d
    79.81%
    30d
    79.81%
    p50
    5.89s
    p99
    52.3s
    Throughput
    21 tps
    Context
    131K
    $/Mtok
    $0.070 / $0.250
    google-vertex/us-central1
  • Novita100% (best)
    Healthy24h 95.95%
    7d
    98.90%
    30d
    98.90%
    p50
    649ms
    p99
    19.2s
    Throughput
    104 tps
    Context
    131K
    $/Mtok
    $0.040 / $0.150
    novita/fp4fp4
  • DeepInfra99.90%
    Healthy24h 99.88%
    7d
    99.53% (best)
    30d
    99.53% (best)
    p50
    284ms
    p99
    11.0s
    Throughput
    86 tps
    Context
    131K
    $/Mtok
    $0.030 / $0.140 (best)
    deepinfra/bf16bf16
  • CoreWeave99.43%
    Healthy24h 99.29%
    7d
    99.23%
    30d
    99.23%
    p50
    283ms
    p99
    5.93s
    Throughput
    87 tps
    Context
    131K
    $/Mtok
    $0.030 / $0.130 (best)
    coreweave/fp4fp4
  • Groq98.76%
    Healthy24h 98.33%
    7d
    98.17%
    30d
    98.17%
    p50
    679ms
    p99
    4.57s
    Throughput
    334 tps (best)
    Context
    131K
    $/Mtok
    $0.075 / $0.300
    groq
  • Parasail98.52%
    Healthy24h 99.02%
    7d
    98.74%
    30d
    98.74%
    p50
    500ms
    p99
    6.67s
    Throughput
    65 tps
    Context
    131K
    $/Mtok
    $0.030 / $0.150 (best)
    parasail/fp4fp4
  • Fireworks95.59%
    Healthy24h 97.47%
    7d
    97.14%
    30d
    97.14%
    p50
    332ms
    p99
    9.21s
    Throughput
    142 tps
    Context
    131K
    $/Mtok
    $0.070 / $0.300
    fireworks
  • Phala91.86%
    Degraded24h 92.96%
    7d
    92.07%
    30d
    92.07%
    p50
    772ms
    p99
    9.55s
    Throughput
    52 tps
    Context
    131K
    $/Mtok
    $0.040 / $0.150
    phala
  • Down24h 97.26%
    7d
    97.98%
    30d
    97.98%
    p50
    1.36s
    p99
    34.1s
    Throughput
    34 tps
    Context
    131K
    $/Mtok
    $0.040 / $0.180
    siliconflow/fp8fp8
  • No data24h
    7d
    30d
    p50
    p99
    Throughput
    Context
    131K
    $/Mtok
    $0.070 / $0.150
    amazon-bedrock/eu-west-1
  • No data24h 96.81%
    7d
    30d
    p50
    246ms (best)
    p99
    867ms (best)
    Throughput
    105 tps
    Context
    131K
    $/Mtok
    $0.050 / $0.200
    together

Measured 2m ago. Latency and throughput are OpenRouter's rolling 30-minute windows. ◆ marks the best value in each column among these endpoints.

Price vs latency

One dot per endpoint at the latest round · further left = faster, lower = cheaper · hover a dot for its tag

History

Availability

Axis starts at 90% so small differences stay visible; it expands to 0 when an endpoint falls below.

Time to first token

Logarithmic axis — the fleet spans two orders of magnitude. P50 is OpenRouter's rolling 30-minute window.

Throughput

Median tokens per second.