modelstatus.dev

Gemma 4 31B

Gemma 4 31B: providers compared

Every endpoint serving this model, side by side. ◆ marks the best value in each column — per model, never fleet-wide.

  • Cerebras100% (best)
    Healthy24h 100%
    7d
    100% (best)
    30d
    100% (best)
    p50
    425ms
    p99
    10.6s
    Throughput
    163 tps
    Context
    131K
    $/Mtok
    $0.990 / $1.49
    cerebras/fp16fp16
  • DeepInfra100% (best)
    Healthy24h 98.61%
    7d
    91.14%
    30d
    91.14%
    p50
    431ms
    p99
    5.29s
    Throughput
    40 tps
    Context
    262K
    $/Mtok
    $0.130 / $0.380
    deepinfra/fp8fp8
  • ModelRun100% (best)
    Healthy24h 99.92%
    7d
    99.67%
    30d
    99.67%
    p50
    87ms (best)
    p99
    1.58s (best)
    Throughput
    164 tps (best)
    Context
    262K
    $/Mtok
    $0.750 / $1.00
    modelrun/fp4fp4
  • Novita100% (best)
    Healthy24h 99.11%
    7d
    80.32%
    30d
    80.32%
    p50
    974ms
    p99
    8.63s
    Throughput
    9.0 tps
    Context
    262K
    $/Mtok
    $0.140 / $0.400
    novita/bf16bf16
  • OpenInference100% (best)
    Healthy24h 98.10%
    7d
    97.61%
    30d
    97.61%
    p50
    973ms
    p99
    30.6s
    Throughput
    18 tps
    Context
    262K
    $/Mtok
    $0.080 / $0.350 (best)
    open-inference/bf16bf16
  • SiliconFlow100% (best)
    Degraded24h 88.66%
    7d
    94.18%
    30d
    94.18%
    p50
    3.88s
    p99
    51.2s
    Throughput
    16 tps
    Context
    262K
    $/Mtok
    $0.130 / $0.400
    siliconflow/fp8fp8
  • Venice100% (best)
    Healthy24h 99.76%
    7d
    99.74%
    30d
    99.74%
    p50
    1.13s
    p99
    18.1s
    Throughput
    34 tps
    Context
    256K
    $/Mtok
    $0.120 / $0.360
    venice/bf16bf16
  • DeepInfra99.93%
    Healthy24h 99.84%
    7d
    99.54%
    30d
    99.54%
    p50
    900ms
    p99
    7.90s
    Throughput
    20 tps
    Context
    262K
    $/Mtok
    $0.090 / $0.340
    deepinfra/turbofp4
  • Friendli99.92%
    Healthy24h 99.90%
    7d
    99.58%
    30d
    99.58%
    p50
    626ms
    p99
    7.83s
    Throughput
    65 tps
    Context
    262K
    $/Mtok
    $0.140 / $0.400
    friendli
  • Crusoe99.60%
    Healthy24h 98.74%
    7d
    98.65%
    30d
    98.65%
    p50
    758ms
    p99
    42.6s
    Throughput
    23 tps
    Context
    262K
    $/Mtok
    $0.140 / $0.400
    crusoe
  • CoreWeave97.88%
    Healthy24h 98.27%
    7d
    98.49%
    30d
    98.49%
    p50
    1.02s
    p99
    26.6s
    Throughput
    28 tps
    Context
    262K
    $/Mtok
    $0.100 / $0.340
    coreweave/bf16bf16
  • Parasail95.90%
    Healthy24h 99.56%
    7d
    99.28%
    30d
    99.28%
    p50
    564ms
    p99
    10.8s
    Throughput
    23 tps
    Context
    262K
    $/Mtok
    $0.150 / $0.400
    parasail/fp8fp8
  • DeepInfra84.85%
    Down24h 77.66%
    7d
    88.77%
    30d
    88.77%
    p50
    2.98s
    p99
    16.3s
    Throughput
    22 tps
    Context
    131K
    $/Mtok
    $0.270 / $0.760
    deepinfra/ultrafp8
  • Down24h 90.47%
    7d
    64.48%
    30d
    64.48%
    p50
    3.05s
    p99
    117s
    Throughput
    10 tps
    Context
    131K
    $/Mtok
    $0.120 / $0.370
    chutes/fp4fp4
  • Down24h 73.32%
    7d
    84.66%
    30d
    84.66%
    p50
    1.66s
    p99
    12.8s
    Throughput
    30 tps
    Context
    262K
    $/Mtok
    $0.150 / $0.460
    phala
  • Degraded24h 98.15%
    7d
    96.94%
    30d
    96.94%
    p50
    2.97s
    p99
    27.1s
    Throughput
    44 tps
    Context
    262K
    $/Mtok
    $0.380 / $1.15
    sambanova
  • No data24h 92.88%
    7d
    95.44%
    30d
    95.44%
    p50
    259ms
    p99
    22.2s
    Throughput
    50 tps
    Context
    262K
    $/Mtok
    $0.280 / $0.860
    together

Measured 2m ago. Latency and throughput are OpenRouter's rolling 30-minute windows. ◆ marks the best value in each column among these endpoints.

Price vs latency

One dot per endpoint at the latest round · further left = faster, lower = cheaper · hover a dot for its tag

History

Availability

Axis starts at 90% so small differences stay visible; it expands to 0 when an endpoint falls below.

Time to first token

Logarithmic axis — the fleet spans two orders of magnitude. P50 is OpenRouter's rolling 30-minute window, averaged over the hour.

Throughput

Median tokens per second.