modelstatus.dev

Gemma 4 31B

Gemma 4 31B: providers compared

Every endpoint serving this model, side by side. ◆ marks the best value in each column — per model, never fleet-wide.

  • Cerebras100% (best)
    Healthy24h 100%
    7d
    100% (best)
    30d
    100% (best)
    p50
    429ms
    p99
    6.88s
    Throughput
    131 tps (best)
    Context
    131K
    $/Mtok
    $0.990 / $1.49
    cerebras/fp16fp16
  • DeepInfra100% (best)
    Healthy24h 98.41%
    7d
    90.99%
    30d
    90.99%
    p50
    444ms
    p99
    5.53s
    Throughput
    42 tps
    Context
    262K
    $/Mtok
    $0.130 / $0.380
    deepinfra/fp8fp8
  • ModelRun100% (best)
    Healthy24h 99.92%
    7d
    99.66%
    30d
    99.66%
    p50
    225ms (best)
    p99
    3.54s (best)
    Throughput
    74 tps
    Context
    262K
    $/Mtok
    $0.750 / $1.00
    modelrun/fp4fp4
  • SambaNova100% (best)
    Degraded24h 97.99%
    7d
    97.09%
    30d
    97.09%
    p50
    3.67s
    p99
    31.3s
    Throughput
    18 tps
    Context
    262K
    $/Mtok
    $0.380 / $1.15
    sambanova
  • Together100% (best)
    Healthy24h 92.36%
    7d
    95.49%
    30d
    95.49%
    p50
    702ms
    p99
    15.2s
    Throughput
    31 tps
    Context
    262K
    $/Mtok
    $0.280 / $0.860
    together
  • Friendli99.87%
    Healthy24h 99.83%
    7d
    99.59%
    30d
    99.59%
    p50
    2.29s
    p99
    12.1s
    Throughput
    39 tps
    Context
    262K
    $/Mtok
    $0.140 / $0.400
    friendli
  • Venice99.78%
    Healthy24h 99.76%
    7d
    99.74%
    30d
    99.74%
    p50
    1.11s
    p99
    18.7s
    Throughput
    36 tps
    Context
    256K
    $/Mtok
    $0.120 / $0.360
    venice/bf16bf16
  • DeepInfra99.76%
    Healthy24h 99.84%
    7d
    99.54%
    30d
    99.54%
    p50
    776ms
    p99
    5.89s
    Throughput
    16 tps
    Context
    262K
    $/Mtok
    $0.090 / $0.340
    deepinfra/turbofp4
  • Healthy24h 97.72%
    7d
    97.58%
    30d
    97.58%
    p50
    1.02s
    p99
    36.6s
    Throughput
    18 tps
    Context
    262K
    $/Mtok
    $0.080 / $0.350 (best)
    open-inference/bf16bf16
  • Novita99.62%
    Healthy24h 99.02%
    7d
    80.00%
    30d
    80.00%
    p50
    977ms
    p99
    8.68s
    Throughput
    13 tps
    Context
    262K
    $/Mtok
    $0.140 / $0.400
    novita/bf16bf16
  • Parasail99.56%
    Healthy24h 99.59%
    7d
    99.29%
    30d
    99.29%
    p50
    714ms
    p99
    9.70s
    Throughput
    20 tps
    Context
    262K
    $/Mtok
    $0.150 / $0.400
    parasail/fp8fp8
  • CoreWeave98.49%
    Degraded24h 98.20%
    7d
    98.51%
    30d
    98.51%
    p50
    1.36s
    p99
    30.8s
    Throughput
    27 tps
    Context
    262K
    $/Mtok
    $0.100 / $0.340
    coreweave/bf16bf16
  • Crusoe93.53%
    Degraded24h 98.81%
    7d
    98.66%
    30d
    98.66%
    p50
    1.18s
    p99
    27.3s
    Throughput
    21 tps
    Context
    262K
    $/Mtok
    $0.140 / $0.400
    crusoe
  • Chutes90.09%
    Degraded24h 89.19%
    7d
    64.54%
    30d
    64.54%
    p50
    2.21s
    p99
    23.2s
    Throughput
    12 tps
    Context
    131K
    $/Mtok
    $0.120 / $0.370
    chutes/fp4fp4
  • DeepInfra87.14%
    Down24h 79.17%
    7d
    88.70%
    30d
    88.70%
    p50
    5.12s
    p99
    20.2s
    Throughput
    16 tps
    Context
    131K
    $/Mtok
    $0.270 / $0.760
    deepinfra/ultrafp8
  • Down24h 84.55%
    7d
    94.20%
    30d
    94.20%
    p50
    2.36s
    p99
    41.8s
    Throughput
    23 tps
    Context
    262K
    $/Mtok
    $0.130 / $0.400
    siliconflow/fp8fp8
  • Phala7.39%
    Down24h 75.22%
    7d
    85.43%
    30d
    85.43%
    p50
    1.38s
    p99
    14.3s
    Throughput
    43 tps
    Context
    262K
    $/Mtok
    $0.150 / $0.460
    phala

Measured 2m ago. Latency and throughput are OpenRouter's rolling 30-minute windows. ◆ marks the best value in each column among these endpoints.

Price vs latency

One dot per endpoint at the latest round · further left = faster, lower = cheaper · hover a dot for its tag

History

Availability

Axis starts at 90% so small differences stay visible; it expands to 0 when an endpoint falls below.

Time to first token

Logarithmic axis — the fleet spans two orders of magnitude. P50 is OpenRouter's rolling 30-minute window.

Throughput

Median tokens per second.