modelstatus.dev

Gemma 4 26B A4B

google/gemma-4-26b-a4b-it

Compare with Gemma 4 31B, Gemma 4 26B A4B (free), Gemma 4 31B (free)

CompareGemma 4 26B A4B
  • DeepInfra99.93%
    Healthy24h 99.10%
    p50
    602ms
    p99
    8.65s
    Throughput
    11 tps
    Context
    262K
    $/Mtok
    $0.070 / $0.340
    deepinfra/fp8fp8
  • NextBit99.67%
    Healthy24h 99.65%
    p50
    766ms
    p99
    4.38s
    Throughput
    25 tps
    Context
    262K
    $/Mtok
    $0.100 / $0.400
    nextbit/bf16bf16
  • Novita99.51%
    Degraded24h 94.16%
    p50
    1.49s
    p99
    9.48s
    Throughput
    14 tps
    Context
    262K
    $/Mtok
    $0.130 / $0.400
    novita/bf16bf16
  • Venice99.51%
    Healthy24h 98.88%
    p50
    1.05s
    p99
    10.3s
    Throughput
    11 tps
    Context
    256K
    $/Mtok
    $0.130 / $0.400
    venice/bf16bf16
  • Cloudflare99.25%
    Degraded24h 93.27%
    p50
    961ms
    p99
    29.6s
    Throughput
    42 tps
    Context
    256K
    $/Mtok
    $0.100 / $0.300
    cloudflare
  • Parasail99.12%
    Healthy24h 98.78%
    p50
    802ms
    p99
    9.46s
    Throughput
    14 tps
    Context
    262K
    $/Mtok
    $0.130 / $0.400
    parasail/bf16bf16
  • SiliconFlow98.13%
    Healthy24h 98.84%
    p50
    1.41s
    p99
    10.8s
    Throughput
    24 tps
    Context
    262K
    $/Mtok
    $0.120 / $0.400
    siliconflow/fp8fp8
  • Google94.35%
    Degraded24h 97.71%
    p50
    668ms
    p99
    12.2s
    Throughput
    26 tps
    Context
    262K
    $/Mtok
    $0.150 / $0.600
    google-vertex/global

Measured 2m ago. Latency and throughput are OpenRouter's rolling 30-minute windows.

Availability

Axis starts at 90% so small differences stay visible; it expands to 0 when an endpoint falls below.

Time to first token

Logarithmic axis — the fleet spans two orders of magnitude. p50 and p99 are OpenRouter's rolling 30-minute windows, averaged over the hour.

Throughput

Median tokens per second.