modelstatus.dev

Gemma 4 26B A4B

google/gemma-4-26b-a4b-it

Compare with Gemma 4 31B, Gemma 4 26B A4B (free), Gemma 4 31B (free)

CompareGemma 4 26B A4B
  • Novita100%
    Healthy24h 94.42%
    p50
    2.01s
    p99
    9.32s
    Throughput
    21 tps
    Context
    262K
    $/Mtok
    $0.130 / $0.400
    novita/bf16bf16
  • SiliconFlow99.87%
    Healthy24h 98.88%
    p50
    1.55s
    p99
    10.8s
    Throughput
    23 tps
    Context
    262K
    $/Mtok
    $0.120 / $0.400
    siliconflow/fp8fp8
  • NextBit99.79%
    Healthy24h 99.65%
    p50
    765ms
    p99
    3.34s
    Throughput
    24 tps
    Context
    262K
    $/Mtok
    $0.100 / $0.400
    nextbit/bf16bf16
  • Cloudflare99.75%
    Degraded24h 94.22%
    p50
    1.04s
    p99
    33.3s
    Throughput
    38 tps
    Context
    256K
    $/Mtok
    $0.100 / $0.300
    cloudflare
  • DeepInfra99.62%
    Healthy24h 99.11%
    p50
    602ms
    p99
    6.96s
    Throughput
    12 tps
    Context
    262K
    $/Mtok
    $0.070 / $0.340
    deepinfra/fp8fp8
  • Venice99.02%
    Healthy24h 98.90%
    p50
    1.13s
    p99
    10.8s
    Throughput
    13 tps
    Context
    256K
    $/Mtok
    $0.130 / $0.400
    venice/bf16bf16
  • Parasail98.65%
    Healthy24h 98.81%
    p50
    1.04s
    p99
    9.91s
    Throughput
    10 tps
    Context
    262K
    $/Mtok
    $0.130 / $0.400
    parasail/bf16bf16
  • Google95.11%
    Healthy24h 97.82%
    p50
    594ms
    p99
    10.5s
    Throughput
    25 tps
    Context
    262K
    $/Mtok
    $0.150 / $0.600
    google-vertex/global

Measured 5m ago. Latency and throughput are OpenRouter's rolling 30-minute windows.

Availability

Axis starts at 90% so small differences stay visible; it expands to 0 when an endpoint falls below.

Time to first token

Logarithmic axis — the fleet spans two orders of magnitude. p50 and p99 are OpenRouter's rolling 30-minute windows.

Throughput

Median tokens per second.