modelstatus.dev

Gemma 4 26B A4B

google/gemma-4-26b-a4b-it

Compare with Gemma 4 31B, Gemma 4 26B A4B (free), Gemma 4 31B (free)

CompareGemma 4 26B A4B
  • Cloudflare100%
    Down24h 92.62%
    p50
    1.65s
    p99
    31.5s
    Throughput
    35 tps
    Context
    256K
    $/Mtok
    $0.100 / $0.300
    cloudflare
  • DeepInfra99.82%
    Healthy24h 99.09%
    p50
    539ms
    p99
    7.15s
    Throughput
    15 tps
    Context
    262K
    $/Mtok
    $0.070 / $0.340
    deepinfra/fp8fp8
  • Venice99.81%
    Healthy24h 98.87%
    p50
    1.02s
    p99
    9.71s
    Throughput
    11 tps
    Context
    256K
    $/Mtok
    $0.130 / $0.400
    venice/bf16bf16
  • NextBit99.64%
    Healthy24h 99.64%
    p50
    664ms
    p99
    3.42s
    Throughput
    36 tps
    Context
    262K
    $/Mtok
    $0.100 / $0.400
    nextbit/bf16bf16
  • SiliconFlow99.55%
    Healthy24h 98.83%
    p50
    1.24s
    p99
    5.69s
    Throughput
    30 tps
    Context
    262K
    $/Mtok
    $0.120 / $0.400
    siliconflow/fp8fp8
  • Parasail98.52%
    Healthy24h 98.77%
    p50
    896ms
    p99
    10.7s
    Throughput
    11 tps
    Context
    262K
    $/Mtok
    $0.130 / $0.400
    parasail/bf16bf16
  • Google96.71%
    Healthy24h 97.65%
    p50
    749ms
    p99
    11.7s
    Throughput
    23 tps
    Context
    262K
    $/Mtok
    $0.150 / $0.600
    google-vertex/global
  • Novita50.10%
    Down24h 93.67%
    p50
    1.42s
    p99
    9.54s
    Throughput
    15 tps
    Context
    262K
    $/Mtok
    $0.130 / $0.400
    novita/bf16bf16

Measured 1m ago. Latency and throughput are OpenRouter's rolling 30-minute windows.

Availability

Axis starts at 90% so small differences stay visible; it expands to 0 when an endpoint falls below.

Time to first token

Logarithmic axis — the fleet spans two orders of magnitude. p50 and p99 are OpenRouter's rolling 30-minute windows.

Throughput

Median tokens per second.