modelstatus.dev

Llama 4 Scout

meta-llama/llama-4-scout

Compare with Llama 4 Maverick, Llama Guard 4 12B, Llama 3.3 70B Instruct

CompareLlama 4 Scout
  • DeepInfra100%
    Healthy24h 99.75%
    p50
    501ms
    p99
    3.32s
    Throughput
    22 tps
    Context
    328K
    $/Mtok
    $0.100 / $0.300
    deepinfra/fp8fp8
  • Google100%
    Healthy24h 99.78%
    p50
    471ms
    p99
    5.16s
    Throughput
    41 tps
    Context
    1.3M
    $/Mtok
    $0.250 / $0.700
    google-vertex/us-east5
  • Novita100%
    Healthy24h 99.52%
    p50
    676ms
    p99
    2.72s
    Throughput
    21 tps
    Context
    131K
    $/Mtok
    $0.180 / $0.590
    novita/bf16bf16
  • Groq98.71%
    Healthy24h 99.82%
    p50
    454ms
    p99
    2.04s
    Throughput
    102 tps
    Context
    131K
    $/Mtok
    $0.110 / $0.340
    groq

Measured 1m ago. Latency and throughput are OpenRouter's rolling 30-minute windows.

Availability

Axis starts at 90% so small differences stay visible; it expands to 0 when an endpoint falls below.

Time to first token

Logarithmic axis — the fleet spans two orders of magnitude. p50 and p99 are OpenRouter's rolling 30-minute windows.

Throughput

Median tokens per second.