modelstatus.dev

Llama 4 Scout

meta-llama/llama-4-scout

Compare with Llama 4 Maverick, Llama Guard 4 12B, Llama 3.3 70B Instruct

CompareLlama 4 Scout
  • DeepInfra100%
    Healthy24h 99.75%
    p50
    489ms
    p99
    3.29s
    Throughput
    23 tps
    Context
    328K
    $/Mtok
    $0.100 / $0.300
    deepinfra/fp8fp8
  • Google100%
    Healthy24h 99.78%
    p50
    476ms
    p99
    4.78s
    Throughput
    40 tps
    Context
    1.3M
    $/Mtok
    $0.250 / $0.700
    google-vertex/us-east5
  • Groq100%
    Healthy24h 99.82%
    p50
    441ms
    p99
    2.11s
    Throughput
    108 tps
    Context
    131K
    $/Mtok
    $0.110 / $0.340
    groq
  • Novita100%
    Healthy24h 99.52%
    p50
    670ms
    p99
    2.54s
    Throughput
    23 tps
    Context
    131K
    $/Mtok
    $0.180 / $0.590
    novita/bf16bf16

Measured 5m ago. Latency and throughput are OpenRouter's rolling 30-minute windows.

Availability

Axis starts at 90% so small differences stay visible; it expands to 0 when an endpoint falls below.

Time to first token

Logarithmic axis — the fleet spans two orders of magnitude. p50 and p99 are OpenRouter's rolling 30-minute windows, averaged over the hour.

Throughput

Median tokens per second.