modelstatus.dev

Llama 4 Scout

meta-llama/llama-4-scout

Compare with Llama 4 Maverick, Llama Guard 4 12B, Llama 3.3 70B Instruct

CompareLlama 4 Scout
  • Google100%
    Healthy24h 99.78%
    p50
    463ms
    p99
    2.48s
    Throughput
    59 tps
    Context
    1.3M
    $/Mtok
    $0.250 / $0.700
    google-vertex/us-east5
  • Novita100%
    Healthy24h 99.53%
    p50
    585ms
    p99
    2.21s
    Throughput
    27 tps
    Context
    131K
    $/Mtok
    $0.180 / $0.590
    novita/bf16bf16
  • DeepInfra99.91%
    Healthy24h 99.75%
    p50
    381ms
    p99
    2.69s
    Throughput
    29 tps
    Context
    328K
    $/Mtok
    $0.100 / $0.300
    deepinfra/fp8fp8
  • Groq99.71%
    Healthy24h 99.82%
    p50
    422ms
    p99
    2.01s
    Throughput
    135 tps
    Context
    131K
    $/Mtok
    $0.110 / $0.340
    groq

Measured 4m ago. Latency and throughput are OpenRouter's rolling 30-minute windows.

Availability

Axis starts at 90% so small differences stay visible; it expands to 0 when an endpoint falls below.

Time to first token

Logarithmic axis — the fleet spans two orders of magnitude. p50 and p99 are OpenRouter's rolling 30-minute windows, averaged over the hour.

Throughput

Median tokens per second.