modelstatus.dev

Llama 4 Scout

meta-llama/llama-4-scout

Compare with Llama 4 Maverick, Llama Guard 4 12B, Llama 3.3 70B Instruct

CompareLlama 4 Scout
  • DeepInfra100%
    Healthy24h 99.75%
    p50
    486ms
    p99
    1.75s
    Throughput
    5.0 tps
    Context
    328K
    $/Mtok
    $0.100 / $0.300
    deepinfra/fp8fp8
  • Google100%
    Healthy24h 99.77%
    p50
    590ms
    p99
    2.73s
    Throughput
    41 tps
    Context
    1.3M
    $/Mtok
    $0.250 / $0.700
    google-vertex/us-east5
  • Novita99.78%
    Healthy24h 99.52%
    p50
    878ms
    p99
    2.79s
    Throughput
    8.0 tps
    Context
    131K
    $/Mtok
    $0.180 / $0.590
    novita/bf16bf16
  • Groq99.62%
    Healthy24h 99.84%
    p50
    461ms
    p99
    2.27s
    Throughput
    83 tps
    Context
    131K
    $/Mtok
    $0.110 / $0.340
    groq

Measured 4m ago. Latency and throughput are OpenRouter's rolling 30-minute windows.

Availability

Axis starts at 90% so small differences stay visible; it expands to 0 when an endpoint falls below.

Time to first token

Logarithmic axis — the fleet spans two orders of magnitude. p50 and p99 are OpenRouter's rolling 30-minute windows.

Throughput

Median tokens per second.