Llama 4 Scout
meta-llama/llama-4-scout
Compare with Llama 4 Maverick, Llama Guard 4 12B, Llama 3.3 70B Instruct
CompareLlama 4 Scout
| Provider | State | Uptime 5m | 24h | p50 latency | p99 | Throughput | Context | $/Mtok |
|---|---|---|---|---|---|---|---|---|
Googlegoogle-vertex/us-east5 | Healthy | 100% | 99.78% | 463ms | 2.48s | 59 tps | 1.3M | $0.250 / $0.700 |
Novitanovita/bf16bf16 | Healthy | 100% | 99.53% | 585ms | 2.21s | 27 tps | 131K | $0.180 / $0.590 |
DeepInfradeepinfra/fp8fp8 | Healthy | 99.91% | 99.75% | 381ms | 2.69s | 29 tps | 328K | $0.100 / $0.300 |
Groqgroq | Healthy | 99.71% | 99.82% | 422ms | 2.01s | 135 tps | 131K | $0.110 / $0.340 |
- Google100%Healthy24h 99.78%
- p50
- 463ms
- p99
- 2.48s
- Throughput
- 59 tps
- Context
- 1.3M
- $/Mtok
- $0.250 / $0.700
google-vertex/us-east5 - Novita100%Healthy24h 99.53%
- p50
- 585ms
- p99
- 2.21s
- Throughput
- 27 tps
- Context
- 131K
- $/Mtok
- $0.180 / $0.590
novita/bf16bf16 - DeepInfra99.91%Healthy24h 99.75%
- p50
- 381ms
- p99
- 2.69s
- Throughput
- 29 tps
- Context
- 328K
- $/Mtok
- $0.100 / $0.300
deepinfra/fp8fp8 - Groq99.71%Healthy24h 99.82%
- p50
- 422ms
- p99
- 2.01s
- Throughput
- 135 tps
- Context
- 131K
- $/Mtok
- $0.110 / $0.340
groq
Measured 4m ago. Latency and throughput are OpenRouter's rolling 30-minute windows.
History
Availability
Axis starts at 90% so small differences stay visible; it expands to 0 when an endpoint falls below.
Time to first token
Logarithmic axis — the fleet spans two orders of magnitude. p50 and p99 are OpenRouter's rolling 30-minute windows, averaged over the hour.
Throughput
Median tokens per second.