Llama 4 Scout
meta-llama/llama-4-scout
Compare with Llama 4 Maverick, Llama Guard 4 12B, Llama 3.3 70B Instruct
CompareLlama 4 Scout
| Provider | State | Uptime 5m | 24h | p50 latency | p99 | Throughput | Context | $/Mtok |
|---|---|---|---|---|---|---|---|---|
DeepInfradeepinfra/fp8fp8 | Healthy | 100% | 99.75% | 501ms | 3.32s | 22 tps | 328K | $0.100 / $0.300 |
Googlegoogle-vertex/us-east5 | Healthy | 100% | 99.78% | 471ms | 5.16s | 41 tps | 1.3M | $0.250 / $0.700 |
Novitanovita/bf16bf16 | Healthy | 100% | 99.52% | 676ms | 2.72s | 21 tps | 131K | $0.180 / $0.590 |
Groqgroq | Healthy | 98.71% | 99.82% | 454ms | 2.04s | 102 tps | 131K | $0.110 / $0.340 |
- DeepInfra100%Healthy24h 99.75%
- p50
- 501ms
- p99
- 3.32s
- Throughput
- 22 tps
- Context
- 328K
- $/Mtok
- $0.100 / $0.300
deepinfra/fp8fp8 - Google100%Healthy24h 99.78%
- p50
- 471ms
- p99
- 5.16s
- Throughput
- 41 tps
- Context
- 1.3M
- $/Mtok
- $0.250 / $0.700
google-vertex/us-east5 - Novita100%Healthy24h 99.52%
- p50
- 676ms
- p99
- 2.72s
- Throughput
- 21 tps
- Context
- 131K
- $/Mtok
- $0.180 / $0.590
novita/bf16bf16 - Groq98.71%Healthy24h 99.82%
- p50
- 454ms
- p99
- 2.04s
- Throughput
- 102 tps
- Context
- 131K
- $/Mtok
- $0.110 / $0.340
groq
Measured 1m ago. Latency and throughput are OpenRouter's rolling 30-minute windows.
History
Availability
Axis starts at 90% so small differences stay visible; it expands to 0 when an endpoint falls below.
Time to first token
Logarithmic axis — the fleet spans two orders of magnitude. p50 and p99 are OpenRouter's rolling 30-minute windows.
Throughput
Median tokens per second.