Llama 4 Scout
meta-llama/llama-4-scout
Compare with Llama 4 Maverick, Llama Guard 4 12B, Llama 3.3 70B Instruct
CompareLlama 4 Scout
| Provider | State | Uptime 5m | 24h | p50 latency | p99 | Throughput | Context | $/Mtok |
|---|---|---|---|---|---|---|---|---|
DeepInfradeepinfra/fp8fp8 | Healthy | 100% | 99.75% | 486ms | 1.75s | 5.0 tps | 328K | $0.100 / $0.300 |
Googlegoogle-vertex/us-east5 | Healthy | 100% | 99.77% | 590ms | 2.73s | 41 tps | 1.3M | $0.250 / $0.700 |
Novitanovita/bf16bf16 | Healthy | 99.78% | 99.52% | 878ms | 2.79s | 8.0 tps | 131K | $0.180 / $0.590 |
Groqgroq | Healthy | 99.62% | 99.84% | 461ms | 2.27s | 83 tps | 131K | $0.110 / $0.340 |
- DeepInfra100%Healthy24h 99.75%
- p50
- 486ms
- p99
- 1.75s
- Throughput
- 5.0 tps
- Context
- 328K
- $/Mtok
- $0.100 / $0.300
deepinfra/fp8fp8 - Google100%Healthy24h 99.77%
- p50
- 590ms
- p99
- 2.73s
- Throughput
- 41 tps
- Context
- 1.3M
- $/Mtok
- $0.250 / $0.700
google-vertex/us-east5 - Novita99.78%Healthy24h 99.52%
- p50
- 878ms
- p99
- 2.79s
- Throughput
- 8.0 tps
- Context
- 131K
- $/Mtok
- $0.180 / $0.590
novita/bf16bf16 - Groq99.62%Healthy24h 99.84%
- p50
- 461ms
- p99
- 2.27s
- Throughput
- 83 tps
- Context
- 131K
- $/Mtok
- $0.110 / $0.340
groq
Measured 4m ago. Latency and throughput are OpenRouter's rolling 30-minute windows.
History
Availability
Axis starts at 90% so small differences stay visible; it expands to 0 when an endpoint falls below.
Time to first token
Logarithmic axis — the fleet spans two orders of magnitude. p50 and p99 are OpenRouter's rolling 30-minute windows.
Throughput
Median tokens per second.