Llama 4 Scout
meta-llama/llama-4-scout
Compare with Llama 4 Maverick, Llama Guard 4 12B, Llama 3.3 70B Instruct
CompareLlama 4 Scout
| Provider | State | Uptime 5m | 24h | p50 latency | p99 | Throughput | Context | $/Mtok |
|---|---|---|---|---|---|---|---|---|
DeepInfradeepinfra/fp8fp8 | Healthy | 100% | 99.75% | 489ms | 3.29s | 23 tps | 328K | $0.100 / $0.300 |
Googlegoogle-vertex/us-east5 | Healthy | 100% | 99.78% | 476ms | 4.78s | 40 tps | 1.3M | $0.250 / $0.700 |
Groqgroq | Healthy | 100% | 99.82% | 441ms | 2.11s | 108 tps | 131K | $0.110 / $0.340 |
Novitanovita/bf16bf16 | Healthy | 100% | 99.52% | 670ms | 2.54s | 23 tps | 131K | $0.180 / $0.590 |
- DeepInfra100%Healthy24h 99.75%
- p50
- 489ms
- p99
- 3.29s
- Throughput
- 23 tps
- Context
- 328K
- $/Mtok
- $0.100 / $0.300
deepinfra/fp8fp8 - Google100%Healthy24h 99.78%
- p50
- 476ms
- p99
- 4.78s
- Throughput
- 40 tps
- Context
- 1.3M
- $/Mtok
- $0.250 / $0.700
google-vertex/us-east5 - Groq100%Healthy24h 99.82%
- p50
- 441ms
- p99
- 2.11s
- Throughput
- 108 tps
- Context
- 131K
- $/Mtok
- $0.110 / $0.340
groq - Novita100%Healthy24h 99.52%
- p50
- 670ms
- p99
- 2.54s
- Throughput
- 23 tps
- Context
- 131K
- $/Mtok
- $0.180 / $0.590
novita/bf16bf16
Measured 5m ago. Latency and throughput are OpenRouter's rolling 30-minute windows.
History
Availability
Axis starts at 90% so small differences stay visible; it expands to 0 when an endpoint falls below.
Time to first token
Logarithmic axis — the fleet spans two orders of magnitude. p50 and p99 are OpenRouter's rolling 30-minute windows, averaged over the hour.
Throughput
Median tokens per second.