Llama 3.1 70B Instruct
meta-llama/llama-3.1-70b-instruct
Compare with Llama 3.1 8B Instruct, Llama 3.2 3B Instruct, Llama 3.2 1B Instruct
CompareLlama 3.1 70B Instruct
| Provider | State | Uptime 5m | 24h | p50 latency | p99 | Throughput | Context | $/Mtok |
|---|---|---|---|---|---|---|---|---|
CoreWeavecoreweave/bf16bf16 | Healthy | 100% | 99.57% | 226ms | 613ms | 27 tps | 128K | $0.800 / $0.800 |
DeepInfradeepinfra/turbofp8 | Healthy | 100% | 99.24% | 236ms | 1.49s | 21 tps | 131K | $0.400 / $0.400 |
Amazon Bedrockamazon-bedrock | Healthy | 99.77% | 99.59% | 352ms | 1.05s | 20 tps | 131K | $0.720 / $0.720 |
- CoreWeave100%Healthy24h 99.57%
- p50
- 226ms
- p99
- 613ms
- Throughput
- 27 tps
- Context
- 128K
- $/Mtok
- $0.800 / $0.800
coreweave/bf16bf16 - DeepInfra100%Healthy24h 99.24%
- p50
- 236ms
- p99
- 1.49s
- Throughput
- 21 tps
- Context
- 131K
- $/Mtok
- $0.400 / $0.400
deepinfra/turbofp8 - Amazon Bedrock99.77%Healthy24h 99.59%
- p50
- 352ms
- p99
- 1.05s
- Throughput
- 20 tps
- Context
- 131K
- $/Mtok
- $0.720 / $0.720
amazon-bedrock
Measured 3m ago. Latency and throughput are OpenRouter's rolling 30-minute windows.
History
Availability
Axis starts at 90% so small differences stay visible; it expands to 0 when an endpoint falls below.
Time to first token
Logarithmic axis — the fleet spans two orders of magnitude. p50 and p99 are OpenRouter's rolling 30-minute windows.
Throughput
Median tokens per second.