Llama 3.1 70B Instruct
meta-llama/llama-3.1-70b-instruct
Compare with Llama 3.1 8B Instruct, Llama 3.2 3B Instruct, Llama 3.2 1B Instruct
CompareLlama 3.1 70B Instruct
| Provider | State | Uptime 5m | 24h | p50 latency | p99 | Throughput | Context | $/Mtok |
|---|---|---|---|---|---|---|---|---|
Amazon Bedrockamazon-bedrock | Healthy | 100% | 99.58% | 434ms | 1.05s | 29 tps | 131K | $0.720 / $0.720 |
CoreWeavecoreweave/bf16bf16 | Healthy | 100% | 99.57% | 225ms | 581ms | 28 tps | 128K | $0.800 / $0.800 |
DeepInfradeepinfra/turbofp8 | Healthy | 99.93% | 99.22% | 229ms | 1.49s | 22 tps | 131K | $0.400 / $0.400 |
- Amazon Bedrock100%Healthy24h 99.58%
- p50
- 434ms
- p99
- 1.05s
- Throughput
- 29 tps
- Context
- 131K
- $/Mtok
- $0.720 / $0.720
amazon-bedrock - CoreWeave100%Healthy24h 99.57%
- p50
- 225ms
- p99
- 581ms
- Throughput
- 28 tps
- Context
- 128K
- $/Mtok
- $0.800 / $0.800
coreweave/bf16bf16 - DeepInfra99.93%Healthy24h 99.22%
- p50
- 229ms
- p99
- 1.49s
- Throughput
- 22 tps
- Context
- 131K
- $/Mtok
- $0.400 / $0.400
deepinfra/turbofp8
Measured 1m ago. Latency and throughput are OpenRouter's rolling 30-minute windows.
History
Availability
Axis starts at 90% so small differences stay visible; it expands to 0 when an endpoint falls below.
Time to first token
Logarithmic axis — the fleet spans two orders of magnitude. p50 and p99 are OpenRouter's rolling 30-minute windows.
Throughput
Median tokens per second.