MiniMax M3: providers compared
Every endpoint serving this model, side by side. ◆ marks the best value in each column — per model, never fleet-wide.
Current state
| State | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Healthy | 100%◆ (best) | 99.68% | 99.53% | 99.53% | 1.31s | 40.3s | 75 tps | 524K | $0.300 / $1.20 | |
| Healthy | 100%◆ (best) | 99.68% | 97.50% | 97.50% | 529ms◆ (best) | 2.56s◆ (best) | 116 tps | 262K | $0.230 / $0.960◆ (best) | |
| Healthy | 100%◆ (best) | 98.05% | 98.27% | 98.27% | 552ms | 6.92s | 31 tps | 524K | $0.280 / $1.10 | |
| Healthy | 100%◆ (best) | 99.50% | 99.35% | 99.35% | 3.29s | 8.36s | 53 tps | 1.0M | $0.240 / $0.960 | |
| Healthy | 100%◆ (best) | 99.80% | 99.77% | 99.77% | 687ms | 18.7s | 119 tps◆ (best) | 1.0M | $0.750 / $3.00 | |
SambaNovasambanova | Healthy | 100%◆ (best) | 98.12% | 98.63% | 98.63% | 2.93s | 13.4s | 117 tps | 1.0M | $0.600 / $2.40 |
| Healthy | 100%◆ (best) | 99.88% | 99.82%◆ (best) | 99.82%◆ (best) | 1.56s | 12.8s | 85 tps | 524K | $0.300 / $1.20 | |
| Healthy | 99.72% | 99.60% | 99.66% | 99.66% | 2.21s | 22.3s | 49 tps | 1.0M | $0.300 / $1.20 | |
| Healthy | 99.34% | 99.51% | 99.64% | 99.64% | 1.21s | 13.0s | 69 tps | 524K | $0.300 / $1.20 | |
Togethertogether | Healthy | 98.98% | 99.63% | 99.76% | 99.76% | 908ms | 9.27s | 60 tps | 524K | $0.300 / $1.20 |
| Healthy | 98.93% | 98.17% | 97.90% | 97.90% | 1.60s | 26.4s | 37 tps | 1.0M | $0.300 / $1.20 | |
| Degraded | 94.00% | 98.92% | 98.47% | 98.47% | 1.57s | 7.34s | 72 tps | 1.0M | $0.300 / $1.20 | |
| Degraded | 92.52% | 96.43% | 97.48% | 97.48% | 2.31s | 21.8s | 21 tps | 256K | $0.255 / $1.02 |
- AtlasCloud100%◆ (best)Healthy24h 99.68%
- 7d
- 99.53%
- 30d
- 99.53%
- p50
- 1.31s
- p99
- 40.3s
- Throughput
- 75 tps
- Context
- 524K
- $/Mtok
- $0.300 / $1.20
atlas-cloud/fp8fp8 - CoreWeave100%◆ (best)Healthy24h 99.68%
- 7d
- 97.50%
- 30d
- 97.50%
- p50
- 529ms◆ (best)
- p99
- 2.56s◆ (best)
- Throughput
- 116 tps
- Context
- 262K
- $/Mtok
- $0.230 / $0.960◆ (best)
coreweave/fp4fp4 - DeepInfra100%◆ (best)Healthy24h 98.05%
- 7d
- 98.27%
- 30d
- 98.27%
- p50
- 552ms
- p99
- 6.92s
- Throughput
- 31 tps
- Context
- 524K
- $/Mtok
- $0.280 / $1.10
deepinfra/fp8fp8 - GMICloud100%◆ (best)Healthy24h 99.50%
- 7d
- 99.35%
- 30d
- 99.35%
- p50
- 3.29s
- p99
- 8.36s
- Throughput
- 53 tps
- Context
- 1.0M
- $/Mtok
- $0.240 / $0.960
gmicloud/fp8fp8 - ModelRun100%◆ (best)Healthy24h 99.80%
- 7d
- 99.77%
- 30d
- 99.77%
- p50
- 687ms
- p99
- 18.7s
- Throughput
- 119 tps◆ (best)
- Context
- 1.0M
- $/Mtok
- $0.750 / $3.00
modelrun/fp4fp4 - SambaNova100%◆ (best)Healthy24h 98.12%
- 7d
- 98.63%
- 30d
- 98.63%
- p50
- 2.93s
- p99
- 13.4s
- Throughput
- 117 tps
- Context
- 1.0M
- $/Mtok
- $0.600 / $2.40
sambanova - Venice100%◆ (best)Healthy24h 99.88%
- 7d
- 99.82%◆ (best)
- 30d
- 99.82%◆ (best)
- p50
- 1.56s
- p99
- 12.8s
- Throughput
- 85 tps
- Context
- 524K
- $/Mtok
- $0.300 / $1.20
venice/fp8fp8 - Novita99.72%Healthy24h 99.60%
- 7d
- 99.66%
- 30d
- 99.66%
- p50
- 2.21s
- p99
- 22.3s
- Throughput
- 49 tps
- Context
- 1.0M
- $/Mtok
- $0.300 / $1.20
novita/fp8fp8 - Minimax99.34%Healthy24h 99.51%
- 7d
- 99.64%
- 30d
- 99.64%
- p50
- 1.21s
- p99
- 13.0s
- Throughput
- 69 tps
- Context
- 524K
- $/Mtok
- $0.300 / $1.20
minimax/fp8fp8 - Together98.98%Healthy24h 99.63%
- 7d
- 99.76%
- 30d
- 99.76%
- p50
- 908ms
- p99
- 9.27s
- Throughput
- 60 tps
- Context
- 524K
- $/Mtok
- $0.300 / $1.20
together - Parasail98.93%Healthy24h 98.17%
- 7d
- 97.90%
- 30d
- 97.90%
- p50
- 1.60s
- p99
- 26.4s
- Throughput
- 37 tps
- Context
- 1.0M
- $/Mtok
- $0.300 / $1.20
parasail/fp8fp8 - StreamLake94.00%Degraded24h 98.92%
- 7d
- 98.47%
- 30d
- 98.47%
- p50
- 1.57s
- p99
- 7.34s
- Throughput
- 72 tps
- Context
- 1.0M
- $/Mtok
- $0.300 / $1.20
streamlake/fp8fp8 - Morph92.52%Degraded24h 96.43%
- 7d
- 97.48%
- 30d
- 97.48%
- p50
- 2.31s
- p99
- 21.8s
- Throughput
- 21 tps
- Context
- 256K
- $/Mtok
- $0.255 / $1.02
morph/fp4fp4
Measured 4m ago. Latency and throughput are OpenRouter's rolling 30-minute windows. ◆ marks the best value in each column among these endpoints.
Price vs latency
One dot per endpoint at the latest round · further left = faster, lower = cheaper · hover a dot for its tag
History
Availability
Axis starts at 90% so small differences stay visible; it expands to 0 when an endpoint falls below.
Time to first token
Logarithmic axis — the fleet spans two orders of magnitude. P50 is OpenRouter's rolling 30-minute window, averaged over the hour.
Throughput
Median tokens per second.