MiniMax M2
minimax/minimax-m2
Compare with MiniMax M2-her, MiniMax M2.1, MiniMax M2.5
CompareMiniMax M2
| Provider | State | Uptime 5m | 24h | p50 latency | p99 | Throughput | Context | $/Mtok |
|---|---|---|---|---|---|---|---|---|
Minimaxminimax/fp8fp8 | Healthy | 100% | 99.92% | 851ms | 3.61s | 46 tps | 205K | $0.255 / $1.02 |
Googlegoogle-vertex | No data | — | 99.98% | 265ms | 4.81s | 19 tps | 197K | $0.300 / $1.20 |
Novitanovita/fp8fp8 | No data | — | 99.62% | 1.64s | 4.45s | 67 tps | 205K | $0.300 / $1.20 |
- Minimax100%Healthy24h 99.92%
- p50
- 851ms
- p99
- 3.61s
- Throughput
- 46 tps
- Context
- 205K
- $/Mtok
- $0.255 / $1.02
minimax/fp8fp8 - Google—No data24h 99.98%
- p50
- 265ms
- p99
- 4.81s
- Throughput
- 19 tps
- Context
- 197K
- $/Mtok
- $0.300 / $1.20
google-vertex - Novita—No data24h 99.62%
- p50
- 1.64s
- p99
- 4.45s
- Throughput
- 67 tps
- Context
- 205K
- $/Mtok
- $0.300 / $1.20
novita/fp8fp8
Measured 1m ago. Latency and throughput are OpenRouter's rolling 30-minute windows.
History
2 endpoints reporting no data are omitted from the charts.
Availability
Axis starts at 90% so small differences stay visible; it expands to 0 when an endpoint falls below.
Time to first token
Logarithmic axis — the fleet spans two orders of magnitude. p50 and p99 are OpenRouter's rolling 30-minute windows.
Throughput
Median tokens per second.