MiMo-V2.5-Pro
xiaomi/mimo-v2.5-pro
Compare with MiMo-V2.5, GLM 5.2, DeepSeek V4 Flash 0731
CompareMiMo-V2.5-Pro
| Provider | State | Uptime 5m | 24h | p50 latency | p99 | Throughput | Context | $/Mtok |
|---|---|---|---|---|---|---|---|---|
AtlasCloudatlas-cloud/fp8fp8 | Degraded | 100% | 98.66% | 3.37s | 72.9s | 23 tps | 1.0M | $0.435 / $0.870 |
DeepInfradeepinfra/fp8fp8 | Healthy | 100% | 97.28% | 1.27s | 39.0s | 45 tps | 1.0M | $1.00 / $3.00 |
DigitalOceandigitalocean | Healthy | 100% | 98.71% | 788ms | 5.28s | 35 tps | 262K | $0.400 / $1.50 |
Novitanovita | Healthy | 100% | 99.39% | 3.14s | 11.3s | 29 tps | 1.0M | $0.480 / $0.960 |
Xiaomixiaomi/fp8fp8 | Healthy | 98.50% | 99.60% | 3.00s | 23.7s | 33 tps | 1.0M | $0.435 / $0.870 |
GMICloudgmicloud/bf16bf16 | Degraded | 96.72% | 94.00% | 4.02s | 8.13s | 15 tps | 1.1M | $0.304 / $0.609 |
StreamLakestreamlake | Degraded | — | 94.23% | 2.33s | 8.88s | 7.0 tps | 1.0M | $0.522 / $1.04 |
- AtlasCloud100%Degraded24h 98.66%
- p50
- 3.37s
- p99
- 72.9s
- Throughput
- 23 tps
- Context
- 1.0M
- $/Mtok
- $0.435 / $0.870
atlas-cloud/fp8fp8 - DeepInfra100%Healthy24h 97.28%
- p50
- 1.27s
- p99
- 39.0s
- Throughput
- 45 tps
- Context
- 1.0M
- $/Mtok
- $1.00 / $3.00
deepinfra/fp8fp8 - DigitalOcean100%Healthy24h 98.71%
- p50
- 788ms
- p99
- 5.28s
- Throughput
- 35 tps
- Context
- 262K
- $/Mtok
- $0.400 / $1.50
digitalocean - Novita100%Healthy24h 99.39%
- p50
- 3.14s
- p99
- 11.3s
- Throughput
- 29 tps
- Context
- 1.0M
- $/Mtok
- $0.480 / $0.960
novita - Xiaomi98.50%Healthy24h 99.60%
- p50
- 3.00s
- p99
- 23.7s
- Throughput
- 33 tps
- Context
- 1.0M
- $/Mtok
- $0.435 / $0.870
xiaomi/fp8fp8 - GMICloud96.72%Degraded24h 94.00%
- p50
- 4.02s
- p99
- 8.13s
- Throughput
- 15 tps
- Context
- 1.1M
- $/Mtok
- $0.304 / $0.609
gmicloud/bf16bf16 - StreamLake—Degraded24h 94.23%
- p50
- 2.33s
- p99
- 8.88s
- Throughput
- 7.0 tps
- Context
- 1.0M
- $/Mtok
- $0.522 / $1.04
streamlake
Measured 1m ago. Latency and throughput are OpenRouter's rolling 30-minute windows.
History
Availability
Axis starts at 90% so small differences stay visible; it expands to 0 when an endpoint falls below.
Time to first token
Logarithmic axis — the fleet spans two orders of magnitude. p50 and p99 are OpenRouter's rolling 30-minute windows, averaged over the hour.
Throughput
Median tokens per second.