Qwen2.5 VL 72B Instruct
qwen/qwen2.5-vl-72b-instruct
Compare with Qwen2.5 7B Instruct, Qwen2.5 72B Instruct, Qwen-Plus
CompareQwen2.5 VL 72B Instruct
| Provider | State | Uptime 5m | 24h | p50 latency | p99 | Throughput | Context | $/Mtok |
|---|---|---|---|---|---|---|---|---|
Parasailparasail/fp8fp8 | Healthy | 100% | 99.82% | 818ms | 3.85s | 26 tps | 128K | $0.800 / $1.00 |
Nebiusnebius/fp8fp8 | Healthy | 99.84% | 94.09% | 1.35s | 5.46s | 10 tps | 32K | $0.250 / $0.750 |
- Parasail100%Healthy24h 99.82%
- p50
- 818ms
- p99
- 3.85s
- Throughput
- 26 tps
- Context
- 128K
- $/Mtok
- $0.800 / $1.00
parasail/fp8fp8 - Nebius99.84%Healthy24h 94.09%
- p50
- 1.35s
- p99
- 5.46s
- Throughput
- 10 tps
- Context
- 32K
- $/Mtok
- $0.250 / $0.750
nebius/fp8fp8
Measured 4m ago. Latency and throughput are OpenRouter's rolling 30-minute windows.
History
Availability
Axis starts at 90% so small differences stay visible; it expands to 0 when an endpoint falls below.
Time to first token
Logarithmic axis — the fleet spans two orders of magnitude. p50 and p99 are OpenRouter's rolling 30-minute windows, averaged over the hour.
Throughput
Median tokens per second.