Nemotron 3 Ultra
nvidia/nemotron-3-ultra-550b-a55b
Compare with Nemotron 3 Ultra (batch), Nemotron 3 Ultra (free), Nemotron 3 Nano 30B A3B
CompareNemotron 3 Ultra
| Provider | State | Uptime 5m | 24h | p50 latency | p99 | Throughput | Context | $/Mtok |
|---|---|---|---|---|---|---|---|---|
BaseTenbaseten/fp4fp4 | Healthy | 100% | 98.86% | 533ms | 5.30s | 55 tps | 203K | $0.600 / $2.40 |
Togethertogether | Healthy | 99.00% | 98.57% | 690ms | 8.23s | 79 tps | 512K | $0.600 / $3.60 |
DeepInfradeepinfra/fp4fp4 | Degraded | 94.77% | 92.13% | 2.45s | 21.9s | 63 tps | 262K | $0.500 / $2.20 |
Venicevenice/fp8fp8 | Degraded | 90.00% | 87.47% | 1.37s | 11.3s | 57 tps | 256K | $0.625 / $3.13 |
- BaseTen100%Healthy24h 98.86%
- p50
- 533ms
- p99
- 5.30s
- Throughput
- 55 tps
- Context
- 203K
- $/Mtok
- $0.600 / $2.40
baseten/fp4fp4 - Together99.00%Healthy24h 98.57%
- p50
- 690ms
- p99
- 8.23s
- Throughput
- 79 tps
- Context
- 512K
- $/Mtok
- $0.600 / $3.60
together - DeepInfra94.77%Degraded24h 92.13%
- p50
- 2.45s
- p99
- 21.9s
- Throughput
- 63 tps
- Context
- 262K
- $/Mtok
- $0.500 / $2.20
deepinfra/fp4fp4 - Venice90.00%Degraded24h 87.47%
- p50
- 1.37s
- p99
- 11.3s
- Throughput
- 57 tps
- Context
- 256K
- $/Mtok
- $0.625 / $3.13
venice/fp8fp8
Measured 4m ago. Latency and throughput are OpenRouter's rolling 30-minute windows.
History
Availability
Axis starts at 90% so small differences stay visible; it expands to 0 when an endpoint falls below.
Time to first token
Logarithmic axis — the fleet spans two orders of magnitude. p50 and p99 are OpenRouter's rolling 30-minute windows, averaged over the hour.
Throughput
Median tokens per second.