Nemotron 3 Ultra
nvidia/nemotron-3-ultra-550b-a55b
Compare with Nemotron 3 Ultra (batch), Nemotron 3 Ultra (free), Nemotron 3 Nano 30B A3B
CompareNemotron 3 Ultra
| Provider | State | Uptime 5m | 24h | p50 latency | p99 | Throughput | Context | $/Mtok |
|---|---|---|---|---|---|---|---|---|
BaseTenbaseten/fp4fp4 | Healthy | 100% | 98.86% | 756ms | 10.2s | 60 tps | 203K | $0.600 / $2.40 |
Togethertogether | Healthy | 100% | 96.99% | 830ms | 8.31s | 91 tps | 512K | $0.600 / $3.60 |
DeepInfradeepinfra/fp4fp4 | Degraded | 91.36% | 92.04% | 2.65s | 45.0s | 54 tps | 262K | $0.500 / $2.20 |
Venicevenice/fp8fp8 | Down | — | 87.04% | 1.40s | 119s | 21 tps | 256K | $0.625 / $3.13 |
- BaseTen100%Healthy24h 98.86%
- p50
- 756ms
- p99
- 10.2s
- Throughput
- 60 tps
- Context
- 203K
- $/Mtok
- $0.600 / $2.40
baseten/fp4fp4 - Together100%Healthy24h 96.99%
- p50
- 830ms
- p99
- 8.31s
- Throughput
- 91 tps
- Context
- 512K
- $/Mtok
- $0.600 / $3.60
together - DeepInfra91.36%Degraded24h 92.04%
- p50
- 2.65s
- p99
- 45.0s
- Throughput
- 54 tps
- Context
- 262K
- $/Mtok
- $0.500 / $2.20
deepinfra/fp4fp4 - Venice—Down24h 87.04%
- p50
- 1.40s
- p99
- 119s
- Throughput
- 21 tps
- Context
- 256K
- $/Mtok
- $0.625 / $3.13
venice/fp8fp8
Measured 53s ago. Latency and throughput are OpenRouter's rolling 30-minute windows.
History
Availability
Axis starts at 90% so small differences stay visible; it expands to 0 when an endpoint falls below.
Time to first token
Logarithmic axis — the fleet spans two orders of magnitude. p50 and p99 are OpenRouter's rolling 30-minute windows, averaged over the hour.
Throughput
Median tokens per second.