Nemotron 3 Ultra
nvidia/nemotron-3-ultra-550b-a55b
Compare with Nemotron 3 Ultra (batch), Nemotron 3 Ultra (free), Nemotron 3 Nano 30B A3B
CompareNemotron 3 Ultra
| Provider | State | Uptime 5m | 24h | p50 latency | p99 | Throughput | Context | $/Mtok |
|---|---|---|---|---|---|---|---|---|
Togethertogether | Healthy | 99.75% | 96.89% | 896ms | 9.72s | 91 tps | 512K | $0.600 / $3.60 |
DeepInfradeepinfra/fp4fp4 | Degraded | 93.76% | 91.95% | 2.85s | 74.7s | 50 tps | 262K | $0.500 / $2.20 |
BaseTenbaseten/fp4fp4 | Degraded | — | 98.96% | 2.09s | 21.8s | 35 tps | 203K | $0.600 / $2.40 |
Venicevenice/fp8fp8 | Down | — | 87.30% | 2.66s | 29.6s | 59 tps | 256K | $0.625 / $3.13 |
- Together99.75%Healthy24h 96.89%
- p50
- 896ms
- p99
- 9.72s
- Throughput
- 91 tps
- Context
- 512K
- $/Mtok
- $0.600 / $3.60
together - DeepInfra93.76%Degraded24h 91.95%
- p50
- 2.85s
- p99
- 74.7s
- Throughput
- 50 tps
- Context
- 262K
- $/Mtok
- $0.500 / $2.20
deepinfra/fp4fp4 - BaseTen—Degraded24h 98.96%
- p50
- 2.09s
- p99
- 21.8s
- Throughput
- 35 tps
- Context
- 203K
- $/Mtok
- $0.600 / $2.40
baseten/fp4fp4 - Venice—Down24h 87.30%
- p50
- 2.66s
- p99
- 29.6s
- Throughput
- 59 tps
- Context
- 256K
- $/Mtok
- $0.625 / $3.13
venice/fp8fp8
Measured 4m ago. Latency and throughput are OpenRouter's rolling 30-minute windows.
History
Availability
Axis starts at 90% so small differences stay visible; it expands to 0 when an endpoint falls below.
Time to first token
Logarithmic axis — the fleet spans two orders of magnitude. p50 and p99 are OpenRouter's rolling 30-minute windows.
Throughput
Median tokens per second.