Nemotron 3 Ultra
nvidia/nemotron-3-ultra-550b-a55b
Compare with Nemotron 3 Ultra (batch), Nemotron 3 Ultra (free), Nemotron 3 Nano 30B A3B
Showing only the venice/fp8 provider · show all 4
CompareNemotron 3 Ultra
| State | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Healthy | 100% | 75.58% | 82.50% | 82.50% | 1.29s | 38.2s | 56 tps | 256K | $0.625 / $3.13 |
- Venice100%Healthy24h 75.58%
- 7d
- 82.50%
- 30d
- 82.50%
- p50
- 1.29s
- p99
- 38.2s
- Throughput
- 56 tps
- Context
- 256K
- $/Mtok
- $0.625 / $3.13
venice/fp8fp8
Measured 2m ago. Latency and throughput are OpenRouter's rolling 30-minute windows.
History
Availability
Axis starts at 90% so small differences stay visible; it expands to 0 when an endpoint falls below.
Time to first token
Logarithmic axis — the fleet spans two orders of magnitude. P50 is OpenRouter's rolling 30-minute window, averaged over the hour.
Throughput
Median tokens per second.