Nemotron 3 Super
nvidia/nemotron-3-super-120b-a12b
Compare with Nemotron 3 Ultra, Nemotron 3 Nano 30B A3B, Nemotron 3.5 Lightning
CompareNemotron 3 Super
| Provider | State | Uptime 5m | 24h | p50 latency | p99 | Throughput | Context | $/Mtok |
|---|---|---|---|---|---|---|---|---|
DigitalOceandigitalocean | Healthy | 100% | 98.93% | 1.54s | 16.6s | 8.0 tps | 1.0M | $0.165 / $0.358 |
DeepInfradeepinfra/bf16bf16 | Healthy | 98.99% | 92.52% | 1.14s | 71.6s | 38 tps | 262K | $0.085 / $0.400 |
Nebiusnebius/fp4fp4 | No data | — | 97.52% | — | — | — | 8K | $0.300 / $0.900 |
- DigitalOcean100%Healthy24h 98.93%
- p50
- 1.54s
- p99
- 16.6s
- Throughput
- 8.0 tps
- Context
- 1.0M
- $/Mtok
- $0.165 / $0.358
digitalocean - DeepInfra98.99%Healthy24h 92.52%
- p50
- 1.14s
- p99
- 71.6s
- Throughput
- 38 tps
- Context
- 262K
- $/Mtok
- $0.085 / $0.400
deepinfra/bf16bf16 - Nebius—No data24h 97.52%
- p50
- —
- p99
- —
- Throughput
- —
- Context
- 8K
- $/Mtok
- $0.300 / $0.900
nebius/fp4fp4
Measured 5m ago. Latency and throughput are OpenRouter's rolling 30-minute windows.
History
1 endpoint reporting no data is omitted from the charts.
Availability
Axis starts at 90% so small differences stay visible; it expands to 0 when an endpoint falls below.
Time to first token
Logarithmic axis — the fleet spans two orders of magnitude. p50 and p99 are OpenRouter's rolling 30-minute windows, averaged over the hour.
Throughput
Median tokens per second.