modelstatus.dev

Nemotron 3 Ultra

nvidia/nemotron-3-ultra-550b-a55b

Compare with Nemotron 3 Ultra (batch), Nemotron 3 Ultra (free), Nemotron 3 Nano 30B A3B

CompareNemotron 3 Ultra
  • BaseTen100%
    Healthy24h 98.86%
    p50
    533ms
    p99
    5.30s
    Throughput
    55 tps
    Context
    203K
    $/Mtok
    $0.600 / $2.40
    baseten/fp4fp4
  • Together99.00%
    Healthy24h 98.57%
    p50
    690ms
    p99
    8.23s
    Throughput
    79 tps
    Context
    512K
    $/Mtok
    $0.600 / $3.60
    together
  • DeepInfra94.77%
    Degraded24h 92.13%
    p50
    2.45s
    p99
    21.9s
    Throughput
    63 tps
    Context
    262K
    $/Mtok
    $0.500 / $2.20
    deepinfra/fp4fp4
  • Venice90.00%
    Degraded24h 87.47%
    p50
    1.37s
    p99
    11.3s
    Throughput
    57 tps
    Context
    256K
    $/Mtok
    $0.625 / $3.13
    venice/fp8fp8

Measured 4m ago. Latency and throughput are OpenRouter's rolling 30-minute windows.

Availability

Axis starts at 90% so small differences stay visible; it expands to 0 when an endpoint falls below.

Time to first token

Logarithmic axis — the fleet spans two orders of magnitude. p50 and p99 are OpenRouter's rolling 30-minute windows, averaged over the hour.

Throughput

Median tokens per second.