modelstatus.dev

Nemotron 3 Ultra

nvidia/nemotron-3-ultra-550b-a55b

Compare with Nemotron 3 Ultra (batch), Nemotron 3 Ultra (free), Nemotron 3 Nano 30B A3B

CompareNemotron 3 Ultra
  • BaseTen100%
    Healthy24h 98.86%
    p50
    756ms
    p99
    10.2s
    Throughput
    60 tps
    Context
    203K
    $/Mtok
    $0.600 / $2.40
    baseten/fp4fp4
  • Together100%
    Healthy24h 96.99%
    p50
    830ms
    p99
    8.31s
    Throughput
    91 tps
    Context
    512K
    $/Mtok
    $0.600 / $3.60
    together
  • DeepInfra91.36%
    Degraded24h 92.04%
    p50
    2.65s
    p99
    45.0s
    Throughput
    54 tps
    Context
    262K
    $/Mtok
    $0.500 / $2.20
    deepinfra/fp4fp4
  • Venice
    Down24h 87.04%
    p50
    1.40s
    p99
    119s
    Throughput
    21 tps
    Context
    256K
    $/Mtok
    $0.625 / $3.13
    venice/fp8fp8

Measured 53s ago. Latency and throughput are OpenRouter's rolling 30-minute windows.

Availability

Axis starts at 90% so small differences stay visible; it expands to 0 when an endpoint falls below.

Time to first token

Logarithmic axis — the fleet spans two orders of magnitude. p50 and p99 are OpenRouter's rolling 30-minute windows, averaged over the hour.

Throughput

Median tokens per second.