modelstatus.dev

Nemotron 3 Ultra

nvidia/nemotron-3-ultra-550b-a55b

Compare with Nemotron 3 Ultra (batch), Nemotron 3 Ultra (free), Nemotron 3 Nano 30B A3B

CompareNemotron 3 Ultra
  • BaseTen100%
    Healthy24h 98.85%
    p50
    528ms
    p99
    5.28s
    Throughput
    55 tps
    Context
    203K
    $/Mtok
    $0.600 / $2.40
    baseten/fp4fp4
  • Together98.95%
    Healthy24h 98.58%
    p50
    673ms
    p99
    8.05s
    Throughput
    75 tps
    Context
    512K
    $/Mtok
    $0.600 / $3.60
    together
  • DeepInfra95.17%
    Degraded24h 92.16%
    p50
    2.22s
    p99
    16.8s
    Throughput
    68 tps
    Context
    262K
    $/Mtok
    $0.500 / $2.20
    deepinfra/fp4fp4
  • Venice95.12%
    Degraded24h 87.49%
    p50
    1.22s
    p99
    9.62s
    Throughput
    28 tps
    Context
    256K
    $/Mtok
    $0.625 / $3.13
    venice/fp8fp8

Measured 2m ago. Latency and throughput are OpenRouter's rolling 30-minute windows.

Availability

Axis starts at 90% so small differences stay visible; it expands to 0 when an endpoint falls below.

Time to first token

Logarithmic axis — the fleet spans two orders of magnitude. p50 and p99 are OpenRouter's rolling 30-minute windows.

Throughput

Median tokens per second.