modelstatus.dev

Nemotron 3.5 Lightning

nvidia/nemotron-3.5-lightningInputtextOutputtext

Compare its 6 providersIncident RSSCompare with Nemotron 3.5 Content Safety, Nemotron 3 Nano 30B A3B, Nemotron 3 Ultra

CompareNemotron 3.5 Lightning
  • Healthy24h 100%
    7d
    99.99%
    30d
    99.95%
    p50
    182ms
    p99
    433ms
    Throughput
    282 tps
    Context
    262K
    $/Mtok
    $0.070 / $0.200
    coreweave/bf16bf16
  • Healthy24h 99.99%
    7d
    99.92%
    30d
    99.83%
    p50
    440ms
    p99
    2.47s
    Throughput
    245 tps
    Context
    262K
    $/Mtok
    $0.060 / $0.160
    deepinfra/bf16bf16
  • No data24h 99.68%
    7d
    99.87%
    30d
    99.54%
    p50
    477ms
    p99
    4.08s
    Throughput
    90 tps
    Context
    262K
    $/Mtok
    $0.039 / $0.180
    darkbloom/int4int4
  • No data24h 99.55%
    7d
    97.54%
    30d
    97.94%
    p50
    294ms
    p99
    3.84s
    Throughput
    208 tps
    Context
    262K
    $/Mtok
    $0.059 / $0.170
    io-net
  • No data24h 99.64%
    7d
    98.67%
    30d
    99.07%
    p50
    312ms
    p99
    1.40s
    Throughput
    287 tps
    Context
    262K
    $/Mtok
    $0.070 / $0.200
    phala
  • No data24h —
    7d
    —
    30d
    —
    p50
    —
    p99
    —
    Throughput
    —
    Context
    1.0M
    $/Mtok
    $0.100 / $0.250
    venice/fp4fp4

Measured 2m ago. Latency and throughput are OpenRouter's rolling 30-minute windows.

1 endpoint has no chartable data in this range.

Availability

Axis starts at 90% so small differences stay visible; it expands to 0 when an endpoint falls below.

Time to first token

Logarithmic axis — the fleet spans two orders of magnitude. P50 is OpenRouter's rolling 30-minute window.

Throughput

Median tokens per second.

Daily availability

One row per endpoint · one dot per day · last 90 days.

worst hour ≥ 95%90–95%< 90%no data

Colour by each endpoint's worst hour of each UTC day · 6 endpoints · 47 days with data, 22 endpoint-days below 90%.

Price changes