modelstatus.dev

DeepSeek V3.2

DeepSeek V3.2: providers compared

Every endpoint serving this model, side by side. ◆ marks the best value in each column — per model, never fleet-wide.

  • Baidu100% (best)
    Healthy24h 99.91%
    7d
    99.94%
    30d
    99.94%
    p50
    999ms
    p99
    3.58s (best)
    Throughput
    50 tps (best)
    Context
    131K
    $/Mtok
    $0.280 / $0.420
    baidu/fp8fp8
  • Friendli100% (best)
    Healthy24h 99.88%
    7d
    92.07%
    30d
    92.07%
    p50
    975ms
    p99
    9.09s
    Throughput
    30 tps
    Context
    164K
    $/Mtok
    $0.500 / $1.50
    friendli
  • Google100% (best)
    Healthy24h 99.42%
    7d
    99.70%
    30d
    99.70%
    p50
    1.32s
    p99
    12.0s
    Throughput
    19 tps
    Context
    164K
    $/Mtok
    $0.560 / $1.68
    google-vertex
  • Novita100% (best)
    Healthy24h 99.98%
    7d
    99.99% (best)
    30d
    99.99% (best)
    p50
    1.35s
    p99
    6.20s
    Throughput
    26 tps
    Context
    164K
    $/Mtok
    $0.269 / $0.400
    novita/fp8fp8
  • Healthy24h 99.96%
    7d
    99.96%
    30d
    99.96%
    p50
    1.58s
    p99
    9.16s
    Throughput
    22 tps
    Context
    164K
    $/Mtok
    $0.260 / $0.380
    atlas-cloud/fp8fp8
  • GMICloud99.94%
    Healthy24h 99.55%
    7d
    99.57%
    30d
    99.57%
    p50
    2.04s
    p99
    6.71s
    Throughput
    5.0 tps
    Context
    164K
    $/Mtok
    $0.209 / $0.310 (best)
    gmicloud/fp8fp8
  • Healthy24h 96.46%
    7d
    97.21%
    30d
    97.21%
    p50
    1.52s
    p99
    14.4s
    Throughput
    12 tps
    Context
    164K
    $/Mtok
    $0.250 / $0.800
    digitalocean
  • Alibaba98.91%
    Healthy24h 97.96%
    7d
    98.40%
    30d
    98.40%
    p50
    844ms (best)
    p99
    12.0s
    Throughput
    48 tps
    Context
    131K
    $/Mtok
    $0.370 / $1.11
    alibaba/fp8fp8
  • Degraded24h 97.02%
    7d
    99.35%
    30d
    99.35%
    p50
    1.39s
    p99
    9.96s
    Throughput
    23 tps
    Context
    128K
    $/Mtok
    $0.214 / $0.322
    streamlake/fp8fp8
  • Phala98.32%
    Healthy24h 98.20%
    7d
    99.06%
    30d
    99.06%
    p50
    2.64s
    p99
    247s
    Throughput
    5.0 tps
    Context
    164K
    $/Mtok
    $1.00 / $1.00
    phala
  • Degraded24h 94.89%
    7d
    99.11%
    30d
    99.11%
    p50
    1.81s
    p99
    11.7s
    Throughput
    24 tps
    Context
    164K
    $/Mtok
    $0.259 / $0.420
    siliconflow/fp8fp8
  • Venice16.33%
    Down24h 82.72%
    7d
    75.07%
    30d
    75.07%
    p50
    3.90s
    p99
    75.1s
    Throughput
    3.0 tps
    Context
    160K
    $/Mtok
    $0.330 / $0.480
    venice
  • No data24h 97.32%
    7d
    93.77%
    30d
    93.77%
    p50
    999ms
    p99
    9.90s
    Throughput
    12 tps
    Context
    164K
    $/Mtok
    $0.260 / $0.380
    deepinfra/fp4fp4
  • No data24h 95.42%
    7d
    96.95%
    30d
    96.95%
    p50
    p99
    Throughput
    Context
    33K
    $/Mtok
    $3.00 / $4.50
    sambanova

Measured 2m ago. Latency and throughput are OpenRouter's rolling 30-minute windows. ◆ marks the best value in each column among these endpoints.

Price vs latency

One dot per endpoint at the latest round · further left = faster, lower = cheaper · hover a dot for its tag

History

Availability

Axis starts at 90% so small differences stay visible; it expands to 0 when an endpoint falls below.

Time to first token

Logarithmic axis — the fleet spans two orders of magnitude. P50 is OpenRouter's rolling 30-minute window, averaged over the hour.

Throughput

Median tokens per second.