modelstatus.dev

DeepSeek V4 Pro 0423

DeepSeek V4 Pro 0423: providers compared

Every endpoint serving this model, side by side. ◆ marks the best value in each column — per model, never fleet-wide.

  • Alibaba100% (best)
    Healthy24h 97.45%
    7d
    98.37%
    30d
    98.37%
    p50
    1.45s
    p99
    6.13s
    Throughput
    51 tps
    Context
    1.0M
    $/Mtok
    $1.42 / $2.83
    alibaba/fp8fp8
  • AtlasCloud100% (best)
    Healthy24h 94.75%
    7d
    96.76%
    30d
    96.76%
    p50
    1.81s
    p99
    15.7s
    Throughput
    38 tps
    Context
    1.0M
    $/Mtok
    $1.68 / $3.38
    atlas-cloud/fp4fp4
  • Fireworks100% (best)
    Healthy24h 95.24%
    7d
    96.49%
    30d
    96.49%
    p50
    1.43s
    p99
    11.7s
    Throughput
    59 tps
    Context
    1.0M
    $/Mtok
    $1.74 / $3.48
    fireworks
  • Parasail100% (best)
    Healthy24h 92.90%
    7d
    92.72%
    30d
    92.72%
    p50
    419ms (best)
    p99
    3.38s
    Throughput
    55 tps
    Context
    1.0M
    $/Mtok
    $1.74 / $3.48
    parasail/fp8fp8
  • Venice100% (best)
    Healthy24h 81.72%
    7d
    84.13%
    30d
    84.13%
    p50
    1.21s
    p99
    9.50s
    Throughput
    51 tps
    Context
    1.0M
    $/Mtok
    $1.65 / $3.30
    venice
  • Baidu99.98%
    Healthy24h 99.03%
    7d
    98.03%
    30d
    98.03%
    p50
    1.10s
    p99
    7.76s
    Throughput
    48 tps
    Context
    1.0M
    $/Mtok
    $0.642 / $1.28 (best)
    baidu/fp8fp8
  • Healthy24h 98.92%
    7d
    98.36%
    30d
    98.36%
    p50
    2.23s
    p99
    17.4s
    Throughput
    32 tps
    Context
    1.0M
    $/Mtok
    $0.650 / $1.30
    streamlake/fp8fp8
  • Ionstream99.83%
    Healthy24h 95.59%
    7d
    96.14%
    30d
    96.14%
    p50
    2.63s
    p99
    15.4s
    Throughput
    25 tps
    Context
    1.0M
    $/Mtok
    $1.13 / $2.26
    ionstream/fp4fp4
  • Healthy24h 95.35%
    7d
    63.07%
    30d
    63.07%
    p50
    2.15s
    p99
    13.8s
    Throughput
    8.0 tps
    Context
    1.0M
    $/Mtok
    $0.870 / $1.74
    digitalocean
  • DeepInfra99.71%
    Healthy24h 99.52%
    7d
    98.71%
    30d
    98.71%
    p50
    805ms
    p99
    8.75s
    Throughput
    38 tps
    Context
    1.0M
    $/Mtok
    $1.30 / $2.60
    deepinfra/fp8fp8
  • BaseTen99.69%
    Healthy24h 99.86%
    7d
    99.88% (best)
    30d
    99.88% (best)
    p50
    429ms
    p99
    4.09s
    Throughput
    62 tps (best)
    Context
    1.0M
    $/Mtok
    $1.74 / $3.48
    baseten/fp4fp4
  • Together99.67%
    Healthy24h 96.20%
    7d
    97.59%
    30d
    97.59%
    p50
    548ms
    p99
    9.48s
    Throughput
    58 tps
    Context
    512K
    $/Mtok
    $1.74 / $3.48
    together
  • Azure99.31%
    Healthy24h 94.65%
    7d
    94.61%
    30d
    94.61%
    p50
    1.56s
    p99
    18.0s
    Throughput
    47 tps
    Context
    1.0M
    $/Mtok
    $1.91 / $3.83
    azure/us
  • Healthy24h 97.96%
    7d
    97.44%
    30d
    97.44%
    p50
    2.44s
    p99
    18.7s
    Throughput
    40 tps
    Context
    1.0M
    $/Mtok
    $1.50 / $3.14
    siliconflow/fp8fp8
  • GMICloud98.74%
    Healthy24h 96.43%
    7d
    97.09%
    30d
    97.09%
    p50
    3.35s
    p99
    8.69s
    Throughput
    26 tps
    Context
    1.0M
    $/Mtok
    $0.792 / $2.38
    gmicloud/fp8fp8
  • Novita98.60%
    Degraded24h 98.69%
    7d
    99.16%
    30d
    99.16%
    p50
    1.79s
    p99
    17.5s
    Throughput
    46 tps
    Context
    1.0M
    $/Mtok
    $1.60 / $3.20
    novita/fp8fp8
  • No data24h 96.61%
    7d
    99.52%
    30d
    99.52%
    p50
    565ms
    p99
    3.30s (best)
    Throughput
    62 tps
    Context
    1.0M
    $/Mtok
    $1.15 / $2.55
    coreweave/fp8fp8
  • No data24h
    7d
    86.30%
    30d
    86.30%
    p50
    p99
    Throughput
    Context
    1.0M
    $/Mtok
    $0.660 / $1.98
    deepseek

Measured 4m ago. Latency and throughput are OpenRouter's rolling 30-minute windows. ◆ marks the best value in each column among these endpoints.

Price vs latency

One dot per endpoint at the latest round · further left = faster, lower = cheaper · hover a dot for its tag

History

Availability

Axis starts at 90% so small differences stay visible; it expands to 0 when an endpoint falls below.

Time to first token

Logarithmic axis — the fleet spans two orders of magnitude. P50 is OpenRouter's rolling 30-minute window.

Throughput

Median tokens per second.