modelstatus.dev

DeepSeek V4 Pro 0423

deepseek/deepseek-v4-pro

Compare with DeepSeek V4 Pro 0813, DeepSeek V4 Flash 0731, DeepSeek V4 Flash 0423

CompareDeepSeek V4 Pro 0423
  • BaseTen100%
    Healthy24h 99.76%
    p50
    419ms
    p99
    3.17s
    Throughput
    59 tps
    Context
    1.0M
    $/Mtok
    $1.74 / $3.48
    baseten/fp4fp4
  • DeepInfra100%
    Healthy24h 96.50%
    p50
    865ms
    p99
    29.9s
    Throughput
    29 tps
    Context
    1.0M
    $/Mtok
    $1.30 / $2.60
    deepinfra/fp8fp8
  • DeepSeek100%
    Healthy24h 74.80%
    p50
    1.10s
    p99
    2.84s
    Throughput
    27 tps
    Context
    1.0M
    $/Mtok
    $0.660 / $1.98
    deepseek
  • Novita99.58%
    Healthy24h 99.15%
    p50
    1.59s
    p99
    78.5s
    Throughput
    58 tps
    Context
    1.0M
    $/Mtok
    $1.44 / $2.88
    novita/fp8fp8
  • Together98.62%
    Healthy24h 91.49%
    p50
    646ms
    p99
    13.8s
    Throughput
    51 tps
    Context
    512K
    $/Mtok
    $1.74 / $3.48
    together
  • SiliconFlow96.48%
    Healthy24h 96.39%
    p50
    2.19s
    p99
    47.9s
    Throughput
    47 tps
    Context
    1.0M
    $/Mtok
    $1.50 / $3.14
    siliconflow/fp8fp8
  • StreamLake94.93%
    Degraded24h 96.29%
    p50
    2.38s
    p99
    16.3s
    Throughput
    36 tps
    Context
    1.0M
    $/Mtok
    $0.657 / $1.31
    streamlake/fp8fp8
  • Baidu90.88%
    Degraded24h 97.12%
    p50
    774ms
    p99
    7.27s
    Throughput
    62 tps
    Context
    1.0M
    $/Mtok
    $0.659 / $1.32
    baidu/fp8fp8
  • Alibaba85.86%
    Down24h 97.40%
    p50
    1.32s
    p99
    10.6s
    Throughput
    62 tps
    Context
    1.0M
    $/Mtok
    $1.42 / $2.83
    alibaba/fp8fp8
  • Ionstream85.48%
    Down24h 95.06%
    p50
    3.44s
    p99
    13.4s
    Throughput
    15 tps
    Context
    1.0M
    $/Mtok
    $1.13 / $2.26
    ionstream/fp4fp4
  • Azure85.37%
    Down24h 93.30%
    p50
    1.75s
    p99
    16.8s
    Throughput
    48 tps
    Context
    1.0M
    $/Mtok
    $1.91 / $3.83
    azure/us
  • GMICloud83.50%
    Down24h 95.49%
    p50
    2.81s
    p99
    10.8s
    Throughput
    34 tps
    Context
    1.0M
    $/Mtok
    $0.696 / $1.39
    gmicloud/fp8fp8
  • Parasail81.22%
    Down24h 89.59%
    p50
    866ms
    p99
    12.1s
    Throughput
    32 tps
    Context
    1.0M
    $/Mtok
    $1.74 / $3.48
    parasail/fp8fp8
  • AtlasCloud75.76%
    Down24h 92.11%
    p50
    1.37s
    p99
    8.53s
    Throughput
    38 tps
    Context
    1.0M
    $/Mtok
    $1.68 / $3.38
    atlas-cloud/fp4fp4
  • CoreWeave
    No data24h 96.82%
    p50
    474ms
    p99
    3.80s
    Throughput
    63 tps
    Context
    1.0M
    $/Mtok
    $1.15 / $2.55
    coreweave/fp8fp8
  • DigitalOcean
    Down24h 93.72%
    p50
    p99
    Throughput
    Context
    1.0M
    $/Mtok
    $0.870 / $1.74
    digitalocean
  • Fireworks
    No data24h 0.00%
    p50
    2.02s
    p99
    12.0s
    Throughput
    41 tps
    Context
    1.0M
    $/Mtok
    $1.74 / $3.48
    fireworks
  • Venice
    Down24h 75.38%
    p50
    2.43s
    p99
    55.8s
    Throughput
    29 tps
    Context
    1.0M
    $/Mtok
    $1.65 / $3.30
    venice

Measured 50s ago. Latency and throughput are OpenRouter's rolling 30-minute windows.

2 endpoints reporting no data are omitted from the charts.

Availability

Axis starts at 90% so small differences stay visible; it expands to 0 when an endpoint falls below.

Time to first token

Logarithmic axis — the fleet spans two orders of magnitude. p50 and p99 are OpenRouter's rolling 30-minute windows, averaged over the hour.

Throughput

Median tokens per second.