modelstatus.dev

DeepSeek V4 Pro 0423

deepseek/deepseek-v4-pro

Compare with DeepSeek V4 Pro 0813, DeepSeek V4 Flash 0731, DeepSeek V4 Flash 0423

CompareDeepSeek V4 Pro 0423
  • BaseTen100%
    Healthy24h 99.76%
    p50
    418ms
    p99
    3.06s
    Throughput
    60 tps
    Context
    1.0M
    $/Mtok
    $1.74 / $3.48
    baseten/fp4fp4
  • DeepInfra100%
    Healthy24h 96.53%
    p50
    866ms
    p99
    28.7s
    Throughput
    29 tps
    Context
    1.0M
    $/Mtok
    $1.30 / $2.60
    deepinfra/fp8fp8
  • DeepSeek99.96%
    Healthy24h 74.76%
    p50
    1.09s
    p99
    2.87s
    Throughput
    27 tps
    Context
    1.0M
    $/Mtok
    $0.660 / $1.98
    deepseek
  • Novita99.73%
    Healthy24h 99.15%
    p50
    1.59s
    p99
    93.1s
    Throughput
    57 tps
    Context
    1.0M
    $/Mtok
    $1.44 / $2.88
    novita/fp8fp8
  • Together97.71%
    Healthy24h 91.50%
    p50
    721ms
    p99
    14.0s
    Throughput
    51 tps
    Context
    512K
    $/Mtok
    $1.74 / $3.48
    together
  • AtlasCloud97.14%
    Down24h 92.08%
    p50
    1.38s
    p99
    9.72s
    Throughput
    37 tps
    Context
    1.0M
    $/Mtok
    $1.68 / $3.38
    atlas-cloud/fp4fp4
  • SiliconFlow96.90%
    Healthy24h 96.33%
    p50
    2.20s
    p99
    67.4s
    Throughput
    47 tps
    Context
    1.0M
    $/Mtok
    $1.50 / $3.14
    siliconflow/fp8fp8
  • StreamLake94.60%
    Degraded24h 96.26%
    p50
    2.38s
    p99
    15.8s
    Throughput
    36 tps
    Context
    1.0M
    $/Mtok
    $0.657 / $1.31
    streamlake/fp8fp8
  • Baidu91.14%
    Degraded24h 97.00%
    p50
    769ms
    p99
    7.58s
    Throughput
    60 tps
    Context
    1.0M
    $/Mtok
    $0.659 / $1.32
    baidu/fp8fp8
  • Alibaba85.52%
    Down24h 97.35%
    p50
    1.32s
    p99
    11.3s
    Throughput
    61 tps
    Context
    1.0M
    $/Mtok
    $1.42 / $2.83
    alibaba/fp8fp8
  • GMICloud84.38%
    Down24h 95.43%
    p50
    2.78s
    p99
    15.8s
    Throughput
    32 tps
    Context
    1.0M
    $/Mtok
    $0.696 / $1.39
    gmicloud/fp8fp8
  • Ionstream80.67%
    Down24h 94.95%
    p50
    3.43s
    p99
    13.5s
    Throughput
    15 tps
    Context
    1.0M
    $/Mtok
    $1.13 / $2.26
    ionstream/fp4fp4
  • Azure72.25%
    Down24h 93.17%
    p50
    1.83s
    p99
    15.4s
    Throughput
    46 tps
    Context
    1.0M
    $/Mtok
    $1.91 / $3.83
    azure/us
  • Parasail72.07%
    Down24h 89.45%
    p50
    866ms
    p99
    12.0s
    Throughput
    31 tps
    Context
    1.0M
    $/Mtok
    $1.74 / $3.48
    parasail/fp8fp8
  • CoreWeave
    No data24h 96.80%
    p50
    474ms
    p99
    3.80s
    Throughput
    63 tps
    Context
    1.0M
    $/Mtok
    $1.15 / $2.55
    coreweave/fp8fp8
  • DigitalOcean
    Down24h 93.69%
    p50
    p99
    Throughput
    Context
    1.0M
    $/Mtok
    $0.870 / $1.74
    digitalocean
  • Fireworks
    No data24h 0.00%
    p50
    2.03s
    p99
    9.99s
    Throughput
    41 tps
    Context
    1.0M
    $/Mtok
    $1.74 / $3.48
    fireworks
  • Venice
    Down24h 75.35%
    p50
    2.34s
    p99
    65.1s
    Throughput
    31 tps
    Context
    1.0M
    $/Mtok
    $1.65 / $3.30
    venice

Measured 32s ago. Latency and throughput are OpenRouter's rolling 30-minute windows.

2 endpoints reporting no data are omitted from the charts.

Availability

Axis starts at 90% so small differences stay visible; it expands to 0 when an endpoint falls below.

Time to first token

Logarithmic axis — the fleet spans two orders of magnitude. p50 and p99 are OpenRouter's rolling 30-minute windows.

Throughput

Median tokens per second.