modelstatus.dev

← DeepSeek V4 Flash 0731

DeepSeek V4 Flash 0731: providers compared

Every endpoint serving this model, side by side. ◆ marks the best value in each column — per model, never fleet-wide.

  • AtlasCloud100%◆ (best)
    Healthy24h 99.89%
    7d
    99.89%
    30d
    99.65%
    p50
    1.37s
    p99
    2.95s
    Throughput
    78 tps
    Context
    1.0M
    $/Mtok
    $0.440 / $1.32
    atlas-cloud/fp4fp4
  • Cloudflare100%◆ (best)
    Healthy24h 99.97%
    7d
    99.95%
    30d
    99.69%
    p50
    1.37s
    p99
    13.1s
    Throughput
    54 tps
    Context
    1.0M
    $/Mtok
    $0.440 / $1.32
    cloudflare
  • GMICloud100%◆ (best)
    Healthy24h 99.99%
    7d
    99.98%
    30d
    96.43%
    p50
    2.04s
    p99
    5.92s
    Throughput
    68 tps
    Context
    1.0M
    $/Mtok
    $0.286 / $0.858
    gmicloud/fp8fp8
  • Morph100%◆ (best)
    Degraded24h 93.20%
    7d
    95.72%
    30d
    98.07%
    p50
    1.55s
    p99
    40.1s
    Throughput
    21 tps
    Context
    1.0M
    $/Mtok
    $0.142 / $0.400
    morph/bf16bf16
  • NextBit100%◆ (best)
    Healthy24h 99.82%
    7d
    99.96%
    30d
    99.54%
    p50
    2.44s
    p99
    11.3s
    Throughput
    67 tps
    Context
    1.0M
    $/Mtok
    $0.352 / $1.06
    nextbit/fp8fp8
  • Novita100%◆ (best)
    Healthy24h 100%
    7d
    99.99%◆ (best)
    30d
    99.99%
    p50
    1.33s
    p99
    7.75s
    Throughput
    50 tps
    Context
    1.0M
    $/Mtok
    $0.409 / $1.23
    novita/fp8fp8
  • OpenInference100%◆ (best)
    Healthy24h 99.64%
    7d
    98.77%
    30d
    98.77%
    p50
    996ms
    p99
    10.4s
    Throughput
    35 tps
    Context
    1.0M
    $/Mtok
    $0.020 / $1.66
    open-inference/fp4fp4
  • Reka100%◆ (best)
    Healthy24h 99.87%
    7d
    99.71%
    30d
    99.73%
    p50
    734ms
    p99
    14.5s
    Throughput
    75 tps
    Context
    262K
    $/Mtok
    $0.021 / $0.528
    reka
  • SiliconFlow100%◆ (best)
    Healthy24h 99.46%
    7d
    99.52%
    30d
    98.72%
    p50
    1.73s
    p99
    8.55s
    Throughput
    54 tps
    Context
    1.0M
    $/Mtok
    $0.220 / $0.660
    siliconflow/fp8fp8
  • Wafer100%◆ (best)
    Healthy24h 99.96%
    7d
    99.99%
    30d
    99.46%
    p50
    539ms
    p99
    1.60s◆ (best)
    Throughput
    80 tps
    Context
    1.0M
    $/Mtok
    $0.200 / $0.840
    wafer/fast
  • CoreWeave99.99%
    Healthy24h 99.99%
    7d
    99.98%
    30d
    99.91%
    p50
    505ms
    p99
    7.91s
    Throughput
    84 tps
    Context
    262K
    $/Mtok
    $0.130 / $0.280
    coreweave/fp8fp8
  • Baidu99.99%
    Healthy24h 99.93%
    7d
    99.22%
    30d
    99.47%
    p50
    865ms
    p99
    5.82s
    Throughput
    108 tps◆ (best)
    Context
    1.0M
    $/Mtok
    $0.440 / $1.32
    baidu/fp8fp8
  • DeepInfra99.98%
    Healthy24h 99.94%
    7d
    99.91%
    30d
    99.75%
    p50
    909ms
    p99
    9.84s
    Throughput
    26 tps
    Context
    1.0M
    $/Mtok
    $0.060 / $0.180
    deepinfra/fp8fp8
  • Inceptron99.93%
    Healthy24h 99.92%
    7d
    99.93%
    30d
    98.97%
    p50
    687ms
    p99
    17.2s
    Throughput
    46 tps
    Context
    1.0M
    $/Mtok
    $0.050 / $0.650
    inceptron/fp4fp4
  • Healthy24h 99.94%
    7d
    99.86%
    30d
    99.65%
    p50
    677ms
    p99
    11.8s
    Throughput
    44 tps
    Context
    1.0M
    $/Mtok
    $0.119 / $0.238
    digitalocean
  • Healthy24h 99.69%
    7d
    99.65%
    30d
    99.12%
    p50
    1.33s
    p99
    13.6s
    Throughput
    39 tps
    Context
    1.0M
    $/Mtok
    $0.044 / $0.132
    streamlake/fp8fp8
  • BaseTen99.72%
    Healthy24h 97.41%
    7d
    99.66%
    30d
    98.37%
    p50
    850ms
    p99
    9.60s
    Throughput
    93 tps
    Context
    1.0M
    $/Mtok
    $0.130 / $0.260
    baseten/fp8fp8
  • Together99.66%
    Healthy24h 99.69%
    7d
    99.70%
    30d
    99.51%
    p50
    944ms
    p99
    31.8s
    Throughput
    39 tps
    Context
    1.0M
    $/Mtok
    $0.140 / $0.280
    together
  • Cohere99.41%
    Healthy24h 99.67%
    7d
    99.69%
    30d
    99.63%
    p50
    471ms◆ (best)
    p99
    12.9s
    Throughput
    91 tps
    Context
    1.0M
    $/Mtok
    $0.140 / $0.280
    cohere
  • Relace99.37%
    Healthy24h 99.90%
    7d
    99.30%
    30d
    99.31%
    p50
    939ms
    p99
    12.0s
    Throughput
    28 tps
    Context
    1.0M
    $/Mtok
    $0.015 / $1.28
    relace/fp4fp4
  • Alibaba99.26%
    Healthy24h 99.05%
    7d
    99.14%
    30d
    99.64%
    p50
    1.41s
    p99
    6.71s
    Throughput
    83 tps
    Context
    1.0M
    $/Mtok
    $0.176 / $0.528
    alibaba
  • Venice99.08%
    Healthy24h 99.55%
    7d
    99.63%
    30d
    99.26%
    p50
    1.30s
    p99
    17.8s
    Throughput
    33 tps
    Context
    1.0M
    $/Mtok
    $0.175 / $0.350
    venice
  • Healthy24h 96.69%
    7d
    97.63%
    30d
    96.49%
    p50
    3.38s
    p99
    19.3s
    Throughput
    25 tps
    Context
    1.0M
    $/Mtok
    $0.019 / $0.420
    sail-research/usfp4
  • Parasail98.71%
    Healthy24h 99.58%
    7d
    99.53%
    30d
    99.63%
    p50
    705ms
    p99
    10.3s
    Throughput
    90 tps
    Context
    1.0M
    $/Mtok
    $0.140 / $0.280
    parasail/fp8fp8
  • Healthy24h 97.92%
    7d
    98.39%
    30d
    98.88%
    p50
    3.11s
    p99
    21.7s
    Throughput
    25 tps
    Context
    1.0M
    $/Mtok
    $0.019 / $0.300
    sail-research/fp4fp4
  • Mancer 297.75%
    Healthy24h 99.31%
    7d
    99.24%
    30d
    97.85%
    p50
    1.33s
    p99
    10.9s
    Throughput
    37 tps
    Context
    1.0M
    $/Mtok
    $0.200 / $0.600
    mancer/fp8fp8
  • No data24h —
    7d
    —
    30d
    —
    p50
    —
    p99
    —
    Throughput
    —
    Context
    1.0M
    $/Mtok
    $0.065 / $0.180
    akashml/fp8fp8
  • No data24h —
    7d
    —
    30d
    —
    p50
    —
    p99
    —
    Throughput
    —
    Context
    1.0M
    $/Mtok
    $0.080 / $0.180
    ambient/fp4fp4
  • No data24h —
    7d
    —
    30d
    —
    p50
    —
    p99
    —
    Throughput
    —
    Context
    262K
    $/Mtok
    $0.064 / $0.128
    decart/fp4fp4
  • No data24h —
    7d
    —
    30d
    99.99%◆ (best)
    p50
    —
    p99
    —
    Throughput
    —
    Context
    1.0M
    $/Mtok
    $0.440 / $1.32
    deepseek
  • No data24h —
    7d
    —
    30d
    88.20%
    p50
    —
    p99
    —
    Throughput
    —
    Context
    1.0M
    $/Mtok
    $0.220 / $0.660
    fireworks
  • No data24h —
    7d
    —
    30d
    99.27%
    p50
    —
    p99
    —
    Throughput
    —
    Context
    1.0M
    $/Mtok
    $0.286 / $0.858
    gmicloud/fp4fp4
  • No data24h —
    7d
    —
    30d
    —
    p50
    —
    p99
    —
    Throughput
    —
    Context
    262K
    $/Mtok
    $0.219 / $0.490
    io-net/fp8fp8
  • No data24h —
    7d
    98.81%
    30d
    98.27%
    p50
    —
    p99
    —
    Throughput
    —
    Context
    1.0M
    $/Mtok
    $0.090 / $0.195
    makora
  • No data24h —
    7d
    —
    30d
    97.50%
    p50
    —
    p99
    —
    Throughput
    —
    Context
    1.0M
    $/Mtok
    $0.142 / $0.400
    morph
  • No data24h —
    7d
    89.48%
    30d
    87.88%
    p50
    —
    p99
    —
    Throughput
    —
    Context
    1.0M
    $/Mtok
    $0.140 / $0.280
    nebius/fp8fp8
  • No data24h —
    7d
    95.77%
    30d
    98.38%
    p50
    —
    p99
    —
    Throughput
    —
    Context
    1.0M
    $/Mtok
    $0.004 / $1.04◆ (best)
    open-inference/fp8fp8
  • No data24h 99.86%
    7d
    99.93%
    30d
    98.49%
    p50
    2.19s
    p99
    11.2s
    Throughput
    41 tps
    Context
    1.0M
    $/Mtok
    $0.308 / $0.924
    phala
  • Reka—
    No data24h —
    7d
    —
    30d
    99.44%
    p50
    —
    p99
    —
    Throughput
    —
    Context
    262K
    $/Mtok
    $0.088 / $0.528
    reka/fp4fp4

Measured 4m ago. Latency and throughput are OpenRouter's rolling 30-minute windows. ◆ marks the best value in each column among these endpoints.

Price vs latency

One dot per endpoint at the latest round · further left = faster, lower = cheaper · hover a dot for its tag

History

Availability

Axis starts at 90% so small differences stay visible; it expands to 0 when an endpoint falls below.

Time to first token

Logarithmic axis — the fleet spans two orders of magnitude. P50 is OpenRouter's rolling 30-minute window.

Throughput

Median tokens per second.