modelstatus.dev

DeepSeek V3.1

deepseek/deepseek-chat-v3.1

Compare with DeepSeek V4 Flash 0731, DeepSeek V4 Flash 0423, DeepSeek V4 Pro 0423

  • CoreWeave100%
    Healthy24h 99.96%
    p50
    363ms
    p99
    1.02s
    Throughput
    45 tps
    Context
    161K
    $/Mtok
    $0.550 / $1.65
    coreweave/fp8fp8
  • Google100%
    Healthy24h 98.34%
    p50
    1.15s
    p99
    3.74s
    Throughput
    87 tps
    Context
    164K
    $/Mtok
    $0.600 / $1.70
    google-vertex/us-west2
  • Novita100%
    Healthy24h 99.98%
    p50
    1.73s
    p99
    7.40s
    Throughput
    23 tps
    Context
    131K
    $/Mtok
    $0.270 / $1.00
    novita/fp8fp8
  • DeepInfra99.87%
    Healthy24h 99.21%
    p50
    905ms
    p99
    10.2s
    Throughput
    9.0 tps
    Context
    164K
    $/Mtok
    $0.250 / $0.950
    deepinfra/fp4fp4
  • AtlasCloud99.81%
    Healthy24h 97.41%
    p50
    1.52s
    p99
    3.79s
    Throughput
    22 tps
    Context
    131K
    $/Mtok
    $0.300 / $0.950
    atlas-cloud/fp8fp8
  • SiliconFlow91.43%
    Degraded24h 97.48%
    p50
    1.34s
    p99
    13.0s
    Throughput
    17 tps
    Context
    164K
    $/Mtok
    $0.270 / $1.00
    siliconflow/fp8fp8
  • Mara
    Down24h 88.59%
    p50
    1.84s
    p99
    19.8s
    Throughput
    122 tps
    Context
    131K
    $/Mtok
    $0.600 / $1.70
    mara
  • SambaNova
    Degraded24h 93.51%
    p50
    1.63s
    p99
    11.1s
    Throughput
    33 tps
    Context
    131K
    $/Mtok
    $0.650 / $1.50
    sambanova/fp8fp8

Measured 1m ago. Latency and throughput are OpenRouter's rolling 30-minute windows.

Availability

Axis starts at 90% so small differences stay visible; it expands to 0 when an endpoint falls below.

Time to first token

Logarithmic axis — the fleet spans two orders of magnitude. p50 and p99 are OpenRouter's rolling 30-minute windows.

Throughput

Median tokens per second.