modelstatus.dev

DeepSeek V3.1

deepseek/deepseek-chat-v3.1

Compare with DeepSeek V3.1 Terminus, DeepSeek V3, DeepSeek V3.2

CompareDeepSeek V3.1
  • AtlasCloud100%
    Healthy24h 97.53%
    p50
    1.71s
    p99
    4.35s
    Throughput
    21 tps
    Context
    131K
    $/Mtok
    $0.300 / $0.950
    atlas-cloud/fp8fp8
  • CoreWeave100%
    Healthy24h 99.96%
    p50
    371ms
    p99
    1.17s
    Throughput
    43 tps
    Context
    161K
    $/Mtok
    $0.550 / $1.65
    coreweave/fp8fp8
  • DeepInfra100%
    Healthy24h 99.19%
    p50
    906ms
    p99
    7.07s
    Throughput
    8.0 tps
    Context
    164K
    $/Mtok
    $0.250 / $0.950
    deepinfra/fp4fp4
  • Novita100%
    Healthy24h 99.98%
    p50
    1.75s
    p99
    8.87s
    Throughput
    18 tps
    Context
    131K
    $/Mtok
    $0.270 / $1.00
    novita/fp8fp8
  • SiliconFlow100%
    Healthy24h 97.46%
    p50
    1.34s
    p99
    9.27s
    Throughput
    14 tps
    Context
    164K
    $/Mtok
    $0.270 / $1.00
    siliconflow/fp8fp8
  • Google93.01%
    Degraded24h 98.40%
    p50
    1.18s
    p99
    3.07s
    Throughput
    61 tps
    Context
    164K
    $/Mtok
    $0.600 / $1.70
    google-vertex/us-west2
  • Mara61.18%
    Down24h 86.37%
    p50
    3.22s
    p99
    16.9s
    Throughput
    23 tps
    Context
    131K
    $/Mtok
    $0.600 / $1.70
    mara
  • SambaNova
    Degraded24h 92.07%
    p50
    1.89s
    p99
    11.7s
    Throughput
    28 tps
    Context
    131K
    $/Mtok
    $0.650 / $1.50
    sambanova/fp8fp8

Measured 3m ago. Latency and throughput are OpenRouter's rolling 30-minute windows.

Availability

Axis starts at 90% so small differences stay visible; it expands to 0 when an endpoint falls below.

Time to first token

Logarithmic axis — the fleet spans two orders of magnitude. p50 and p99 are OpenRouter's rolling 30-minute windows.

Throughput

Median tokens per second.