modelstatus.dev

DeepSeek V3.1

deepseek/deepseek-chat-v3.1

Compare with DeepSeek V3.1 Terminus, DeepSeek V3, DeepSeek V3.2

CompareDeepSeek V3.1
  • AtlasCloud100%
    Healthy24h 97.53%
    p50
    1.72s
    p99
    4.54s
    Throughput
    22 tps
    Context
    131K
    $/Mtok
    $0.300 / $0.950
    atlas-cloud/fp8fp8
  • CoreWeave100%
    Healthy24h 99.96%
    p50
    371ms
    p99
    1.13s
    Throughput
    43 tps
    Context
    161K
    $/Mtok
    $0.550 / $1.65
    coreweave/fp8fp8
  • DeepInfra100%
    Healthy24h 99.19%
    p50
    951ms
    p99
    9.54s
    Throughput
    8.0 tps
    Context
    164K
    $/Mtok
    $0.250 / $0.950
    deepinfra/fp4fp4
  • Google100%
    Healthy24h 98.43%
    p50
    1.18s
    p99
    2.63s
    Throughput
    67 tps
    Context
    164K
    $/Mtok
    $0.600 / $1.70
    google-vertex/us-west2
  • Novita100%
    Healthy24h 99.98%
    p50
    1.75s
    p99
    8.88s
    Throughput
    18 tps
    Context
    131K
    $/Mtok
    $0.270 / $1.00
    novita/fp8fp8
  • SiliconFlow99.52%
    Healthy24h 97.45%
    p50
    1.37s
    p99
    9.53s
    Throughput
    15 tps
    Context
    164K
    $/Mtok
    $0.270 / $1.00
    siliconflow/fp8fp8
  • Mara54.17%
    Down24h 86.49%
    p50
    3.27s
    p99
    18.5s
    Throughput
    24 tps
    Context
    131K
    $/Mtok
    $0.600 / $1.70
    mara
  • SambaNova
    Degraded24h 92.10%
    p50
    1.77s
    p99
    12.0s
    Throughput
    31 tps
    Context
    131K
    $/Mtok
    $0.650 / $1.50
    sambanova/fp8fp8

Measured 1m ago. Latency and throughput are OpenRouter's rolling 30-minute windows.

Availability

Axis starts at 90% so small differences stay visible; it expands to 0 when an endpoint falls below.

Time to first token

Logarithmic axis — the fleet spans two orders of magnitude. p50 and p99 are OpenRouter's rolling 30-minute windows.

Throughput

Median tokens per second.