modelstatus.dev

DeepSeek V3.1

deepseek/deepseek-chat-v3.1

Compare with DeepSeek V3.1 Terminus, DeepSeek V3, DeepSeek V3.2

CompareDeepSeek V3.1
  • AtlasCloud100%
    Healthy24h 97.51%
    p50
    1.73s
    p99
    4.93s
    Throughput
    20 tps
    Context
    131K
    $/Mtok
    $0.300 / $0.950
    atlas-cloud/fp8fp8
  • CoreWeave100%
    Healthy24h 99.96%
    p50
    371ms
    p99
    1.02s
    Throughput
    44 tps
    Context
    161K
    $/Mtok
    $0.550 / $1.65
    coreweave/fp8fp8
  • DeepInfra100%
    Healthy24h 99.21%
    p50
    776ms
    p99
    7.50s
    Throughput
    11 tps
    Context
    164K
    $/Mtok
    $0.250 / $0.950
    deepinfra/fp4fp4
  • Google100%
    Down24h 98.17%
    p50
    729ms
    p99
    8.77s
    Throughput
    95 tps
    Context
    164K
    $/Mtok
    $0.600 / $1.70
    google-vertex/us-west2
  • Novita100%
    Healthy24h 99.98%
    p50
    1.76s
    p99
    7.54s
    Throughput
    22 tps
    Context
    131K
    $/Mtok
    $0.270 / $1.00
    novita/fp8fp8
  • SiliconFlow100%
    Healthy24h 97.51%
    p50
    1.48s
    p99
    8.37s
    Throughput
    15 tps
    Context
    164K
    $/Mtok
    $0.270 / $1.00
    siliconflow/fp8fp8
  • Mara64.38%
    Down24h 85.85%
    p50
    3.06s
    p99
    19.5s
    Throughput
    79 tps
    Context
    131K
    $/Mtok
    $0.600 / $1.70
    mara
  • SambaNova
    No data24h 91.99%
    p50
    3.02s
    p99
    20.5s
    Throughput
    26 tps
    Context
    131K
    $/Mtok
    $0.650 / $1.50
    sambanova/fp8fp8

Measured 52s ago. Latency and throughput are OpenRouter's rolling 30-minute windows.

1 endpoint reporting no data is omitted from the charts.

Availability

Axis starts at 90% so small differences stay visible; it expands to 0 when an endpoint falls below.

Time to first token

Logarithmic axis — the fleet spans two orders of magnitude. p50 and p99 are OpenRouter's rolling 30-minute windows, averaged over the hour.

Throughput

Median tokens per second.