modelstatus.dev

Kimi K2.5

moonshotai/kimi-k2.5

Compare with Kimi K2.6, Kimi K2.7 Code, Kimi K3

CompareKimi K2.5
  • DigitalOcean100%
    Healthy24h 99.45%
    p50
    628ms
    p99
    3.90s
    Throughput
    20 tps
    Context
    262K
    $/Mtok
    $0.375 / $2.02
    digitalocean
  • Moonshot AI100%
    Healthy24h 99.98%
    p50
    2.04s
    p99
    11.5s
    Throughput
    32 tps
    Context
    262K
    $/Mtok
    $0.600 / $3.00
    moonshotai/int4int4
  • SiliconFlow100%
    Healthy24h 99.91%
    p50
    1.45s
    p99
    23.5s
    Throughput
    41 tps
    Context
    262K
    $/Mtok
    $0.450 / $2.25
    siliconflow/int4int4
  • StreamLake100%
    Healthy24h 99.91%
    p50
    1.13s
    p99
    9.37s
    Throughput
    35 tps
    Context
    256K
    $/Mtok
    $0.540 / $2.70
    streamlake/fp8fp8
  • Novita99.81%
    Healthy24h 98.35%
    p50
    1.13s
    p99
    10.7s
    Throughput
    37 tps
    Context
    262K
    $/Mtok
    $0.570 / $2.85
    novita
  • AtlasCloud99.61%
    Healthy24h 98.46%
    p50
    1.40s
    p99
    9.49s
    Throughput
    25 tps
    Context
    262K
    $/Mtok
    $0.490 / $2.50
    atlas-cloud/int4int4
  • DeepInfra99.10%
    Healthy24h 99.30%
    p50
    779ms
    p99
    15.9s
    Throughput
    31 tps
    Context
    262K
    $/Mtok
    $0.450 / $2.25
    deepinfra/fp4fp4
  • Venice69.91%
    Down24h 96.31%
    p50
    1.94s
    p99
    14.0s
    Throughput
    29 tps
    Context
    256K
    $/Mtok
    $0.532 / $3.32
    venice
  • Amazon Bedrock
    No data24h 97.03%
    p50
    1.40s
    p99
    8.30s
    Throughput
    35 tps
    Context
    262K
    $/Mtok
    $0.600 / $3.00
    amazon-bedrock/us-east-2
  • Phala
    No data24h 85.44%
    p50
    2.50s
    p99
    12.9s
    Throughput
    42 tps
    Context
    262K
    $/Mtok
    $0.600 / $3.00
    phala

Measured 2m ago. Latency and throughput are OpenRouter's rolling 30-minute windows.

2 endpoints reporting no data are omitted from the charts.

Availability

Axis starts at 90% so small differences stay visible; it expands to 0 when an endpoint falls below.

Time to first token

Logarithmic axis — the fleet spans two orders of magnitude. p50 and p99 are OpenRouter's rolling 30-minute windows, averaged over the hour.

Throughput

Median tokens per second.