modelstatus.dev

Inkling

thinkingmachines/inkling

Compare with Inkling Small, Inkling (batch), GLM 5.2

CompareInkling
  • Healthy24h 99.94%
    7d
    99.99%
    30d
    99.99%
    p50
    349ms
    p99
    2.80s
    Throughput
    95 tps
    Context
    1.0M
    $/Mtok
    $1.00 / $4.05
    baseten/fp8fp8
  • Healthy24h 99.50%
    7d
    98.81%
    30d
    98.81%
    p50
    2.94s
    p99
    32.7s
    Throughput
    11 tps
    Context
    524K
    $/Mtok
    $0.950 / $4.05
    deepinfra/fp8fp8
  • Together98.11%
    Healthy24h 98.61%
    7d
    99.28%
    30d
    99.28%
    p50
    1.32s
    p99
    9.78s
    Throughput
    33 tps
    Context
    524K
    $/Mtok
    $1.00 / $4.05
    together

Measured 2m ago. Latency and throughput are OpenRouter's rolling 30-minute windows.

Availability

Axis starts at 90% so small differences stay visible; it expands to 0 when an endpoint falls below.

Time to first token

Logarithmic axis — the fleet spans two orders of magnitude. P50 is OpenRouter's rolling 30-minute window.

Throughput

Median tokens per second.