modelstatus.dev

Inkling

thinkingmachines/inkling

Compare with Inkling Small, Inkling (batch), GLM 5.2

CompareInkling
  • Healthy24h 99.98%
    p50
    379ms
    p99
    1.34s
    Throughput
    152 tps
    Context
    1.0M
    $/Mtok
    $1.00 / $4.05
    baseten/fp8fp8
  • Healthy24h 99.45%
    p50
    361ms
    p99
    3.57s
    Throughput
    105 tps
    Context
    524K
    $/Mtok
    $1.00 / $4.05
    together
  • DeepInfra98.75%
    Healthy24h 99.69%
    p50
    749ms
    p99
    13.0s
    Throughput
    45 tps
    Context
    524K
    $/Mtok
    $0.950 / $4.05
    deepinfra/fp8fp8

Measured 4m ago. Latency and throughput are OpenRouter's rolling 30-minute windows.

Availability

Axis starts at 90% so small differences stay visible; it expands to 0 when an endpoint falls below.

Time to first token

Logarithmic axis — the fleet spans two orders of magnitude. P50 is OpenRouter's rolling 30-minute window, averaged over the hour.

Throughput

Median tokens per second.