modelstatus.dev

Inkling Small

thinkingmachines/inkling-small

Compare with Inkling, Inkling (batch), GLM 5.2

CompareInkling Small
  • Healthy24h 98.07%
    7d
    98.30%
    30d
    98.30%
    p50
    414ms
    p99
    1.52s
    Throughput
    165 tps
    Context
    524K
    $/Mtok
    $0.450 / $1.20
    deepinfra/fp8fp8
  • No data24h 98.94%
    7d
    97.68%
    30d
    97.68%
    p50
    377ms
    p99
    1.99s
    Throughput
    202 tps
    Context
    1.0M
    $/Mtok
    $0.500 / $1.20
    baseten/fp8fp8
  • No data24h 98.85%
    7d
    94.51%
    30d
    94.51%
    p50
    446ms
    p99
    6.20s
    Throughput
    89 tps
    Context
    524K
    $/Mtok
    $0.500 / $1.20
    together

Measured 4m ago. Latency and throughput are OpenRouter's rolling 30-minute windows.

2 endpoints reporting no data are omitted from the charts.

Availability

Axis starts at 90% so small differences stay visible; it expands to 0 when an endpoint falls below.

Time to first token

Logarithmic axis — the fleet spans two orders of magnitude. P50 is OpenRouter's rolling 30-minute window.

Throughput

Median tokens per second.