Inkling Small
thinkingmachines/inkling-small
Compare with Inkling, Inkling (batch), GLM 5.2
CompareInkling Small
| Provider | State | Uptime 5m | 24h | p50 latency | p99 | Throughput | Context | $/Mtok |
|---|---|---|---|---|---|---|---|---|
DeepInfradeepinfra/fp8fp8 | Healthy | 100% | 99.06% | 694ms | 2.01s | 137 tps | 524K | $0.450 / $1.20 |
Togethertogether | No data | — | 99.17% | 2.10s | 18.0s | 27 tps | 524K | $0.500 / $1.20 |
- DeepInfra100%Healthy24h 99.06%
- p50
- 694ms
- p99
- 2.01s
- Throughput
- 137 tps
- Context
- 524K
- $/Mtok
- $0.450 / $1.20
deepinfra/fp8fp8 - Together—No data24h 99.17%
- p50
- 2.10s
- p99
- 18.0s
- Throughput
- 27 tps
- Context
- 524K
- $/Mtok
- $0.500 / $1.20
together
Measured 1m ago. Latency and throughput are OpenRouter's rolling 30-minute windows.
History
1 endpoint reporting no data is omitted from the charts.
Availability
Axis starts at 90% so small differences stay visible; it expands to 0 when an endpoint falls below.
Time to first token
Logarithmic axis — the fleet spans two orders of magnitude. p50 and p99 are OpenRouter's rolling 30-minute windows.
Throughput
Median tokens per second.