Inkling Small
thinkingmachines/inkling-small
Compare with Inkling, Inkling (batch), GLM 5.2
CompareInkling Small
| Provider | State | Uptime 5m | 24h | p50 latency | p99 | Throughput | Context | $/Mtok |
|---|---|---|---|---|---|---|---|---|
DeepInfradeepinfra/fp8fp8 | Healthy | 100% | 99.05% | 424ms | 1.86s | 98 tps | 524K | $0.450 / $1.20 |
Togethertogether | No data | — | 99.62% | 3.18s | 12.7s | 31 tps | 524K | $0.500 / $1.20 |
- DeepInfra100%Healthy24h 99.05%
- p50
- 424ms
- p99
- 1.86s
- Throughput
- 98 tps
- Context
- 524K
- $/Mtok
- $0.450 / $1.20
deepinfra/fp8fp8 - Together—No data24h 99.62%
- p50
- 3.18s
- p99
- 12.7s
- Throughput
- 31 tps
- Context
- 524K
- $/Mtok
- $0.500 / $1.20
together
Measured 5m ago. Latency and throughput are OpenRouter's rolling 30-minute windows.
History
1 endpoint reporting no data is omitted from the charts.
Availability
Axis starts at 90% so small differences stay visible; it expands to 0 when an endpoint falls below.
Time to first token
Logarithmic axis — the fleet spans two orders of magnitude. p50 and p99 are OpenRouter's rolling 30-minute windows.
Throughput
Median tokens per second.