Inkling
thinkingmachines/inkling
Compare with Inkling Small, Inkling (batch), GLM 5.2
CompareInkling
| Provider | State | Uptime 5m | 24h | p50 latency | p99 | Throughput | Context | $/Mtok |
|---|---|---|---|---|---|---|---|---|
DeepInfradeepinfra/fp8fp8 | Healthy | 100% | 99.61% | 522ms | 2.46s | 100 tps | 524K | $0.950 / $4.05 |
Togethertogether | Healthy | 100% | 98.20% | 455ms | 1.54s | 122 tps | 524K | $1.00 / $4.05 |
BaseTenbaseten/fp8fp8 | No data | — | 99.95% | 329ms | 962ms | 233 tps | 1.0M | $1.00 / $4.05 |
- DeepInfra100%Healthy24h 99.61%
- p50
- 522ms
- p99
- 2.46s
- Throughput
- 100 tps
- Context
- 524K
- $/Mtok
- $0.950 / $4.05
deepinfra/fp8fp8 - Together100%Healthy24h 98.20%
- p50
- 455ms
- p99
- 1.54s
- Throughput
- 122 tps
- Context
- 524K
- $/Mtok
- $1.00 / $4.05
together - BaseTen—No data24h 99.95%
- p50
- 329ms
- p99
- 962ms
- Throughput
- 233 tps
- Context
- 1.0M
- $/Mtok
- $1.00 / $4.05
baseten/fp8fp8
Measured 3m ago. Latency and throughput are OpenRouter's rolling 30-minute windows.
History
1 endpoint reporting no data is omitted from the charts.
Availability
Axis starts at 90% so small differences stay visible; it expands to 0 when an endpoint falls below.
Time to first token
Logarithmic axis — the fleet spans two orders of magnitude. p50 and p99 are OpenRouter's rolling 30-minute windows.
Throughput
Median tokens per second.