Speed

measured live 2026-07-27 · model gemma-4-31b · point-in-time, under that evening's load · back to the matrix

What a real agentic call — 20,000 tokens of context in, 256 tokens out — costs in wall-clock, per provider. Two things dominate it: whether the provider reuses context it has already seen, and how fast it reads uncached context. The generation speed providers advertise barely matters at agentic context sizes; the spread between working providers is roughly 13×.

inferencecanary · gemma-4-31b · measured 2026-07-27 inferencecanary.com
Provider Caching Floor (ms) Floor high‑water (ms) Read (tok/s) Write (tok/s) TTFT @20K Turn, cold Turn, hit
Cerebras Yes3/3 hits 197 197 13.9K 1855 1.6s 1.8s 0.7s
Lightning Yes3/3 hits 186 186 21.5K 339 1.1s 1.9s 1.2s
AWS Bedrock (Chat Completions) No0/8 hits 385 406 12.2K 198 2.0s 3.3s
AWS Bedrock (Responses) No0/8 hits 363 371 6.8K 3.3s
SambaNova No0/8 hits 1255 1681 8.6K 189 3.6s 4.9s
Together AI Sometimes2/3 hits 284 284 6.2K 83 3.5s 6.6s 3.4s
DeepInfra Sometimes1/3 hits 418 15,071 3.7K 11 5.8s 29.7s 24.8s
Parasail Sometimes1/3 hits 289 289 2.5K 15 8.3s 25.1s 17.4s
Novita No0/8 hits 439 70,035 3.5–6K 9–12 ~4–6s ~30s
Google (native) testing 638 638 5.3
Google (OpenAI-compat) testing 593 593 5.3

What the columns mean

The measurement: a fresh ~19K-token agentic transcript is read cold, then again warm one second later. A unique marker makes the cold read provably unseen, so a warm answer arriving under 3× the provider's floor can only be cache reuse. TTFT = time to first token.

Notes

Speeds are point-in-time under real load, not a benchmark under ideal conditions — read speed on one provider moved from 7.0K to 12.2K tok/s between same-night runs. The caching class per provider also appears on the grade matrix.