gemma-4-31b, provider by provider

All cells probed live · latest probe 2026-07-27 · how grading works

One model, 11 hosted endpoints. The capabilities below are defined by the model's own chat template — the text program, shipped with the model, that turns a conversation into the exact input the model reads. Every provider is graded on whether it delivers them.

inferencecanary · gemma-4-31b · probed 2026-07-27 inferencecanary.com
Provider Harness Billing Caching 20K-token agentic turn Faithfulness
Speed Cost
Google (native)reference generateContent — Google's own protocol A− 1 yellow testing TBD coming soon
Parasail OpenAI chat/completions A− 1 yellow Sometimes 17.4–25.1s 0.13–0.31¢ coming soon
Lightning OpenAI chat/completions B 2 yellow Yes 1.2s 0.29¢ coming soon
Cerebras OpenAI chat/completions B− 1 red · 1 yellow Yes 0.7s 2.0¢ coming soon
Together AI OpenAI chat/completions B− 1 red · 1 yellow Sometimes 3.4–6.6s 0.80¢ coming soon
Novita OpenAI chat/completions C− 2 red · 1 yellow No ~30s 0.29¢ coming soon
SambaNova OpenAI chat/completions C− 2 red · 1 yellow RED2.18× overcharge No 4.9s 0.79¢ coming soon
DeepInfra OpenAI chat/completions C− 2 red · 1 yellow Sometimes 24.8–29.7s 0.27¢ coming soon
AWS Bedrock (Chat Completions) OpenAI chat/completions on Bedrock Mantle D 2 red · 2 yellow footnote on card No 3.3s 0.29¢ coming soon
AWS Bedrock (Responses) OpenAI Responses on Bedrock Mantle D− 2 red · 3 yellow footnote on card No 0.29¢ coming soon
Google (OpenAI-compat) OpenAI chat/completions shim over generateContent D− 3 red · 1 yellow footnote on card testing TBD coming soon

What the model can do — and who delivers it

Every graded endpoint

Vendor roll-ups: Google · Parasail · Lightning AI · Cerebras · Together AI · Novita AI · SambaNova · DeepInfra · AWS.