Hosted gemma-4-31b, graded.

All cells probed live · latest probe 2026-07-27 · model gemma-4-31b · how grading works

A provider has one job: surface the model's capabilities to application developers. The model's capabilities are defined by its chat template. Everything the template defines and the provider doesn't deliver — under the best dialect they offer — is documented here, with a reproducible probe behind every grade.

red · billing SambaNova bills identical requests inflated token counts — 2.18× average overcharge red · harness Four providers drop replayed reasoning: the model never sees its own prior thinking red · harness Google's OpenAI-compat shim has no reasoning channel; thought text leaks into content red · harness Images in tool results: rejected, 500'd, or silently blinded on 7 of 11 providers
inferencecanary · gemma-4-31b · probed 2026-07-27 inferencecanary.com
Provider Harness Billing Caching 20K-token agentic turn Faithfulness
Speed Cost
Google (native)reference generateContent — Google's own protocol A− 1 yellow testing TBD coming soon
Parasail OpenAI chat/completions A− 1 yellow Sometimes 17.4–25.1s 0.13–0.31¢ coming soon
Lightning OpenAI chat/completions B 2 yellow Yes 1.2s 0.29¢ coming soon
Cerebras OpenAI chat/completions B− 1 red · 1 yellow Yes 0.7s 2.0¢ coming soon
Together AI OpenAI chat/completions B− 1 red · 1 yellow Sometimes 3.4–6.6s 0.80¢ coming soon
Novita OpenAI chat/completions C− 2 red · 1 yellow No ~30s 0.29¢ coming soon
SambaNova OpenAI chat/completions C− 2 red · 1 yellow RED2.18× overcharge No 4.9s 0.79¢ coming soon
DeepInfra OpenAI chat/completions C− 2 red · 1 yellow Sometimes 24.8–29.7s 0.27¢ coming soon
AWS Bedrock (Chat Completions) OpenAI chat/completions on Bedrock Mantle D 2 red · 2 yellow footnote on card No 3.3s 0.29¢ coming soon
AWS Bedrock (Responses) OpenAI Responses on Bedrock Mantle D− 2 red · 3 yellow footnote on card No 0.29¢ coming soon
Google (OpenAI-compat) OpenAI chat/completions shim over generateContent D− 3 red · 1 yellow footnote on card testing TBD coming soon

How grading works

Grades are computed, not opined. Every flag is a documented defect with live evidence and a stated repro; the letter is arithmetic over the flags. This section is the rubric; each finding documents the probe that produced it.

Axes

The harness ladder

Every provider starts at A. Each red flag — a defect that breaks or corrupts a real application — subtracts a full letter. Each yellow flag — a defect that degrades an application or violates the dialect contract with a workaround — subtracts half a step. The scale floors at F: A, A−, B, B−, C, C−, D, D−, F.

The reference provider (Google native) is graded under the same rules as everyone else — it holds an A− on its own yellow flag, not an exemption.