gemma-4-31b, provider by provider
All cells probed live · latest probe 2026-07-27 · how grading works
One model, 11 hosted endpoints. The capabilities below are defined by the model's own chat template — the text program, shipped with the model, that turns a conversation into the exact input the model reads. Every provider is graded on whether it delivers them.
inferencecanary · gemma-4-31b · probed 2026-07-27
inferencecanary.com
| Provider | Harness | Billing | Caching | 20K-token agentic turn | Faithfulness | |
|---|---|---|---|---|---|---|
| Speed | Cost | |||||
| Google (native)reference generateContent — Google's own protocol | A− 1 yellow | ✓ | testing | — | TBD | coming soon |
| Parasail OpenAI chat/completions | A− 1 yellow | ✓ | Sometimes | 17.4–25.1s | 0.13–0.31¢ | coming soon |
| Lightning OpenAI chat/completions | B 2 yellow | ✓ | Yes | 1.2s | 0.29¢ | coming soon |
| Cerebras OpenAI chat/completions | B− 1 red · 1 yellow | ✓ | Yes | 0.7s | 2.0¢ | coming soon |
| Together AI OpenAI chat/completions | B− 1 red · 1 yellow | ✓ | Sometimes | 3.4–6.6s | 0.80¢ | coming soon |
| Novita OpenAI chat/completions | C− 2 red · 1 yellow | ✓ | No | ~30s | 0.29¢ | coming soon |
| SambaNova OpenAI chat/completions | C− 2 red · 1 yellow | RED2.18× overcharge | No | 4.9s | 0.79¢ | coming soon |
| DeepInfra OpenAI chat/completions | C− 2 red · 1 yellow | ✓ | Sometimes | 24.8–29.7s | 0.27¢ | coming soon |
| AWS Bedrock (Chat Completions) OpenAI chat/completions on Bedrock Mantle | D 2 red · 2 yellow | ✓footnote on card | No | 3.3s | 0.29¢ | coming soon |
| AWS Bedrock (Responses) OpenAI Responses on Bedrock Mantle | D− 2 red · 3 yellow | ✓footnote on card | No | — | 0.29¢ | coming soon |
| Google (OpenAI-compat) OpenAI chat/completions shim over generateContent | D− 3 red · 1 yellow | ✓footnote on card | testing | — | TBD | coming soon |
What the model can do — and who delivers it
- Think in a separate channel. The template defines a reasoning phase ahead of the answer, surfaced by every dialect as a structured reasoning field — except one: Google's own OpenAI-compat shim has no reasoning channel at all, leaking thought text into the answer field.
- Keep its reasoning across tool calls. The template renders the model's prior thinking back to it on later turns. Four providers delete it before the model can read it.
- Read images returned by tools. The reference protocol delivers them and three OpenAI-compat providers prove the standard shape works — yet 7 of 11 providers reject, 500, or silently blind them, closing off agentic multimodal.
- Call tools — several at once. Every provider but the two Bedrock APIs
emits parallel calls; Bedrock never does.
Constraining the choice with
tool_choiceis honored on eight providers and ignored, validated-but-unenforced, or enforced by after-the-fact error on the other three. - Return byte-exact tool arguments. Nine providers agree with the reference; both Bedrock APIs apply an extra unescape pass, corrupting any argument whose bytes matter.
Every graded endpoint
- A− Google (native) reference — generateContent — Google's own protocol
- A− Parasail — OpenAI chat/completions
- B Lightning — OpenAI chat/completions
- B− Cerebras — OpenAI chat/completions
- B− Together AI — OpenAI chat/completions
- C− Novita — OpenAI chat/completions
- C− SambaNova — OpenAI chat/completions
- C− DeepInfra — OpenAI chat/completions
- D AWS Bedrock (Chat Completions) — OpenAI chat/completions on Bedrock Mantle
- D− AWS Bedrock (Responses) — OpenAI Responses on Bedrock Mantle
- D− Google (OpenAI-compat) — OpenAI chat/completions shim over generateContent
Vendor roll-ups: Google · Parasail · Lightning AI · Cerebras · Together AI · Novita AI · SambaNova · DeepInfra · AWS.