# gemma-4-31b, provider by provider — inferencecanary

All cells probed live; latest probe 2026-07-27.
HTML version: [https://inferencecanary.com/gemma-4-31b/](https://inferencecanary.com/gemma-4-31b/)

One model, 11 hosted endpoints. The capabilities below are defined by the
model's own chat template — the text program, shipped with the model, that turns a conversation
into the exact input the model reads. Every provider is graded on whether it delivers them.

## The grade sheet

| Provider | Dialect | Harness | Billing | Caching | 20K-token agentic turn — speed | 20K-token agentic turn — cost | Faithfulness |
|---|---|---|---|---|---|---|---|
| [Google (native)](https://inferencecanary.com/google/gemma-4-31b/index.md) (reference) | generateContent — Google's own protocol | A− | OK | testing | — | TBD | coming soon |
| [Parasail](https://inferencecanary.com/parasail/gemma-4-31b/index.md) | OpenAI chat/completions | A− | OK | Sometimes | 17.4–25.1s | 0.13–0.31¢ | coming soon |
| [Lightning](https://inferencecanary.com/lightning/gemma-4-31b/index.md) | OpenAI chat/completions | B | OK | Yes | 1.2s | 0.29¢ | coming soon |
| [Cerebras](https://inferencecanary.com/cerebras/gemma-4-31b/index.md) | OpenAI chat/completions | B− | OK | Yes | 0.7s | 2.0¢ | coming soon |
| [Together AI](https://inferencecanary.com/togetherai/gemma-4-31b/index.md) | OpenAI chat/completions | B− | OK | Sometimes | 3.4–6.6s | 0.80¢ | coming soon |
| [Novita](https://inferencecanary.com/novita/gemma-4-31b/index.md) | OpenAI chat/completions | C− | OK | No | ~30s | 0.29¢ | coming soon |
| [SambaNova](https://inferencecanary.com/sambanova/gemma-4-31b/index.md) | OpenAI chat/completions | C− | RED (2.18× overcharge) | No | 4.9s | 0.79¢ | coming soon |
| [DeepInfra](https://inferencecanary.com/deepinfra/gemma-4-31b/index.md) | OpenAI chat/completions | C− | OK | Sometimes | 24.8–29.7s | 0.27¢ | coming soon |
| [AWS Bedrock (Chat Completions)](https://inferencecanary.com/aws/gemma-4-31b/index.md) | OpenAI chat/completions on Bedrock Mantle | D | OK | No | 3.3s | 0.29¢ | coming soon |
| [AWS Bedrock (Responses)](https://inferencecanary.com/aws-responses/gemma-4-31b/index.md) | OpenAI Responses on Bedrock Mantle | D− | OK | No | — | 0.29¢ | coming soon |
| [Google (OpenAI-compat)](https://inferencecanary.com/google-openai/gemma-4-31b/index.md) | OpenAI chat/completions shim over generateContent | D− | OK | testing | — | TBD | coming soon |


## What the model can do — and who delivers it

- **Think in a separate channel.** The template defines a reasoning phase ahead of the answer,
  surfaced by every dialect as a structured reasoning field — except one: Google's own
  OpenAI-compat shim has no reasoning channel at all, leaking thought text into the answer
  field ([finding](https://inferencecanary.com/findings/index.md#google-openai-no-reasoning-channel)).
- **Keep its reasoning across tool calls.** The template renders the model's prior thinking
  back to it on later turns. Four providers delete it before the model can read it
  ([full writeup](https://inferencecanary.com/replayed-reasoning-dropped/index.md)).
- **Read images returned by tools.** The reference protocol delivers them and three
  OpenAI-compat providers prove the standard shape works — yet 7 of 11 providers
  reject, 500, or silently blind them, closing off agentic multimodal
  ([finding](https://inferencecanary.com/findings/index.md#tool-result-images)).
- **Call tools — several at once.** Every provider but the two Bedrock APIs emits parallel
  calls; Bedrock never does ([finding](https://inferencecanary.com/findings/index.md#bedrock-single-tool-call)).
  Constraining the choice with `tool_choice` is honored on eight providers and ignored,
  validated-but-unenforced, or enforced by after-the-fact error on the other three.
- **Return byte-exact tool arguments.** Nine providers agree with the reference; both Bedrock
  APIs apply an extra unescape pass, corrupting any argument whose bytes matter
  ([finding](https://inferencecanary.com/findings/index.md#bedrock-argument-corruption)).

## Every graded endpoint

- **A−** [Google (native)](https://inferencecanary.com/google/gemma-4-31b/index.md) (reference) — generateContent — Google's own protocol
- **A−** [Parasail](https://inferencecanary.com/parasail/gemma-4-31b/index.md) — OpenAI chat/completions
- **B** [Lightning](https://inferencecanary.com/lightning/gemma-4-31b/index.md) — OpenAI chat/completions
- **B−** [Cerebras](https://inferencecanary.com/cerebras/gemma-4-31b/index.md) — OpenAI chat/completions
- **B−** [Together AI](https://inferencecanary.com/togetherai/gemma-4-31b/index.md) — OpenAI chat/completions
- **C−** [Novita](https://inferencecanary.com/novita/gemma-4-31b/index.md) — OpenAI chat/completions
- **C−** [SambaNova](https://inferencecanary.com/sambanova/gemma-4-31b/index.md) — OpenAI chat/completions
- **C−** [DeepInfra](https://inferencecanary.com/deepinfra/gemma-4-31b/index.md) — OpenAI chat/completions
- **D** [AWS Bedrock (Chat Completions)](https://inferencecanary.com/aws/gemma-4-31b/index.md) — OpenAI chat/completions on Bedrock Mantle
- **D−** [AWS Bedrock (Responses)](https://inferencecanary.com/aws-responses/gemma-4-31b/index.md) — OpenAI Responses on Bedrock Mantle
- **D−** [Google (OpenAI-compat)](https://inferencecanary.com/google-openai/gemma-4-31b/index.md) — OpenAI chat/completions shim over generateContent

Vendor roll-ups: [Google](https://inferencecanary.com/google/index.md) · [Parasail](https://inferencecanary.com/parasail/index.md) · [Lightning AI](https://inferencecanary.com/lightning/index.md) · [Cerebras](https://inferencecanary.com/cerebras/index.md) · [Together AI](https://inferencecanary.com/togetherai/index.md) · [Novita AI](https://inferencecanary.com/novita/index.md) · [SambaNova](https://inferencecanary.com/sambanova/index.md) · [DeepInfra](https://inferencecanary.com/deepinfra/index.md) · [AWS](https://inferencecanary.com/aws/index.md)
