# Novita — gemma-4-31b grade card

OpenAI chat/completions · serves `google/gemma-4-31b-it` as gemma-4-31b · probed live 2026-07-27

HTML version: [https://inferencecanary.com/novita/gemma-4-31b/](https://inferencecanary.com/novita/gemma-4-31b/) · matrix: [https://inferencecanary.com/index.md](https://inferencecanary.com/index.md)

## Grades


| Axis | Grade | Basis |
|---|---|---|
| Harness | C− | A − 2 red − 1 yellow, per the published ladder |
| Billing | OK | Reported usage consistent with bytes on the wire. |
| Caching | No (0/8 hits) | No cold→warm trial hit — every call re-reads (and re-bills) the full context. Speed numbers: [https://inferencecanary.com/speed/index.md](https://inferencecanary.com/speed/index.md) |
| Faithfulness | not yet run | Battery not yet run on this provider. |

## Flags

- **[red · harness] Images in tool results: rejected, 500'd, or silently blinded on 7 of 11 providers** — The model reads images returned by tools — the reference provider delivers them, perception-judged, and three OpenAI-compat providers prove the standard shape works. Seven providers fail on schema choice, not model limits: four reject the request (400/422), SambaNova returns a deterministic 500, and AWS Bedrock (Responses) is the worst class — it returns 200 and the model never sees the image, answering confidently about content it never saw. Google's own shim rejects a capability Google's native protocol serves. (details: [/findings/index.md](https://inferencecanary.com/findings/index.md#tool-result-images))
- **[red · harness] Novita silently ignores tool_choice** — "required" and forced-specific constraints are accepted and have no effect on decoding; only "none" is honored. The application consequence matches DeepInfra's: workflows that depend on forced tool calls silently receive prose. (details: [/findings/index.md](https://inferencecanary.com/findings/index.md#novita-tool-choice-ignored))
- **[yellow · harness] Tool results replayed out of order are mis-paired on most providers** — The chat dialect's contract is that a tool message is matched to its call by tool_call_id, in any order. After two parallel calls, replaying the results reversed swaps the data between them on these providers — the template renders by name, and these providers render positionally, dropping their own dialect's id contract. An async harness that appends results in completion order silently swaps payloads. One provider pairs correctly, proving the fix is implementable provider-side. (details: [/findings/index.md](https://inferencecanary.com/findings/index.md#tool-results-out-of-order))

