# Lightning — gemma-4-31b grade card

OpenAI chat/completions · serves `lightning-ai/gemma-4-31B-it` as gemma-4-31b · probed live 2026-07-27

HTML version: [https://inferencecanary.com/lightning/gemma-4-31b/](https://inferencecanary.com/lightning/gemma-4-31b/) · matrix: [https://inferencecanary.com/index.md](https://inferencecanary.com/index.md)

## Grades


| Axis | Grade | Basis |
|---|---|---|
| Harness | B | A − 0 red − 2 yellow, per the published ladder |
| Billing | OK | Reported usage consistent with bytes on the wire. |
| Caching | Yes (3/3 hits) | Reused already-seen context on 3/3 cold→warm trials. Speed numbers: [https://inferencecanary.com/speed/index.md](https://inferencecanary.com/speed/index.md) |
| Faithfulness | 20/29 | Agreement with the reference provider over the 29-cell deterministic battery (temperature 0, thinking disabled). |

## Flags

- **[yellow · harness] Tool results replayed out of order are mis-paired on most providers** — The chat dialect's contract is that a tool message is matched to its call by tool_call_id, in any order. After two parallel calls, replaying the results reversed swaps the data between them on these providers — the template renders by name, and these providers render positionally, dropping their own dialect's id contract. An async harness that appends results in completion order silently swaps payloads. One provider pairs correctly, proving the fix is implementable provider-side. (details: [/findings/index.md](https://inferencecanary.com/findings/index.md#tool-results-out-of-order))
- **[yellow · harness] Lightning leaks template reasoning markers into content on truncation** — At tight max_tokens, literal template markers appear in content; unconstrained budgets come back clean. Output is polluted exactly when the response is cut short. (details: [/findings/index.md](https://inferencecanary.com/findings/index.md#lightning-marker-leak))

