DeepInfra
OpenAI chat/completions · serves google/gemma-4-31B-it as gemma-4-31b ·
probed live 2026-07-27 · back to the matrix
Billing
✓
Reported usage consistent with bytes on the wire.
Caching
Sometimes
Reused already-seen context on 1/3 cold→warm trials — a cache exists but a hit is not dependable. Speed numbers →
Faithfulness
27/29
Agreement with the reference provider over the 29-cell deterministic battery
(temperature 0, thinking disabled). 0–10 score arrives with the locked rubric.
Flags
-
red harness
Images in tool results: rejected, 500'd, or silently blinded on 7 of 11 providers
The model reads images returned by tools — the reference provider delivers them, perception-judged, and three OpenAI-compat providers prove the standard shape works. Seven providers fail on schema choice, not model... -
red harness
DeepInfra validates tool_choice but does not enforce it
Invalid values are rejected with 400/422 — every surface signal says the capability exists — but a valid constraint has zero effect on decoding: "required" returns prose with zero calls, and forcing a tool by name yields... -
yellow harness
Tool results replayed out of order are mis-paired on most providers
The chat dialect's contract is that a tool message is matched to its call by tool_call_id, in any order. After two parallel calls, replaying the results reversed swaps the data between them on these providers — the...
Every flag links to its finding — evidence, repro, and disclosure records live there; the deepest findings have full writeup pages.