← Leaderboard

Gemma-4 31B IT 4bit MLX

Tested Rank #23 of 39 · 3/8 GAUNTLET progress
Generalist — 81.7 / 100 G Agentic — 100 / 100 A Understanding — pending U Needle — pending N Thinking — pending T Live — pending L Engineering — pending E Throughput — 8.6 / 100 T

Hollow markers & dashed spokes: axis not yet scored

G
81.7
A
100
U
N
T
n/d
L
E
T
8.6

Specification

Parameters
31.27B
Architecture
gemma4
Size on disk
17.18 GB
Quantization
4bit
Format
MLX
Class
Standard
Reasoning (CoT)
Yes — emits reasoning tokens
Internal ID
M33
Mean speed
20.6 tok/s across suites
Model card
huggingface.co

Suite results

Suite Score Avg / 20 Tests tok/s Run
General capability (13-task real-workload suite) 1a 212.3 / 260 19.3 11/13 17.6 run 2026-09-06
Test Category Score tok/s
A1 Dev/Ops Scripting 17/20 25.6
A2 Dev/Ops Scripting 18/20 23.3
A4 Dev/Ops Scripting 19/20 20.1
A5 Dev/Ops Scripting 20/20 9.2
A6 Dev/Ops Scripting 20/20 19.1
B1 Document Processing 20/20 13.9
B2 Document Processing 20/20 14.6
C1 Content Production 20/20 17
C2 Content Production 18/20 16.4
C3 Content Production 20/20 13.4
C4 Content Production 20/20 21.1

Category mean Dev/Ops Scripting 18.8Document Processing 20Content Production 19.5

Agentic tool-calling & protocol adherence 1b 160 / 160 20 8/8 23.5 run 2026-09-06
Test Category Score tok/s
WA1 Tool Calling & Protocol Adherence 20/20 24.7
WA2 Tool Calling & Protocol Adherence 20/20 25.4
WA3 Tool Calling & Protocol Adherence 20/20 25.3
WA4 Tool Calling & Protocol Adherence 20/20 23.6
WA5 Tool Calling & Protocol Adherence 20/20 25.9
WB1 Agent Robustness & State Management 20/20 21.1
WB2 Agent Robustness & State Management 20/20 22.3
WB3 Agent Robustness & State Management 20/20 19.9

Category mean Tool Calling & Protocol Adherence 20Agent Robustness & State Management 20

Each test is LLM-judged 0–20 against a fixed rubric; suite max = tests × 20. Rows with a expand to the per-test breakdown. Where fewer tests ran than the suite total, unrun tests count as zero toward the axis score. See methodology.