← Leaderboard

Gemma 3 4B QAT 4bit MLXSMALL

Tested Rank #26 of 51 · 7/8 GAUNTLET progress
Generalist — 57.5 / 100 G Agentic — 86 / 100 A Understanding — 49.3 / 100 U Needle — 46.8 / 100 N Thinking — pending T Live — 37.9 / 100 L Engineering — 51 / 100 E Throughput — 46.3 / 100 T

Hollow markers & dashed spokes: axis not yet scored

G
57.5
A
86
U
49.3
N
46.8
T
n/d
L
37.9
E
51
T
46.3

Specification

Parameters
4B
Architecture
gemma3
Size on disk
2.8 GB
Quantization
QAT 4bit
Format
MLX
Runtime
LM Studio
Class
Small (under 8B total parameters — phone/edge class)
Reasoning (CoT)
No
Internal ID
M16
Mean speed
162.6 tok/s across suites
Stall census
0 stalls in 13 observed tests (0.0%)
Model card
lmstudio.ai

Suite results

Suite Score Avg / 20 Tests n tok/s Run
General capability (13-task real-workload suite) 1a 149.5 / 260 11.5 13/13 n=1 171.8 run 2026-08-04
Test Category Score tok/s
A1 Dev/Ops Scripting 8/20 167.1
A2 Dev/Ops Scripting 12/20 166.2
A3 Dev/Ops Scripting 5/20 161.7
A4 Dev/Ops Scripting 5/20 145.4
A5 Dev/Ops Scripting 8/20 166.6
A6 Dev/Ops Scripting 12/20 165.8
B1 Document Processing 18/20 158.3
B2 Document Processing 20/20 163.6
B3 Document Processing 8/20 167.4
C1 Content Production 18/20 167
C2 Content Production 10/20 164.8
C3 Content Production 15/20 167.1
C4 Content Production 11/20 162.5

Category mean Dev/Ops Scripting 8.3Document Processing 15.3Content Production 13.5

Agentic tool-calling & protocol adherence 1b 137.6 / 160 17.2 8/8 n=1 171.3 run 2026-09-22
Test Category Score tok/s
WA1 Tool Calling & Protocol Adherence 20/20 102.8
WA2 Tool Calling & Protocol Adherence 20/20 111.1
WA3 Tool Calling & Protocol Adherence 13/20 115.7
WA4 Tool Calling & Protocol Adherence 20/20 96.2
WA5 Tool Calling & Protocol Adherence 19/20 139.6
WB1 Agent Robustness & State Management 18/20 166.1
WB2 Agent Robustness & State Management 20/20 167.1
WB3 Agent Robustness & State Management 8/20 130.5

Category mean Tool Calling & Protocol Adherence 18.4Agent Robustness & State Management 15.3

Coding depth 1c 91.8 / 180 10.2 9/9 n=1 169.5 run 2026-09-29
Test Category Score tok/s
K1 Coding — Concurrency 16/20 145.5
K2 Coding — Algorithms 8/20 166.3
K3 Coding — Security Review 8/20 168
K4 Coding — Performance 18/20 167.2
K5 Coding — Refactoring Under Constraints 10/20 167.9
K6 Coding — Test Writing 12/20 163.4
K7 Coding — React Component (typed) 8/20 165.4
K8 Coding — TypeScript Type System 6/20 161.9
K9 Coding — React Debugging 6/20 167

Category mean Coding — Concurrency 16Coding — Algorithms 8Coding — Security Review 8Coding — Performance 18Coding — Refactoring Under Constraints 10Coding — Test Writing 12Coding — React Component (typed) 8Coding — TypeScript Type System 6Coding — React Debugging 6

Doc/OCR vision 1d 87.2 / 160 10.9 8/8 n=1 169.6 run 2026-09-22
Test Category Score tok/s
D1 Doc/OCR — Transcription 6/20 119.4
D2 Doc/OCR — Document QA 11/20 87.8
D3 Doc/OCR — Field Extraction 3/20 133.7
D4 Doc/OCR — Document QA 13/20 118
D5 Doc/OCR — Transcription 6/20 124.4
D6 Doc/OCR — Field Extraction 17/20 126
D7 Doc/OCR — Degraded Input 11/20 137.8
D8 Doc/OCR — Degraded Input 20/20 133.1

Category mean Doc/OCR — Transcription 6Doc/OCR — Document QA 12Doc/OCR — Field Extraction 10Doc/OCR — Degraded Input 15.5

Doc/OCR — real-degraded tier 1d2 31.2 / 80 7.8 4/4 n=1 176.5 run 2026-09-20
Test Category Score tok/s
D9 Doc/OCR — Real Degraded Transcription 3/20 85.5
D10 Doc/OCR — Handwriting Extraction 16/20 104.3
D11 Doc/OCR — Degraded Table Transcription 3/20 118
D12 Doc/OCR — Degraded Table QA 9/20 74.6

Category mean Doc/OCR — Real Degraded Transcription 3Doc/OCR — Handwriting Extraction 16Doc/OCR — Degraded Table Transcription 3Doc/OCR — Degraded Table QA 9

Content-production depth 1e 73.2 / 120 12.2 6/6 n=1 150.9 run 2026-09-24
Test Category Score tok/s
E1 SOW & Requirements Decomposition 15/20 129.3
E2 SOW & Requirements Decomposition 18/20 148.4
E3 Executive & Proposal Content 12/20 137.9
E4 Executive & Proposal Content 6/20 142.9
E5 Executive & Proposal Content 18/20 144.7
E6 Executive & Proposal Content 4/20 159.1

Category mean SOW & Requirements Decomposition 16.5Executive & Proposal Content 10

Long-context retrieval & synthesis 1g 76.2 / 120 12.7 6/6 n=1 152.3 run 2026-09-22
Test Category Score tok/s
L1 Long-Context — Needle Retrieval 20/20 19.7
L2 Long-Context — Grounded QA with Distractors 7/20 15.4
L3 Long-Context — Ops Log Reasoning 6/20 45.7
L4 Long-Context — Faithful Summarization 6/20 67.2
L5 Long-Context — Instruction Retention 17/20 16.9
L6 Long-Context — Full-Haystack QA 20/20 5.4

Category mean Long-Context — Needle Retrieval 20Long-Context — Grounded QA with Distractors 7Long-Context — Ops Log Reasoning 6Long-Context — Faithful Summarization 6Long-Context — Instruction Retention 17Long-Context — Full-Haystack QA 20

Long-context multi-needle (MRCR) 1g2 8.1 / 60 2.7 3/3 n=1 137.3 run 2026-09-18
Test Category Score tok/s
L7 Long-Context — Multi-Needle Retrieval (MRCR) 0/20 58.2
L8 Long-Context — Multi-Needle Retrieval (MRCR) 4/20 24.4
L9 Long-Context — Multi-Needle Retrieval (MRCR) 4/20 107.3
Live one-shot builds (runtime-verified) 1h 19.2 / 80 4.8 4/4 n=1 169.3 run 2026-09-18
Test Category Score tok/s
H1 Web-Estate One-Shot Builds 7/20 166.6
H2 Web-Estate One-Shot Builds 6/20 167.2
H3 Web-Estate One-Shot Builds 6/20 167.1
H4 Web-Estate One-Shot Builds 0/20 165.4
Brand-constrained static build (negative-constraint adherence) 1j 49 / 100 9.8 5/5 n=1 157.4 run 2026-09-22
Test Category Score tok/s
J1 Brand-Constrained Static Build 10/20 140.3
J3 Brand-Constrained Static Build 12/20 146.7
J4 Brand-Constrained Static Build 12/20 147
J5 Brand-Constrained Static Build 7/20 153.7
J6 Brand-Constrained Static Build 8/20 166.1

Each test is LLM-judged 0–20 against a fixed rubric; suite max = tests × 20. An ADJ chip marks a cell carrying a subjective adjustment from the maintainer; the judged score is the one shown, is unchanged, and remains what the axes, ranks and tie bands are computed from. Hover for the reason. Rows with a ▶ expand to the per-test breakdown. Where fewer tests ran than the suite total, unrun tests count as zero toward the axis score. See methodology.