← Leaderboard

Gemma-4 E4B 4B 8bit MLXSMALL

Keeper RAN THE GAUNTLET Rank #9 of 51 · 8/8 GAUNTLET progress
Generalist — 84.5 / 100 G Agentic — 94.5 / 100 A Understanding — 82.3 / 100 U Needle — 83.8 / 100 N Thinking — 100 / 100 T Live — 79 / 100 L Engineering — 85.5 / 100 E Throughput — 22.3 / 100 T
G
84.5
A
94.5
U
82.3
N
83.8
T
100
L
79
E
85.5
T
22.3

Specification

Parameters
4B
Architecture
gemma4
Size on disk
8.97 GB
Quantization
8bit
Format
MLX
Runtime
LM Studio
Class
Small (under 8B total parameters — phone/edge class)
Reasoning (CoT)
Yes — emits reasoning tokens
Internal ID
M1b
Mean speed
78.3 tok/s across suites
Stall census
0 stalls in 21 observed tests (0.0%)
Reasoning appetite
721 tokens mean · 2,398 max
Model card
lmstudio.ai

Suite results

Suite Score Avg / 20 Tests n tok/s Run
General capability (13-task real-workload suite) 1a 219.7 / 260 16.9 13/13 n=2.2 83.4 run 2026-09-02
Test Category Score tok/s
A1 Dev/Ops Scripting 9/20 80.4
A2 Dev/Ops Scripting 16/20 83.2
A3 Dev/Ops Scripting 10/20 82.8
A4 Dev/Ops Scripting 19/20 82.9
A5 Dev/Ops Scripting 13/20 80.8
A6 Dev/Ops Scripting 20/20 82.5
B1 Document Processing 20/20 80.5
B2 Document Processing 20/20 82.3
B3 Document Processing 20/20 82.6
C1 Content Production 19/20 82.6
C2 Content Production 18/20 81.5
C3 Content Production 20/20 81
C4 Content Production 16/20 81.1

Category mean Dev/Ops Scripting 14.5Document Processing 20Content Production 18.3

Agentic tool-calling & protocol adherence 1b 151.2 / 160 18.9 8/8 n=1 84.8 run 2026-07-28
Test Category Score tok/s
WA1 Tool Calling & Protocol Adherence 20/20 36.2
WA2 Tool Calling & Protocol Adherence 20/20 73.2
WA3 Tool Calling & Protocol Adherence 20/20 61.5
WA4 Tool Calling & Protocol Adherence 20/20 65.1
WA5 Tool Calling & Protocol Adherence 20/20 76.9
WB1 Agent Robustness & State Management 20/20 83.2
WB2 Agent Robustness & State Management 20/20 82.6
WB3 Agent Robustness & State Management 11/20 81.6

Category mean Tool Calling & Protocol Adherence 20Agent Robustness & State Management 17

Coding depth 1c 153.9 / 180 17.1 9/9 n=1 86.1 run 2026-09-16
Test Category Score tok/s
K1 Coding — Concurrency 20/20 84.1
K2 Coding — Algorithms 20/20 84.9
K3 Coding — Security Review 14/20 84.4
K4 Coding — Performance 19/20 84.8
K5 Coding — Refactoring Under Constraints 20/20 84.5
K6 Coding — Test Writing 16/20 84.8
K7 Coding — React Component (typed) 12/20 84.9
K8 Coding — TypeScript Type System 14/20 84.5
K9 Coding — React Debugging 19/20 84.3

Category mean Coding — Concurrency 20Coding — Algorithms 20Coding — Security Review 14Coding — Performance 19Coding — Refactoring Under Constraints 20Coding — Test Writing 16Coding — React Component (typed) 12Coding — TypeScript Type System 14Coding — React Debugging 19

Doc/OCR vision 1d 150.4 / 160 18.8 8/8 n=1 78.3 run 2026-09-20
Test Category Score tok/s
D1 Doc/OCR — Transcription 20/20 61.6
D2 Doc/OCR — Document QA 20/20 81.5
D3 Doc/OCR — Field Extraction 20/20 70.5
D4 Doc/OCR — Document QA 20/20 70.7
D5 Doc/OCR — Transcription 10/20 71.4
D6 Doc/OCR — Field Extraction 20/20 71
D7 Doc/OCR — Degraded Input 20/20 74.7
D8 Doc/OCR — Degraded Input 20/20 83

Category mean Doc/OCR — Transcription 15Doc/OCR — Document QA 20Doc/OCR — Field Extraction 20Doc/OCR — Degraded Input 20

Doc/OCR — real-degraded tier 1d2 47.2 / 80 11.8 4/4 n=1 82.1 run 2026-09-20
Test Category Score tok/s
D9 Doc/OCR — Real Degraded Transcription 6/20 69.4
D10 Doc/OCR — Handwriting Extraction 17/20 79
D11 Doc/OCR — Degraded Table Transcription 12/20 81.9
D12 Doc/OCR — Degraded Table QA 12/20 81

Category mean Doc/OCR — Real Degraded Transcription 6Doc/OCR — Handwriting Extraction 17Doc/OCR — Degraded Table Transcription 12Doc/OCR — Degraded Table QA 12

Content-production depth 1e 106.2 / 120 17.7 6/6 n=1 82.7 run 2026-09-16
Test Category Score tok/s
E1 SOW & Requirements Decomposition 18/20 79.3
E2 SOW & Requirements Decomposition 20/20 79.9
E3 Executive & Proposal Content 20/20 80
E4 Executive & Proposal Content 14/20 82.2
E5 Executive & Proposal Content 20/20 81.5
E6 Executive & Proposal Content 14/20 83.5

Category mean SOW & Requirements Decomposition 19Executive & Proposal Content 17

Long-context retrieval & synthesis 1g 120 / 120 20 6/6 n=1 70.4 run 2026-09-30
Test Category Score tok/s
L1 Long-Context — Needle Retrieval 20/20 59.1
L2 Long-Context — Grounded QA with Distractors 20/20 66.1
L3 Long-Context — Ops Log Reasoning 20/20 55.1
L4 Long-Context — Faithful Summarization 20/20 59.8
L5 Long-Context — Instruction Retention 20/20 53.1
L6 Long-Context — Full-Haystack QA 20/20 35.9

Category mean Long-Context — Needle Retrieval 20Long-Context — Grounded QA with Distractors 20Long-Context — Ops Log Reasoning 20Long-Context — Faithful Summarization 20Long-Context — Instruction Retention 20Long-Context — Full-Haystack QA 20

Long-context multi-needle (MRCR) 1g2 30.9 / 60 10.3 3/3 n=1 51.7 run 2026-09-30
Test Category Score tok/s
L7 Long-Context — Multi-Needle Retrieval (MRCR) 20/20 47.5
L8 Long-Context — Multi-Needle Retrieval (MRCR) 7/20 31.7
L9 Long-Context — Multi-Needle Retrieval (MRCR) 4/20 30.1
Live one-shot builds (runtime-verified) 1h 59.2 / 80 14.8 4/4 n=1 78.3 run 2026-09-09
Test Category Score tok/s
H1 Web-Estate One-Shot Builds 17/20 78.1
H2 Web-Estate One-Shot Builds 14/20 81.8
H3 Web-Estate One-Shot Builds 13/20 78.4
H4 Web-Estate One-Shot Builds 15/20 75
Brand-constrained static build (negative-constraint adherence) 1j 83 / 100 16.6 5/5 n=1 85.7 run 2026-09-16
Test Category Score tok/s
J1 Brand-Constrained Static Build 17/20 84.2
J3 Brand-Constrained Static Build 18/20 84.5
J4 Brand-Constrained Static Build 20/20 84.1
J5 Brand-Constrained Static Build 15/20 84.5
J6 Brand-Constrained Static Build 13/20 84.7

Each test is LLM-judged 0–20 against a fixed rubric; suite max = tests × 20. An ADJ chip marks a cell carrying a subjective adjustment from the maintainer; the judged score is the one shown, is unchanged, and remains what the axes, ranks and tie bands are computed from. Hover for the reason. Rows with a ▶ expand to the per-test breakdown. Where fewer tests ran than the suite total, unrun tests count as zero toward the axis score. See methodology.