← Leaderboard

Gemma-4 26B-A4B 8bit MLX

Daily driver RAN THE GAUNTLET Rank #2 of 51 · 8/8 GAUNTLET progress
Generalist — 95.5 / 100 G Agentic — 99.5 / 100 A Understanding — 81 / 100 U Needle — 77.8 / 100 N Thinking — 88 / 100 T Live — 68 / 100 L Engineering — 98.5 / 100 E Throughput — 19.8 / 100 T
G
95.5
A
99.5
U
81
N
77.8
T
88
L
68
E
98.5
T
19.8

Specification

Parameters
26B total / 4B active per token
Architecture
gemma4
Size on disk
28 GB
Quantization
8bit
Format
MLX
Runtime
LM Studio
Class
Standard
Reasoning (CoT)
Yes — emits reasoning tokens
Internal ID
M2
Mean speed
69.6 tok/s across suites
Stall census
10 stalls in 83 observed tests (12.0%)
Reasoning appetite
3,860 tokens mean · 16,381 max
Model card
lmstudio.ai

Suite results

Suite Score Avg / 20 Tests n tok/s Run
General capability (13-task real-workload suite) 1a 248.3 / 260 19.1 13/13 n=4.9 47.1 run 2026-09-02
Test Category Score tok/s
A1 Dev/Ops Scripting 20/20 86.6
A2 Dev/Ops Scripting 18/20 85.8
A3 Dev/Ops Scripting 18/20 74.7
A4 Dev/Ops Scripting 18/20 72.4
A5 Dev/Ops Scripting 20/20 72.7
A6 Dev/Ops Scripting 20/20 73.3
B1 Document Processing 20/20 75
B2 Document Processing 20/20 77.3
B3 Document Processing 20/20 77.6
C1 Content Production 18/20 77.6
C2 Content Production 18/20 76.7
C3 Content Production 20/20 75.8
C4 Content Production 19/20 74.6

Category mean Dev/Ops Scripting 19Document Processing 20Content Production 18.8

Agentic tool-calling & protocol adherence 1b 159.2 / 160 19.9 8/8 n=4 83.9 run 2026-09-12
Test Category Score tok/s
WA1 Tool Calling & Protocol Adherence 20/20 55.7
WA2 Tool Calling & Protocol Adherence 20/20 74.2
WA3 Tool Calling & Protocol Adherence 20/20 69.1
WA4 Tool Calling & Protocol Adherence 20/20 65.8
WA5 Tool Calling & Protocol Adherence 20/20 71.6
WB1 Agent Robustness & State Management 20/20 87.2
WB2 Agent Robustness & State Management 20/20 85.7
WB3 Agent Robustness & State Management 19/20 85.9

Category mean Tool Calling & Protocol Adherence 20Agent Robustness & State Management 19.7

Coding depth 1c 177.3 / 180 19.7 9/9 n=1 72.4 run 2026-09-15
Test Category Score tok/s
K1 Coding — Concurrency 20/20 77.2
K2 Coding — Algorithms 20/20 74.9
K3 Coding — Security Review 20/20 69.2
K4 Coding — Performance 18/20 63.5
K5 Coding — Refactoring Under Constraints 20/20 57.8
K6 Coding — Test Writing 20/20 54.3
K7 Coding — React Component (typed) 20/20 80.3
K8 Coding — TypeScript Type System 19/20 82.7
K9 Coding — React Debugging 20/20 78.9

Category mean Coding — Concurrency 20Coding — Algorithms 20Coding — Security Review 20Coding — Performance 18Coding — Refactoring Under Constraints 20Coding — Test Writing 20Coding — React Component (typed) 20Coding — TypeScript Type System 19Coding — React Debugging 20

Doc/OCR vision 1d 158.4 / 160 19.8 8/8 n=1.1 82 run 2026-08-06
Test Category Score tok/s
D1 Doc/OCR — Transcription 20/20 83.2
D2 Doc/OCR — Document QA 20/20 82.3
D3 Doc/OCR — Field Extraction 20/20 78.9
D4 Doc/OCR — Document QA 20/20 80.8
D5 Doc/OCR — Transcription 18/20 80.2
D6 Doc/OCR — Field Extraction 20/20 71.4
D7 Doc/OCR — Degraded Input 20/20 74.3
D8 Doc/OCR — Degraded Input 20/20 72.5

Category mean Doc/OCR — Transcription 19Doc/OCR — Document QA 20Doc/OCR — Field Extraction 20Doc/OCR — Degraded Input 20

Doc/OCR — real-degraded tier 1d2 36 / 80 9 4/4 n=1 68.9 run 2026-08-14
Test Category Score tok/s
D9 Doc/OCR — Real Degraded Transcription 8/20 82.6
D10 Doc/OCR — Handwriting Extraction 0/20 62.1
D11 Doc/OCR — Degraded Table Transcription 13/20 61.7
D12 Doc/OCR — Degraded Table QA 15/20 65.8

Category mean Doc/OCR — Real Degraded Transcription 8Doc/OCR — Handwriting Extraction 0Doc/OCR — Degraded Table Transcription 13Doc/OCR — Degraded Table QA 15

Content-production depth 1e 114 / 120 19 6/6 n=1.2 81.5 run 2026-08-02
Test Category Score tok/s
E1 SOW & Requirements Decomposition 20/20 84.4
E2 SOW & Requirements Decomposition 20/20 82.7
E3 Executive & Proposal Content 20/20 78.4
E4 Executive & Proposal Content 15/20 83
E5 Executive & Proposal Content 19/20 79.2
E6 Executive & Proposal Content 20/20 74.1

Category mean SOW & Requirements Decomposition 20Executive & Proposal Content 18.5

Long-context retrieval & synthesis 1g 120 / 120 20 6/6 n=1.2 71 run 2026-08-06
Test Category Score tok/s
L1 Long-Context — Needle Retrieval 20/20 60.9
L2 Long-Context — Grounded QA with Distractors 20/20 65.4
L3 Long-Context — Ops Log Reasoning 20/20 48.8
L4 Long-Context — Faithful Summarization 20/20 54.6
L5 Long-Context — Instruction Retention 20/20 57.1
L6 Long-Context — Full-Haystack QA 20/20 33.9

Category mean Long-Context — Needle Retrieval 20Long-Context — Grounded QA with Distractors 20Long-Context — Ops Log Reasoning 20Long-Context — Faithful Summarization 20Long-Context — Instruction Retention 20Long-Context — Full-Haystack QA 20

Long-context multi-needle (MRCR) 1g2 20.1 / 60 6.7 3/3 n=1 50.5 run 2026-08-14
Test Category Score tok/s
L7 Long-Context — Multi-Needle Retrieval (MRCR) 20/20 61
L8 Long-Context — Multi-Needle Retrieval (MRCR) 0/20 41.5
L9 Long-Context — Multi-Needle Retrieval (MRCR) 0/20 36
Live one-shot builds (runtime-verified) 1h 68 / 80 17 4/4 n=1 79.8 run 2026-08-15
Test Category Score tok/s
H1 Web-Estate One-Shot Builds 15/20 73.5
H2 Web-Estate One-Shot Builds 19/20 78.2
H3 Web-Estate One-Shot Builds 17/20 79.4
H4 Web-Estate One-Shot Builds 17/20 74.5
Production replay (real agent workload) 1i 141.6 / 240 11.8 12/12 n=1 56.6 run 2026-08-16
Test Category Score tok/s
I1 Nightly Fix Replay 0/20 59.2
I2 Nightly Fix Replay 0/20 47.5
I3 Nightly Fix Replay 8/20 36.7
I4 Nightly Fix Replay 15/20 51.7
I5 Nightly Fix Replay 13/20 51.1
I6 Nightly Fix Replay 0/20 50.9
I7 Audit Judgment 15/20 59.3
I8 Audit Judgment 20/20 65.2
I9 Cron Dependency Reasoning 19/20 56.7
I10 Cron Dependency Reasoning 15/20 56.8
I11 Curation Judgment 20/20 62.5
I12 Curation Judgment 17/20 61

Category mean Nightly Fix Replay 6Audit Judgment 17.5Cron Dependency Reasoning 17Curation Judgment 18.5

Brand-constrained static build (negative-constraint adherence) 1j 76 / 100 15.2 5/5 n=1 72.2 run 2026-09-13
Test Category Score tok/s
J1 Brand-Constrained Static Build 12/20 64.3
J3 Brand-Constrained Static Build 20/20 68.6
J4 Brand-Constrained Static Build 20/20 69.8
J5 Brand-Constrained Static Build 10/20 77
J6 Brand-Constrained Static Build 14/20 78.4

Each test is LLM-judged 0–20 against a fixed rubric; suite max = tests × 20. An ADJ chip marks a cell carrying a subjective adjustment from the maintainer; the judged score is the one shown, is unchanged, and remains what the axes, ranks and tie bands are computed from. Hover for the reason. Rows with a ▶ expand to the per-test breakdown. Where fewer tests ran than the suite total, unrun tests count as zero toward the axis score. See methodology.