← Leaderboard

Gemma-4 26B-A4B 8bit MLX

Daily driver RAN THE GAUNTLET Rank #2 of 33 · 8/8 GAUNTLET progress
Generalist — 95.5 / 100 G Agentic — 98 / 100 A Understanding — 79.7 / 100 U Needle — 77.8 / 100 N Thinking — 88 / 100 T Live — 55.5 / 100 L Engineering — 91 / 100 E Throughput — 40.6 / 100 T
G
95.5
A
98
U
79.7
N
77.8
T
88
L
55.5
E
91
T
40.6

Specification

Parameters
26B total / 4B active per token
Architecture
gemma4
Size on disk
28 GB
Quantization
8bit
Format
MLX
Reasoning (CoT)
No
Internal ID
M2
Mean speed
66.2 tok/s across suites
Stall census
10 stalls in 83 observed tests (12.0%)
Reasoning appetite
3,860 tokens mean · 16,381 max
Model card
lmstudio.ai

Suite results

Suite Score Avg / 20 Tests tok/s
General capability (13-task real-workload suite) 1a 248.3 / 260 19.1 13/13 78.1
Agentic tool-calling & protocol adherence 1b 156.8 / 160 19.6 8/8 70.5
Coding depth 1c 109.2 / 120 18.2 6/6 66.1
Doc/OCR vision 1d 156 / 160 19.5 8/8 78
Doc/OCR — real-degraded tier 1d2 35.2 / 80 8.8 4/4 68
Content-production depth 1e 111 / 120 18.5 6/6 80.9
Long-context retrieval & synthesis 1g 120 / 120 20 6/6 53.4
Long-context multi-needle (MRCR) 1g2 20.1 / 60 6.7 3/3 46.2
Live one-shot builds (runtime-verified) 1h 62.4 / 80 15.6 4/4
Production replay (real agent workload) 1i 115.2 / 240 9.6 12/12 54.9

Each test is LLM-judged 0–20 against a fixed rubric; suite max = tests × 20. Where fewer tests ran than the suite total, unrun tests count as zero toward the axis score. See methodology.