← Leaderboard

Gemma-4 26B-A4B MLX (Ollama, MBP)

Tested Rank #16 of 51 · 7/8 GAUNTLET progress
Generalist — 93.5 / 100 G Agentic — 87.5 / 100 A Understanding — 73 / 100 U Needle — 77.5 / 100 N Thinking — pending T Live — 49.7 / 100 L Engineering — 73.5 / 100 E Throughput — 34.4 / 100 T

Hollow markers & dashed spokes: axis not yet scored

G
93.5
A
87.5
U
73
N
77.5
T
n/d
L
49.7
E
73.5
T
34.4

Runtime cohort \u2014 read this before comparing

This model was benchmarked on Ollama. Most of the board was benchmarked on LM Studio. The two are not interchangeable and the scores are published side by side deliberately rather than merged.

Why both exist: on the benchmark machine, LM Studio drops generation streams on a minority of attempts overall (15.9% of tests first-attempt across 1,526 tests) but heavily on specific models \u2014 several exceed 45%, two exceed 90%. A dropped stream is indistinguishable from a bad answer in a tally, so unattended benchmarking moved to Ollama, which has recorded 0 stream deaths in 177 tests.

What the offset looks like: across every suite where a byte-identical pair has been scored on both runtimes — 23 paired observations over 4 pairs — Ollama runs 0.1 points higher on average, with mean |Δ| 0.6 and sd 1.2. The single largest gap between two byte-identical arms is 4 points — bigger than most of the gaps at the top of the leaderboard. Read this as the noise floor of the whole benchmark rather than as a conversion factor, and treat cross-runtime gaps under 1 point as noise.

Specification

Parameters
26B total / 4B active per token
Architecture
gemma4
Size on disk
17.55 GB
Quantization
Format
MLX
Runtime
Ollama — scores are NOT directly comparable to the LM Studio cohort; see the runtime note below
Class
Standard
Reasoning (CoT)
Yes — emits reasoning tokens
Internal ID
MO1
Mean speed
120.8 tok/s across suites
Model card
ollama.com

Suite results

Suite Score Avg / 20 Tests n tok/s Run
General capability (13-task real-workload suite) 1a 243.1 / 260 18.7 13/13 n=2.1 103.7 run 2026-09-25
Test Category Score tok/s
A1 Dev/Ops Scripting 18.5/20 125.4
A2 Dev/Ops Scripting 18.5/20 115.2
A3 Dev/Ops Scripting 17/20 103.3
A4 Dev/Ops Scripting 19/20 105.2
A5 Dev/Ops Scripting 20/20 95.7
A6 Dev/Ops Scripting 20/20 101.5
B1 Document Processing 20/20 98.8
B2 Document Processing 20/20 104
B3 Document Processing 20/20 91.4
C1 Content Production 18/20 81
C2 Content Production 18/20 81.8
C3 Content Production 20/20 86.7
C4 Content Production 14.5/20 88.6

Category mean Dev/Ops Scripting 18.8Document Processing 20Content Production 17.6

Agentic tool-calling & protocol adherence 1b 140 / 160 17.5 8/8 n=2 146.8 run 2026-09-25
Test Category Score tok/s
WA1 Tool Calling & Protocol Adherence 20/20 129.6
WA2 Tool Calling & Protocol Adherence 20/20 157.2
WA3 Tool Calling & Protocol Adherence 20/20 151.3
WA4 Tool Calling & Protocol Adherence 20/20 159.6
WA5 Tool Calling & Protocol Adherence 20/20 132.3
WB1 Agent Robustness & State Management 20/20 126.7
WB2 Agent Robustness & State Management 20/20 138.7

Category mean Tool Calling & Protocol Adherence 20Agent Robustness & State Management 20

Coding depth 1c 132.3 / 180 14.7 9/9 n=1.9 117.9 run 2026-09-13
Test Category Score tok/s
K1 Coding — Concurrency 20/20 111
K3 Coding — Security Review 17/20 95.4
K4 Coding — Performance 19/20 104
K5 Coding — Refactoring Under Constraints 20/20 106.4
K6 Coding — Test Writing 17/20 115.2
K8 Coding — TypeScript Type System 19/20 98.6
K9 Coding — React Debugging 20/20 99.5

Category mean Coding — Concurrency 20Coding — Security Review 17Coding — Performance 19Coding — Refactoring Under Constraints 20Coding — Test Writing 17Coding — TypeScript Type System 19Coding — React Debugging 20

Doc/OCR vision 1d 158.4 / 160 19.8 8/8 n=2 164.2 run 2026-09-23
Test Category Score tok/s
D1 Doc/OCR — Transcription 20/20 143
D2 Doc/OCR — Document QA 20/20 137.4
D3 Doc/OCR — Field Extraction 20/20 165.2
D4 Doc/OCR — Document QA 19/20 155.6
D5 Doc/OCR — Transcription 19/20 150.9
D6 Doc/OCR — Field Extraction 20/20 165.7
D7 Doc/OCR — Degraded Input 20/20 157
D8 Doc/OCR — Degraded Input 20/20 158.4

Category mean Doc/OCR — Transcription 19.5Doc/OCR — Document QA 19.5Doc/OCR — Field Extraction 20Doc/OCR — Degraded Input 20

Doc/OCR — real-degraded tier 1d2 16.8 / 80 4.2 4/4 n=2 140 run 2026-09-27
Test Category Score tok/s
D12 Doc/OCR — Degraded Table QA 17/20 108.8
Content-production depth 1e 88.2 / 120 14.7 6/6 n=2 106.5 run 2026-09-25
Test Category Score tok/s
E1 SOW & Requirements Decomposition 18/20 108.2
E3 Executive & Proposal Content 17.5/20 93.1
E4 Executive & Proposal Content 16/20 93.8
E5 Executive & Proposal Content 17.5/20 94.9
E6 Executive & Proposal Content 19/20 105.6

Category mean SOW & Requirements Decomposition 18Executive & Proposal Content 17.5

Long-context retrieval & synthesis 1g 119.4 / 120 19.9 6/6 n=2.8 115.4 run 2026-09-27
Test Category Score tok/s
L1 Long-Context — Needle Retrieval 20/20 133.5
L2 Long-Context — Grounded QA with Distractors 20/20 115.6
L3 Long-Context — Ops Log Reasoning 20/20 109.2
L4 Long-Context — Faithful Summarization 20/20 105.3
L5 Long-Context — Instruction Retention 20/20 120.3
L6 Long-Context — Full-Haystack QA 20/20 108.6

Category mean Long-Context — Needle Retrieval 20Long-Context — Grounded QA with Distractors 20Long-Context — Ops Log Reasoning 20Long-Context — Faithful Summarization 20Long-Context — Instruction Retention 20Long-Context — Full-Haystack QA 20

Long-context multi-needle (MRCR) 1g2 20.1 / 60 6.7 3/3 n=2 96.8 run 2026-09-25
Test Category Score tok/s
L7 Long-Context — Multi-Needle Retrieval (MRCR) 20/20 100.6
Live one-shot builds (runtime-verified) 1h 66.4 / 80 16.6 4/4 n=1.5 115.8 run 2026-09-25
Test Category Score tok/s
H1 Web-Estate One-Shot Builds 16/20 121.8
H2 Web-Estate One-Shot Builds 15/20 112.6
H3 Web-Estate One-Shot Builds 15.5/20 110.8
H4 Web-Estate One-Shot Builds 20/20 122
Production replay (real agent workload) 1i 88.8 / 240 7.4 12/12 n=1 103.2 run 2026-09-26
Test Category Score tok/s
I7 Audit Judgment 15/20 105.6
I8 Audit Judgment 18/20 96.5
I9 Cron Dependency Reasoning 18/20 93.9
I11 Curation Judgment 20/20 93.7
I12 Curation Judgment 18/20 95.4

Category mean Audit Judgment 16.5Cron Dependency Reasoning 18Curation Judgment 19

Brand-constrained static build (negative-constraint adherence) 1j 53.5 / 100 10.7 5/5 n=1.6 118.7 run 2026-09-25
Test Category Score tok/s
J1 Brand-Constrained Static Build 13.5/20 128.6
J3 Brand-Constrained Static Build 18/20 119.9
J5 Brand-Constrained Static Build 9/20 106.9
J6 Brand-Constrained Static Build 13/20 114.6

Each test is LLM-judged 0–20 against a fixed rubric; suite max = tests × 20. An ADJ chip marks a cell carrying a subjective adjustment from the maintainer; the judged score is the one shown, is unchanged, and remains what the axes, ranks and tie bands are computed from. Hover for the reason. Rows with a ▶ expand to the per-test breakdown. Where fewer tests ran than the suite total, unrun tests count as zero toward the axis score. See methodology.