Hollow markers & dashed spokes: axis not yet scored
Runtime cohort \u2014 read this before comparing
This model was benchmarked on Ollama. Most of the board was benchmarked on LM Studio. The two are not interchangeable and the scores are published side by side deliberately rather than merged.
Why both exist: on the benchmark machine, LM Studio drops generation streams on a minority of attempts overall (15.9% of tests first-attempt across 1,526 tests) but heavily on specific models \u2014 several exceed 45%, two exceed 90%. A dropped stream is indistinguishable from a bad answer in a tally, so unattended benchmarking moved to Ollama, which has recorded 0 stream deaths in 177 tests.
What the offset looks like: across every suite where a byte-identical pair has been scored on both runtimes — 23 paired observations over 4 pairs — Ollama runs 0.1 points higher on average, with mean |Δ| 0.6 and sd 1.2. The single largest gap between two byte-identical arms is 4 points — bigger than most of the gaps at the top of the leaderboard. Read this as the noise floor of the whole benchmark rather than as a conversion factor, and treat cross-runtime gaps under 1 point as noise.
Specification
- Parameters
- 26B total / 4B active per token
- Architecture
- gemma4
- Size on disk
- 17.55 GB
- Quantization
- Format
- MLX
- Runtime
- Ollama — scores are NOT directly comparable to the LM Studio cohort; see the runtime note below
- Class
- Standard
- Reasoning (CoT)
- Yes — emits reasoning tokens
- Internal ID
- MO1
- Mean speed
- 120.8 tok/s across suites
- Model card
- ollama.com
Suite results
General capability (13-task real-workload suite) 1a 243.1 / 260 18.7 13/13 n=2.1 103.7 run 2026-09-25
| Test | Category | Score | tok/s | |
|---|---|---|---|---|
| A1 | Dev/Ops Scripting | 18.5/20 | 125.4 | |
| A2 | Dev/Ops Scripting | 18.5/20 | 115.2 | |
| A3 | Dev/Ops Scripting | 17/20 | 103.3 | |
| A4 | Dev/Ops Scripting | 19/20 | 105.2 | |
| A5 | Dev/Ops Scripting | 20/20 | 95.7 | |
| A6 | Dev/Ops Scripting | 20/20 | 101.5 | |
| B1 | Document Processing | 20/20 | 98.8 | |
| B2 | Document Processing | 20/20 | 104 | |
| B3 | Document Processing | 20/20 | 91.4 | |
| C1 | Content Production | 18/20 | 81 | |
| C2 | Content Production | 18/20 | 81.8 | |
| C3 | Content Production | 20/20 | 86.7 | |
| C4 | Content Production | 14.5/20 | 88.6 |
Category mean Dev/Ops Scripting 18.8Document Processing 20Content Production 17.6
Agentic tool-calling & protocol adherence 1b 140 / 160 17.5 8/8 n=2 146.8 run 2026-09-25
| Test | Category | Score | tok/s | |
|---|---|---|---|---|
| WA1 | Tool Calling & Protocol Adherence | 20/20 | 129.6 | |
| WA2 | Tool Calling & Protocol Adherence | 20/20 | 157.2 | |
| WA3 | Tool Calling & Protocol Adherence | 20/20 | 151.3 | |
| WA4 | Tool Calling & Protocol Adherence | 20/20 | 159.6 | |
| WA5 | Tool Calling & Protocol Adherence | 20/20 | 132.3 | |
| WB1 | Agent Robustness & State Management | 20/20 | 126.7 | |
| WB2 | Agent Robustness & State Management | 20/20 | 138.7 |
Category mean Tool Calling & Protocol Adherence 20Agent Robustness & State Management 20
Coding depth 1c 132.3 / 180 14.7 9/9 n=1.9 117.9 run 2026-09-13
| Test | Category | Score | tok/s | |
|---|---|---|---|---|
| K1 | Coding — Concurrency | 20/20 | 111 | |
| K3 | Coding — Security Review | 17/20 | 95.4 | |
| K4 | Coding — Performance | 19/20 | 104 | |
| K5 | Coding — Refactoring Under Constraints | 20/20 | 106.4 | |
| K6 | Coding — Test Writing | 17/20 | 115.2 | |
| K8 | Coding — TypeScript Type System | 19/20 | 98.6 | |
| K9 | Coding — React Debugging | 20/20 | 99.5 |
Category mean Coding — Concurrency 20Coding — Security Review 17Coding — Performance 19Coding — Refactoring Under Constraints 20Coding — Test Writing 17Coding — TypeScript Type System 19Coding — React Debugging 20
Doc/OCR vision 1d 158.4 / 160 19.8 8/8 n=2 164.2 run 2026-09-23
| Test | Category | Score | tok/s | |
|---|---|---|---|---|
| D1 | Doc/OCR — Transcription | 20/20 | 143 | |
| D2 | Doc/OCR — Document QA | 20/20 | 137.4 | |
| D3 | Doc/OCR — Field Extraction | 20/20 | 165.2 | |
| D4 | Doc/OCR — Document QA | 19/20 | 155.6 | |
| D5 | Doc/OCR — Transcription | 19/20 | 150.9 | |
| D6 | Doc/OCR — Field Extraction | 20/20 | 165.7 | |
| D7 | Doc/OCR — Degraded Input | 20/20 | 157 | |
| D8 | Doc/OCR — Degraded Input | 20/20 | 158.4 |
Category mean Doc/OCR — Transcription 19.5Doc/OCR — Document QA 19.5Doc/OCR — Field Extraction 20Doc/OCR — Degraded Input 20
Doc/OCR — real-degraded tier 1d2 16.8 / 80 4.2 4/4 n=2 140 run 2026-09-27
| Test | Category | Score | tok/s | |
|---|---|---|---|---|
| D12 | Doc/OCR — Degraded Table QA | 17/20 | 108.8 |
Content-production depth 1e 88.2 / 120 14.7 6/6 n=2 106.5 run 2026-09-25
| Test | Category | Score | tok/s | |
|---|---|---|---|---|
| E1 | SOW & Requirements Decomposition | 18/20 | 108.2 | |
| E3 | Executive & Proposal Content | 17.5/20 | 93.1 | |
| E4 | Executive & Proposal Content | 16/20 | 93.8 | |
| E5 | Executive & Proposal Content | 17.5/20 | 94.9 | |
| E6 | Executive & Proposal Content | 19/20 | 105.6 |
Category mean SOW & Requirements Decomposition 18Executive & Proposal Content 17.5
Long-context retrieval & synthesis 1g 119.4 / 120 19.9 6/6 n=2.8 115.4 run 2026-09-27
| Test | Category | Score | tok/s | |
|---|---|---|---|---|
| L1 | Long-Context — Needle Retrieval | 20/20 | 133.5 | |
| L2 | Long-Context — Grounded QA with Distractors | 20/20 | 115.6 | |
| L3 | Long-Context — Ops Log Reasoning | 20/20 | 109.2 | |
| L4 | Long-Context — Faithful Summarization | 20/20 | 105.3 | |
| L5 | Long-Context — Instruction Retention | 20/20 | 120.3 | |
| L6 | Long-Context — Full-Haystack QA | 20/20 | 108.6 |
Category mean Long-Context — Needle Retrieval 20Long-Context — Grounded QA with Distractors 20Long-Context — Ops Log Reasoning 20Long-Context — Faithful Summarization 20Long-Context — Instruction Retention 20Long-Context — Full-Haystack QA 20
Long-context multi-needle (MRCR) 1g2 20.1 / 60 6.7 3/3 n=2 96.8 run 2026-09-25
| Test | Category | Score | tok/s | |
|---|---|---|---|---|
| L7 | Long-Context — Multi-Needle Retrieval (MRCR) | 20/20 | 100.6 |
Live one-shot builds (runtime-verified) 1h 66.4 / 80 16.6 4/4 n=1.5 115.8 run 2026-09-25
| Test | Category | Score | tok/s | |
|---|---|---|---|---|
| H1 | Web-Estate One-Shot Builds | 16/20 | 121.8 | |
| H2 | Web-Estate One-Shot Builds | 15/20 | 112.6 | |
| H3 | Web-Estate One-Shot Builds | 15.5/20 | 110.8 | |
| H4 | Web-Estate One-Shot Builds | 20/20 | 122 |
Production replay (real agent workload) 1i 88.8 / 240 7.4 12/12 n=1 103.2 run 2026-09-26
| Test | Category | Score | tok/s | |
|---|---|---|---|---|
| I7 | Audit Judgment | 15/20 | 105.6 | |
| I8 | Audit Judgment | 18/20 | 96.5 | |
| I9 | Cron Dependency Reasoning | 18/20 | 93.9 | |
| I11 | Curation Judgment | 20/20 | 93.7 | |
| I12 | Curation Judgment | 18/20 | 95.4 |
Category mean Audit Judgment 16.5Cron Dependency Reasoning 18Curation Judgment 19
Brand-constrained static build (negative-constraint adherence) 1j 53.5 / 100 10.7 5/5 n=1.6 118.7 run 2026-09-25
| Test | Category | Score | tok/s | |
|---|---|---|---|---|
| J1 | Brand-Constrained Static Build | 13.5/20 | 128.6 | |
| J3 | Brand-Constrained Static Build | 18/20 | 119.9 | |
| J5 | Brand-Constrained Static Build | 9/20 | 106.9 | |
| J6 | Brand-Constrained Static Build | 13/20 | 114.6 |
Each test is LLM-judged 0–20 against a fixed rubric; suite max = tests × 20. An ADJ chip marks a cell carrying a subjective adjustment from the maintainer; the judged score is the one shown, is unchanged, and remains what the axes, ranks and tie bands are computed from. Hover for the reason. Rows with a ▶ expand to the per-test breakdown. Where fewer tests ran than the suite total, unrun tests count as zero toward the axis score. See methodology.