← Leaderboard

Qwen3.8-27B nvfp4 MLX (Ollama, MBP)

Tested Rank #15 of 51 · 7/8 GAUNTLET progress
Generalist — 95 / 100 G Agentic — 100 / 100 A Understanding — 93 / 100 U Needle — 89.8 / 100 N Thinking — pending T Live — 48.6 / 100 L Engineering — 97 / 100 E Throughput — 13.7 / 100 T

Hollow markers & dashed spokes: axis not yet scored

G
95
A
100
U
93
N
89.8
T
n/d
L
48.6
E
97
T
13.7

Runtime cohort \u2014 read this before comparing

This model was benchmarked on Ollama. Most of the board was benchmarked on LM Studio. The two are not interchangeable and the scores are published side by side deliberately rather than merged.

Why both exist: on the benchmark machine, LM Studio drops generation streams on a minority of attempts overall (15.9% of tests first-attempt across 1,526 tests) but heavily on specific models \u2014 several exceed 45%, two exceed 90%. A dropped stream is indistinguishable from a bad answer in a tally, so unattended benchmarking moved to Ollama, which has recorded 0 stream deaths in 177 tests.

What the offset looks like: across every suite where a byte-identical pair has been scored on both runtimes — 23 paired observations over 4 pairs — Ollama runs 0.1 points higher on average, with mean |Δ| 0.6 and sd 1.2. The single largest gap between two byte-identical arms is 4 points — bigger than most of the gaps at the top of the leaderboard. Read this as the noise floor of the whole benchmark rather than as a conversion factor, and treat cross-runtime gaps under 1 point as noise.

Specification

Parameters
27B dense
Architecture
qwen3_5
Size on disk
16.5 GB
Quantization
nvfp4
Format
MLX
Runtime
Ollama — scores are NOT directly comparable to the LM Studio cohort; see the runtime note below
Class
Standard
Reasoning (CoT)
Yes — emits reasoning tokens
Internal ID
MO2
Mean speed
48.3 tok/s across suites
Model card
ollama.com

Suite results

Suite Score Avg / 20 Tests n tok/s Run
General capability (13-task real-workload suite) 1a 247 / 260 19 13/13 n=1.9 36.8 run 2026-09-26
Test Category Score tok/s
A1 Dev/Ops Scripting 20/20 51
A2 Dev/Ops Scripting 20/20 47.5
A3 Dev/Ops Scripting 12/20 35.9
A4 Dev/Ops Scripting 20/20 24.3
A5 Dev/Ops Scripting 20/20 28.6
A6 Dev/Ops Scripting 20/20 30.2
B1 Document Processing 20/20 40.4
B2 Document Processing 20/20 35
B3 Document Processing 20/20 32.4
C1 Content Production 20/20 31.1
C2 Content Production 17/20 38.1
C3 Content Production 19/20 38.5
C4 Content Production 19/20 46

Category mean Dev/Ops Scripting 18.7Document Processing 20Content Production 18.8

Agentic tool-calling & protocol adherence 1b 160 / 160 20 8/8 n=2 60.2 run 2026-09-26
Test Category Score tok/s
WA1 Tool Calling & Protocol Adherence 20/20 75.1
WA2 Tool Calling & Protocol Adherence 20/20 60.2
WA3 Tool Calling & Protocol Adherence 20/20 67.9
WA4 Tool Calling & Protocol Adherence 20/20 72.7
WA5 Tool Calling & Protocol Adherence 20/20 48.5
WB1 Agent Robustness & State Management 20/20 47.6
WB2 Agent Robustness & State Management 20/20 58.3
WB3 Agent Robustness & State Management 20/20 51.5

Category mean Tool Calling & Protocol Adherence 20Agent Robustness & State Management 20

Coding depth 1c 174.6 / 180 19.4 9/9 n=2 45.3 run 2026-09-13
Test Category Score tok/s
K1 Coding — Concurrency 20/20 36.7
K2 Coding — Algorithms 20/20 40.1
K3 Coding — Security Review 16/20 38.5
K4 Coding — Performance 20/20 37.9
K5 Coding — Refactoring Under Constraints 20/20 39.2
K6 Coding — Test Writing 20/20 46.6
K7 Coding — React Component (typed) 20/20 46.8
K8 Coding — TypeScript Type System 20/20 48.4
K9 Coding — React Debugging 20/20 45.3

Category mean Coding — Concurrency 20Coding — Algorithms 20Coding — Security Review 16Coding — Performance 20Coding — Refactoring Under Constraints 20Coding — Test Writing 20Coding — React Component (typed) 20Coding — TypeScript Type System 20Coding — React Debugging 20

Doc/OCR vision 1d 160 / 160 20 8/8 n=1 63.8 run 2026-09-23
Test Category Score tok/s
D1 Doc/OCR — Transcription 20/20 69.8
D2 Doc/OCR — Document QA 20/20 71.1
D3 Doc/OCR — Field Extraction 20/20 73.1
D4 Doc/OCR — Document QA 20/20 64.4
D5 Doc/OCR — Transcription 20/20 69.2
D6 Doc/OCR — Field Extraction 20/20 57.2
D7 Doc/OCR — Degraded Input 20/20 53.5
D8 Doc/OCR — Degraded Input 20/20 52.1

Category mean Doc/OCR — Transcription 20Doc/OCR — Document QA 20Doc/OCR — Field Extraction 20Doc/OCR — Degraded Input 20

Doc/OCR — real-degraded tier 1d2 63.2 / 80 15.8 4/4 n=1 49 run 2026-09-21
Test Category Score tok/s
D9 Doc/OCR — Real Degraded Transcription 15/20 38.6
D10 Doc/OCR — Handwriting Extraction 17/20 47.6
D11 Doc/OCR — Degraded Table Transcription 14/20 54.5
D12 Doc/OCR — Degraded Table QA 17/20 55.5

Category mean Doc/OCR — Real Degraded Transcription 15Doc/OCR — Handwriting Extraction 17Doc/OCR — Degraded Table Transcription 14Doc/OCR — Degraded Table QA 17

Content-production depth 1e 115.8 / 120 19.3 6/6 n=2 44.5 run 2026-09-26
Test Category Score tok/s
E1 SOW & Requirements Decomposition 20/20 41.7
E2 SOW & Requirements Decomposition 20/20 45.6
E3 Executive & Proposal Content 18/20 40.2
E4 Executive & Proposal Content 19/20 47.5
E5 Executive & Proposal Content 18/20 53.5
E6 Executive & Proposal Content 20/20 38.6

Category mean SOW & Requirements Decomposition 20Executive & Proposal Content 18.8

Long-context retrieval & synthesis 1g 117.6 / 120 19.6 6/6 n=3 53.2 run 2026-09-27
Test Category Score tok/s
L1 Long-Context — Needle Retrieval 20/20 69.4
L2 Long-Context — Grounded QA with Distractors 20/20 62.9
L3 Long-Context — Ops Log Reasoning 20/20 46.4
L4 Long-Context — Faithful Summarization 20/20 52.5
L5 Long-Context — Instruction Retention 17/20 47.6
L6 Long-Context — Full-Haystack QA 20/20 40.4

Category mean Long-Context — Needle Retrieval 20Long-Context — Grounded QA with Distractors 20Long-Context — Ops Log Reasoning 20Long-Context — Faithful Summarization 20Long-Context — Instruction Retention 17Long-Context — Full-Haystack QA 20

Long-context multi-needle (MRCR) 1g2 44.1 / 60 14.7 3/3 n=2 43.7 run 2026-09-26
Test Category Score tok/s
L7 Long-Context — Multi-Needle Retrieval (MRCR) 20/20 48.1
L8 Long-Context — Multi-Needle Retrieval (MRCR) 20/20 44.5
L9 Long-Context — Multi-Needle Retrieval (MRCR) 4/20 38.6
Live one-shot builds (runtime-verified) 1h 18 / 80 4.5 4/4 n=1 39.8 run 2026-09-17
Test Category Score tok/s
H4 Web-Estate One-Shot Builds 18/20 42.7
Brand-constrained static build (negative-constraint adherence) 1j 69.5 / 100 13.9 5/5 n=2 46.4 run 2026-09-26
Test Category Score tok/s
J3 Brand-Constrained Static Build 20/20 43.5
J4 Brand-Constrained Static Build 20/20 48.3
J5 Brand-Constrained Static Build 13.5/20 49.3
J6 Brand-Constrained Static Build 16/20 51.6

Each test is LLM-judged 0–20 against a fixed rubric; suite max = tests × 20. An ADJ chip marks a cell carrying a subjective adjustment from the maintainer; the judged score is the one shown, is unchanged, and remains what the axes, ranks and tie bands are computed from. Hover for the reason. Rows with a ▶ expand to the per-test breakdown. Where fewer tests ran than the suite total, unrun tests count as zero toward the axis score. See methodology.