← Leaderboard

Qwen3.8-27B Q6_K GGUF lmstudio-community (Ollama, MBP - identical-file arm of M25b)

Tested Rank #12 of 51 · 7/8 GAUNTLET progress
Generalist — 98.5 / 100 G Agentic — 100 / 100 A Understanding — 94.7 / 100 U Needle — 91.2 / 100 N Thinking — pending T Live — 95 / 100 L Engineering — 97 / 100 E Throughput — 4.4 / 100 T

Hollow markers & dashed spokes: axis not yet scored

G
98.5
A
100
U
94.7
N
91.2
T
n/d
L
95
E
97
T
4.4

Runtime cohort \u2014 read this before comparing

This model was benchmarked on Ollama. Most of the board was benchmarked on LM Studio. The two are not interchangeable and the scores are published side by side deliberately rather than merged.

Why both exist: on the benchmark machine, LM Studio drops generation streams on a minority of attempts overall (15.9% of tests first-attempt across 1,526 tests) but heavily on specific models \u2014 several exceed 45%, two exceed 90%. A dropped stream is indistinguishable from a bad answer in a tally, so unattended benchmarking moved to Ollama, which has recorded 0 stream deaths in 177 tests.

What the offset looks like: across every suite where a byte-identical pair has been scored on both runtimes — 23 paired observations over 4 pairs — Ollama runs 0.1 points higher on average, with mean |Δ| 0.6 and sd 1.2. The single largest gap between two byte-identical arms is 4 points — bigger than most of the gaps at the top of the leaderboard. Read this as the noise floor of the whole benchmark rather than as a conversion factor, and treat cross-runtime gaps under 1 point as noise.

Specification

Parameters
27B dense
Architecture
qwen35
Size on disk
22.4 GB
Quantization
Q6_K
Format
GGUF
Runtime
Ollama — scores are NOT directly comparable to the LM Studio cohort; see the runtime note below
Class
Standard
Reasoning (CoT)
Yes — emits reasoning tokens
Internal ID
MO6
Mean speed
15.6 tok/s across suites
Model card
huggingface.co

Suite results

Suite Score Avg / 20 Tests n tok/s Run
General capability (13-task real-workload suite) 1a 256.1 / 260 19.7 13/13 n=1 16.1 run 2026-09-12
Test Category Score tok/s
A1 Dev/Ops Scripting 20/20 18.4
A2 Dev/Ops Scripting 20/20 18.6
A3 Dev/Ops Scripting 20/20 16.6
A4 Dev/Ops Scripting 19/20 15.9
A5 Dev/Ops Scripting 20/20 16
A6 Dev/Ops Scripting 20/20 16.2
B1 Document Processing 20/20 16.7
B2 Document Processing 20/20 16.8
B3 Document Processing 20/20 15
C1 Content Production 20/20 14.9
C2 Content Production 18/20 14.3
C3 Content Production 20/20 14.4
C4 Content Production 19/20 15.6

Category mean Dev/Ops Scripting 19.8Document Processing 20Content Production 19.3

Agentic tool-calling & protocol adherence 1b 160 / 160 20 8/8 n=1 18.5 run 2026-09-12
Test Category Score tok/s
WA1 Tool Calling & Protocol Adherence 20/20 17.8
WA2 Tool Calling & Protocol Adherence 20/20 17.1
WA3 Tool Calling & Protocol Adherence 20/20 17.3
WA4 Tool Calling & Protocol Adherence 20/20 19.2
WA5 Tool Calling & Protocol Adherence 20/20 21.7
WB1 Agent Robustness & State Management 20/20 19.8
WB2 Agent Robustness & State Management 20/20 17.9
WB3 Agent Robustness & State Management 20/20 17.2

Category mean Tool Calling & Protocol Adherence 20Agent Robustness & State Management 20

Coding depth 1c 174.6 / 180 19.4 9/9 n=1 15.7 run 2026-09-12
Test Category Score tok/s
K1 Coding — Concurrency 20/20 16.7
K2 Coding — Algorithms 20/20 16.5
K3 Coding — Security Review 16/20 16.3
K4 Coding — Performance 20/20 15.6
K5 Coding — Refactoring Under Constraints 20/20 14.9
K6 Coding — Test Writing 20/20 14.8
K7 Coding — React Component (typed) 19/20 15
K8 Coding — TypeScript Type System 20/20 16.1
K9 Coding — React Debugging 20/20 15.5

Category mean Coding — Concurrency 20Coding — Algorithms 20Coding — Security Review 16Coding — Performance 20Coding — Refactoring Under Constraints 20Coding — Test Writing 20Coding — React Component (typed) 19Coding — TypeScript Type System 20Coding — React Debugging 20

Doc/OCR vision 1d 160 / 160 20 8/8 n=1 15.2 run 2026-09-23
Test Category Score tok/s
D1 Doc/OCR — Transcription 20/20 19.9
D2 Doc/OCR — Document QA 20/20 18.8
D3 Doc/OCR — Field Extraction 20/20 14.4
D4 Doc/OCR — Document QA 20/20 13.1
D5 Doc/OCR — Transcription 20/20 13.2
D6 Doc/OCR — Field Extraction 20/20 13.6
D7 Doc/OCR — Degraded Input 20/20 14
D8 Doc/OCR — Degraded Input 20/20 14.6

Category mean Doc/OCR — Transcription 20Doc/OCR — Document QA 20Doc/OCR — Field Extraction 20Doc/OCR — Degraded Input 20

Doc/OCR — real-degraded tier 1d2 67.2 / 80 16.8 4/4 n=1 14.9 run 2026-09-23
Test Category Score tok/s
D9 Doc/OCR — Real Degraded Transcription 13/20 16.6
D10 Doc/OCR — Handwriting Extraction 17/20 14.9
D11 Doc/OCR — Degraded Table Transcription 17/20 14.3
D12 Doc/OCR — Degraded Table QA 20/20 13.8

Category mean Doc/OCR — Real Degraded Transcription 13Doc/OCR — Handwriting Extraction 17Doc/OCR — Degraded Table Transcription 17Doc/OCR — Degraded Table QA 20

Content-production depth 1e 118.8 / 120 19.8 6/6 n=1 15.3 run 2026-09-22
Test Category Score tok/s
E1 SOW & Requirements Decomposition 20/20 15.6
E2 SOW & Requirements Decomposition 20/20 15.8
E3 Executive & Proposal Content 20/20 15
E4 Executive & Proposal Content 20/20 14.5
E5 Executive & Proposal Content 19/20 15.3
E6 Executive & Proposal Content 20/20 15.5

Category mean SOW & Requirements Decomposition 20Executive & Proposal Content 19.8

Long-context retrieval & synthesis 1g 120 / 120 20 6/6 n=1 16 run 2026-09-22
Test Category Score tok/s
L1 Long-Context — Needle Retrieval 20/20 18.4
L2 Long-Context — Grounded QA with Distractors 20/20 17.5
L3 Long-Context — Ops Log Reasoning 20/20 15.9
L4 Long-Context — Faithful Summarization 20/20 15.3
L5 Long-Context — Instruction Retention 20/20 15.2
L6 Long-Context — Full-Haystack QA 20/20 13.6

Category mean Long-Context — Needle Retrieval 20Long-Context — Grounded QA with Distractors 20Long-Context — Ops Log Reasoning 20Long-Context — Faithful Summarization 20Long-Context — Instruction Retention 20Long-Context — Full-Haystack QA 20

Long-context multi-needle (MRCR) 1g2 44.1 / 60 14.7 3/3 n=1 12.6 run 2026-09-17
Test Category Score tok/s
L7 Long-Context — Multi-Needle Retrieval (MRCR) 20/20 13.8
L8 Long-Context — Multi-Needle Retrieval (MRCR) 20/20 12.4
L9 Long-Context — Multi-Needle Retrieval (MRCR) 4/20 11.7
Live one-shot builds (runtime-verified) 1h 80 / 80 20 4/4 n=1 16.1 run 2026-09-18
Test Category Score tok/s
H1 Web-Estate One-Shot Builds 20/20 17.5
H2 Web-Estate One-Shot Builds 20/20 15.6
H3 Web-Estate One-Shot Builds 20/20 15.2
H4 Web-Estate One-Shot Builds 20/20 15.9
Brand-constrained static build (negative-constraint adherence) 1j 91 / 100 18.2 5/5 n=1 15.9 run 2026-09-18
Test Category Score tok/s
J1 Brand-Constrained Static Build 11/20 16
J3 Brand-Constrained Static Build 20/20 15.9
J4 Brand-Constrained Static Build 20/20 16.4
J5 Brand-Constrained Static Build 20/20 15.9
J6 Brand-Constrained Static Build 20/20 15.5

Each test is LLM-judged 0–20 against a fixed rubric; suite max = tests × 20. An ADJ chip marks a cell carrying a subjective adjustment from the maintainer; the judged score is the one shown, is unchanged, and remains what the axes, ranks and tie bands are computed from. Hover for the reason. Rows with a ▶ expand to the per-test breakdown. Where fewer tests ran than the suite total, unrun tests count as zero toward the axis score. See methodology.