← Leaderboard

Qwen3.8-27B Uncensored HauhauCS Aggressive Q5_K_P GGUF (Ollama, MBP — identical-file arm of M29)

Tested Rank #28 of 51 · 6/8 GAUNTLET progress
Generalist — 97.5 / 100 G Agentic — 97 / 100 A Understanding — pending U Needle — 91.2 / 100 N Thinking — pending T Live — 94.6 / 100 L Engineering — 99 / 100 E Throughput — 4.4 / 100 T

Hollow markers & dashed spokes: axis not yet scored

G
97.5
A
97
U
—
N
91.2
T
n/d
L
94.6
E
99
T
4.4

Runtime cohort \u2014 read this before comparing

This model was benchmarked on Ollama. Most of the board was benchmarked on LM Studio. The two are not interchangeable and the scores are published side by side deliberately rather than merged.

Why both exist: on the benchmark machine, LM Studio drops generation streams on a minority of attempts overall (15.9% of tests first-attempt across 1,526 tests) but heavily on specific models \u2014 several exceed 45%, two exceed 90%. A dropped stream is indistinguishable from a bad answer in a tally, so unattended benchmarking moved to Ollama, which has recorded 0 stream deaths in 177 tests.

What the offset looks like: across every suite where a byte-identical pair has been scored on both runtimes — 23 paired observations over 4 pairs — Ollama runs 0.1 points higher on average, with mean |Δ| 0.6 and sd 1.2. The single largest gap between two byte-identical arms is 4 points — bigger than most of the gaps at the top of the leaderboard. Read this as the noise floor of the whole benchmark rather than as a conversion factor, and treat cross-runtime gaps under 1 point as noise.

Specification

Parameters
27B dense
Architecture
qwen3_5
Size on disk
20.2 GB
Quantization
Q5_K_P
Format
GGUF
Runtime
Ollama — scores are NOT directly comparable to the LM Studio cohort; see the runtime note below
Class
Standard
Reasoning (CoT)
Yes — emits reasoning tokens
Internal ID
MO9
Mean speed
15.6 tok/s across suites
Model card
huggingface.co

Suite results

Suite Score Avg / 20 Tests n tok/s Run
General capability (13-task real-workload suite) 1a 253.5 / 260 19.5 13/13 n=1 15.3 run 2026-09-25
Test Category Score tok/s
A1 Dev/Ops Scripting 18/20 20.3
A2 Dev/Ops Scripting 19/20 19.6
A3 Dev/Ops Scripting 20/20 15.6
A4 Dev/Ops Scripting 20/20 13.7
A5 Dev/Ops Scripting 20/20 14.1
A6 Dev/Ops Scripting 20/20 14.6
B1 Document Processing 20/20 15.1
B2 Document Processing 20/20 14.7
B3 Document Processing 20/20 13.3
C1 Content Production 20/20 14.5
C2 Content Production 18/20 14.6
C3 Content Production 18/20 14.4
C4 Content Production 20/20 14.4

Category mean Dev/Ops Scripting 19.5Document Processing 20Content Production 19

Agentic tool-calling & protocol adherence 1b 155.2 / 160 19.4 8/8 n=1 21 run 2026-09-29
Test Category Score tok/s
WA1 Tool Calling & Protocol Adherence 20/20 25.1
WA2 Tool Calling & Protocol Adherence 20/20 25.7
WA3 Tool Calling & Protocol Adherence 20/20 25.8
WA4 Tool Calling & Protocol Adherence 20/20 25.8
WA5 Tool Calling & Protocol Adherence 20/20 23.3
WB1 Agent Robustness & State Management 20/20 16.8
WB2 Agent Robustness & State Management 20/20 12.9
WB3 Agent Robustness & State Management 15/20 12.8

Category mean Tool Calling & Protocol Adherence 20Agent Robustness & State Management 18.3

Coding depth 1c 178.2 / 180 19.8 9/9 n=1 14.7 run 2026-10-02
Test Category Score tok/s
K1 Coding — Concurrency 20/20 18.3
K2 Coding — Algorithms 20/20 13.5
K3 Coding — Security Review 19/20 14.1
K4 Coding — Performance 19/20 14.4
K5 Coding — Refactoring Under Constraints 20/20 14.6
K6 Coding — Test Writing 20/20 14.5
K7 Coding — React Component (typed) 20/20 14.5
K8 Coding — TypeScript Type System 20/20 14.6
K9 Coding — React Debugging 20/20 13.9

Category mean Coding — Concurrency 20Coding — Algorithms 20Coding — Security Review 19Coding — Performance 19Coding — Refactoring Under Constraints 20Coding — Test Writing 20Coding — React Component (typed) 20Coding — TypeScript Type System 20Coding — React Debugging 20

Content-production depth 1e 118.8 / 120 19.8 6/6 n=1 14.7 run 2026-10-02
Test Category Score tok/s
E1 SOW & Requirements Decomposition 20/20 16.4
E2 SOW & Requirements Decomposition 20/20 14.2
E3 Executive & Proposal Content 19/20 14.1
E4 Executive & Proposal Content 20/20 14.8
E5 Executive & Proposal Content 20/20 14.3
E6 Executive & Proposal Content 20/20 14.6

Category mean SOW & Requirements Decomposition 20Executive & Proposal Content 19.8

Long-context retrieval & synthesis 1g 120 / 120 20 6/6 n=1 16.1 run 2026-09-29
Test Category Score tok/s
L1 Long-Context — Needle Retrieval 20/20 21.9
L2 Long-Context — Grounded QA with Distractors 20/20 20.2
L3 Long-Context — Ops Log Reasoning 20/20 16.6
L4 Long-Context — Faithful Summarization 20/20 12.9
L5 Long-Context — Instruction Retention 20/20 12.3
L6 Long-Context — Full-Haystack QA 20/20 12.4

Category mean Long-Context — Needle Retrieval 20Long-Context — Grounded QA with Distractors 20Long-Context — Ops Log Reasoning 20Long-Context — Faithful Summarization 20Long-Context — Instruction Retention 20Long-Context — Full-Haystack QA 20

Long-context multi-needle (MRCR) 1g2 44.1 / 60 14.7 3/3 n=1 12.8 run 2026-09-28
Test Category Score tok/s
L7 Long-Context — Multi-Needle Retrieval (MRCR) 20/20 14.7
L8 Long-Context — Multi-Needle Retrieval (MRCR) 20/20 12.2
L9 Long-Context — Multi-Needle Retrieval (MRCR) 4/20 11.6
Live one-shot builds (runtime-verified) 1h 79.2 / 80 19.8 4/4 n=1 15.5 run 2026-10-02
Test Category Score tok/s
H1 Web-Estate One-Shot Builds 20/20 17.7
H2 Web-Estate One-Shot Builds 20/20 14.9
H3 Web-Estate One-Shot Builds 19/20 14.6
H4 Web-Estate One-Shot Builds 20/20 14.7
Brand-constrained static build (negative-constraint adherence) 1j 91 / 100 18.2 5/5 n=1 14.6 run 2026-09-29
Test Category Score tok/s
J1 Brand-Constrained Static Build 20/20 15.9
J3 Brand-Constrained Static Build 20/20 13.9
J4 Brand-Constrained Static Build 20/20 14.7
J5 Brand-Constrained Static Build 20/20 14.5
J6 Brand-Constrained Static Build 11/20 14.1

Each test is LLM-judged 0–20 against a fixed rubric; suite max = tests × 20. An ADJ chip marks a cell carrying a subjective adjustment from the maintainer; the judged score is the one shown, is unchanged, and remains what the axes, ranks and tie bands are computed from. Hover for the reason. Rows with a ▶ expand to the per-test breakdown. Where fewer tests ran than the suite total, unrun tests count as zero toward the axis score. See methodology.