← Leaderboard

Qwen3.8-27B GSQ-RCO IQ3_S GGUF ISTA-DASLab (Ollama, MBP — identical-file arm of M35)

Daily-driver trial Rank #13 of 51 · 7/8 GAUNTLET progress
Generalist — 98 / 100 G Agentic — 100 / 100 A Understanding — 92 / 100 U Needle — 91.2 / 100 N Thinking — pending T Live — 84.2 / 100 L Engineering — 99 / 100 E Throughput — 6.2 / 100 T

Hollow markers & dashed spokes: axis not yet scored

G
98
A
100
U
92
N
91.2
T
n/d
L
84.2
E
99
T
6.2

Runtime cohort \u2014 read this before comparing

This model was benchmarked on Ollama. Most of the board was benchmarked on LM Studio. The two are not interchangeable and the scores are published side by side deliberately rather than merged.

Why both exist: on the benchmark machine, LM Studio drops generation streams on a minority of attempts overall (15.9% of tests first-attempt across 1,526 tests) but heavily on specific models \u2014 several exceed 45%, two exceed 90%. A dropped stream is indistinguishable from a bad answer in a tally, so unattended benchmarking moved to Ollama, which has recorded 0 stream deaths in 177 tests.

What the offset looks like: across every suite where a byte-identical pair has been scored on both runtimes — 23 paired observations over 4 pairs — Ollama runs 0.1 points higher on average, with mean |Δ| 0.6 and sd 1.2. The single largest gap between two byte-identical arms is 4 points — bigger than most of the gaps at the top of the leaderboard. Read this as the noise floor of the whole benchmark rather than as a conversion factor, and treat cross-runtime gaps under 1 point as noise.

Specification

Parameters
27B dense
Architecture
qwen3_5
Size on disk
12.7 GB
Quantization
IQ3_S
Format
GGUF
Runtime
Ollama — scores are NOT directly comparable to the LM Studio cohort; see the runtime note below
Class
Standard
Reasoning (CoT)
Yes — emits reasoning tokens
Internal ID
MO5
Mean speed
21.7 tok/s across suites
Model card
huggingface.co

Suite results

Suite Score Avg / 20 Tests n tok/s Run
General capability (13-task real-workload suite) 1a 254.8 / 260 19.6 13/13 n=2 23.6 run 2026-09-13
Test Category Score tok/s
A1 Dev/Ops Scripting 20/20 24
A2 Dev/Ops Scripting 20/20 23.8
A3 Dev/Ops Scripting 20/20 23.7
A4 Dev/Ops Scripting 19/20 23.5
A5 Dev/Ops Scripting 20/20 23.9
A6 Dev/Ops Scripting 20/20 24.6
B1 Document Processing 20/20 25.1
B2 Document Processing 20/20 25
B3 Document Processing 20/20 22
C1 Content Production 20/20 22.1
C2 Content Production 18/20 22.9
C3 Content Production 18/20 23.3
C4 Content Production 20/20 23.4

Category mean Dev/Ops Scripting 19.8Document Processing 20Content Production 19

Agentic tool-calling & protocol adherence 1b 160 / 160 20 8/8 n=1 28.7 run 2026-09-13
Test Category Score tok/s
WA1 Tool Calling & Protocol Adherence 20/20 29.3
WA2 Tool Calling & Protocol Adherence 20/20 32.8
WA3 Tool Calling & Protocol Adherence 20/20 32.8
WA4 Tool Calling & Protocol Adherence 20/20 33.2
WA5 Tool Calling & Protocol Adherence 20/20 29.1
WB1 Agent Robustness & State Management 20/20 23.4
WB2 Agent Robustness & State Management 20/20 25.2
WB3 Agent Robustness & State Management 20/20 23.6

Category mean Tool Calling & Protocol Adherence 20Agent Robustness & State Management 20

Coding depth 1c 178.2 / 180 19.8 9/9 n=1 20.8 run 2026-09-13
Test Category Score tok/s
K1 Coding — Concurrency 19/20 25.5
K2 Coding — Algorithms 20/20 24
K3 Coding — Security Review 19/20 21.4
K4 Coding — Performance 20/20 17.5
K5 Coding — Refactoring Under Constraints 20/20 17.9
K6 Coding — Test Writing 20/20 20.1
K7 Coding — React Component (typed) 20/20 20.7
K8 Coding — TypeScript Type System 20/20 20.6
K9 Coding — React Debugging 20/20 19.2

Category mean Coding — Concurrency 19Coding — Algorithms 20Coding — Security Review 19Coding — Performance 20Coding — Refactoring Under Constraints 20Coding — Test Writing 20Coding — React Component (typed) 20Coding — TypeScript Type System 20Coding — React Debugging 20

Doc/OCR vision 1d 160 / 160 20 8/8 n=1 21.9 run 2026-09-23
Test Category Score tok/s
D1 Doc/OCR — Transcription 20/20 27
D2 Doc/OCR — Document QA 20/20 25.6
D3 Doc/OCR — Field Extraction 20/20 24.4
D4 Doc/OCR — Document QA 20/20 18.5
D5 Doc/OCR — Transcription 20/20 18.6
D6 Doc/OCR — Field Extraction 20/20 19.4
D7 Doc/OCR — Degraded Input 20/20 20.7
D8 Doc/OCR — Degraded Input 20/20 20.8

Category mean Doc/OCR — Transcription 20Doc/OCR — Document QA 20Doc/OCR — Field Extraction 20Doc/OCR — Degraded Input 20

Doc/OCR — real-degraded tier 1d2 60.8 / 80 15.2 4/4 n=1 24.6 run 2026-09-23
Test Category Score tok/s
D9 Doc/OCR — Real Degraded Transcription 13/20 25.9
D10 Doc/OCR — Handwriting Extraction 17/20 25.1
D11 Doc/OCR — Degraded Table Transcription 17/20 25.1
D12 Doc/OCR — Degraded Table QA 14/20 22.1

Category mean Doc/OCR — Real Degraded Transcription 13Doc/OCR — Handwriting Extraction 17Doc/OCR — Degraded Table Transcription 17Doc/OCR — Degraded Table QA 14

Content-production depth 1e 120 / 120 20 6/6 n=1 22 run 2026-09-22
Test Category Score tok/s
E1 SOW & Requirements Decomposition 20/20 25
E2 SOW & Requirements Decomposition 20/20 22.2
E3 Executive & Proposal Content 20/20 20.9
E4 Executive & Proposal Content 20/20 21.1
E5 Executive & Proposal Content 20/20 20.9
E6 Executive & Proposal Content 20/20 22

Category mean SOW & Requirements Decomposition 20Executive & Proposal Content 20

Long-context retrieval & synthesis 1g 120 / 120 20 6/6 n=1 16.7 run 2026-09-20
Test Category Score tok/s
L1 Long-Context — Needle Retrieval 20/20 18.5
L2 Long-Context — Grounded QA with Distractors 20/20 17.5
L3 Long-Context — Ops Log Reasoning 20/20 18
L4 Long-Context — Faithful Summarization 20/20 16.2
L5 Long-Context — Instruction Retention 20/20 16
L6 Long-Context — Full-Haystack QA 20/20 13.8

Category mean Long-Context — Needle Retrieval 20Long-Context — Grounded QA with Distractors 20Long-Context — Ops Log Reasoning 20Long-Context — Faithful Summarization 20Long-Context — Instruction Retention 20Long-Context — Full-Haystack QA 20

Long-context multi-needle (MRCR) 1g2 44.1 / 60 14.7 3/3 n=1 16.5 run 2026-09-17
Test Category Score tok/s
L7 Long-Context — Multi-Needle Retrieval (MRCR) 20/20 18.5
L8 Long-Context — Multi-Needle Retrieval (MRCR) 20/20 15.7
L9 Long-Context — Multi-Needle Retrieval (MRCR) 4/20 15.2
Live one-shot builds (runtime-verified) 1h 79.2 / 80 19.8 4/4 n=1 20.1 run 2026-09-17
Test Category Score tok/s
H1 Web-Estate One-Shot Builds 20/20 20.6
H2 Web-Estate One-Shot Builds 20/20 20
H3 Web-Estate One-Shot Builds 19/20 19.3
H4 Web-Estate One-Shot Builds 20/20 20.4
Production replay (real agent workload) 1i 189.6 / 240 15.8 12/12 n=1 19.6 run 2026-09-26
Test Category Score tok/s
I1 Nightly Fix Replay 8/20 18.1
I2 Nightly Fix Replay 11/20 19.5
I3 Nightly Fix Replay 9/20 14.1
I4 Nightly Fix Replay 15/20 19.7
I5 Nightly Fix Replay 12/20 19.1
I6 Nightly Fix Replay 16/20 19.6
I7 Audit Judgment 18/20 20.9
I8 Audit Judgment 20/20 21.5
I9 Cron Dependency Reasoning 20/20 20.1
I10 Cron Dependency Reasoning 20/20 19.7
I11 Curation Judgment 20/20 20.9
I12 Curation Judgment 20/20 21.4

Category mean Nightly Fix Replay 11.8Audit Judgment 19Cron Dependency Reasoning 20Curation Judgment 20

Brand-constrained static build (negative-constraint adherence) 1j 85 / 100 17 5/5 n=1 24.5 run 2026-09-14
Test Category Score tok/s
J1 Brand-Constrained Static Build 20/20 26.4
J3 Brand-Constrained Static Build 20/20 26.8
J4 Brand-Constrained Static Build 20/20 26.5
J5 Brand-Constrained Static Build 11/20 22.5
J6 Brand-Constrained Static Build 14/20 20.5

Each test is LLM-judged 0–20 against a fixed rubric; suite max = tests × 20. An ADJ chip marks a cell carrying a subjective adjustment from the maintainer; the judged score is the one shown, is unchanged, and remains what the axes, ranks and tie bands are computed from. Hover for the reason. Rows with a ▶ expand to the per-test breakdown. Where fewer tests ran than the suite total, unrun tests count as zero toward the axis score. See methodology.