← Leaderboard

Qwen3.8-27B Q6_K GGUF

Keeper RAN THE GAUNTLET Rank #1 of 51 · 8/8 GAUNTLET progress
Generalist — 99 / 100 G Agentic — 96 / 100 A Understanding — 93.3 / 100 U Needle — 91.2 / 100 N Thinking — 96.5 / 100 T Live — 83 / 100 L Engineering — 86.7 / 100 E Throughput — 5.7 / 100 T
G
99
A
96
U
93.3
N
91.2
T
96.5
L
83
E
86.7
T
5.7

Specification

Parameters
27B dense
Architecture
qwen3_5
Size on disk
23.36 GB
Quantization
Q6_K
Format
GGUF
Runtime
LM Studio
Class
Standard
Reasoning (CoT)
Yes — emits reasoning tokens
Internal ID
M25b
Mean speed
20 tok/s across suites
Stall census
3 stalls in 86 observed tests (3.5%)
Reasoning appetite
3,233 tokens mean · 31,137 max
Model card
huggingface.co

Suite results

Suite Score Avg / 20 Tests n tok/s Run
General capability (13-task real-workload suite) 1a 257.4 / 260 19.8 13/13 n=4.2 17.9 run 2026-08-15
Test Category Score tok/s
A1 Dev/Ops Scripting 20/20 18.6
A2 Dev/Ops Scripting 20/20 17.9
A3 Dev/Ops Scripting 20/20 17.4
A4 Dev/Ops Scripting 19/20 18.2
A5 Dev/Ops Scripting 20/20 18.5
A6 Dev/Ops Scripting 20/20 17.9
B1 Document Processing 20/20 20.8
B2 Document Processing 20/20 21.7
B3 Document Processing 20/20 18.7
C1 Content Production 20/20 16.7
C2 Content Production 19/20 17.1
C3 Content Production 20/20 15.9
C4 Content Production 19/20 17.4

Category mean Dev/Ops Scripting 19.8Document Processing 20Content Production 19.5

Agentic tool-calling & protocol adherence 1b 153.6 / 160 19.2 8/8 n=1 25 run 2026-08-15
Test Category Score tok/s
WA1 Tool Calling & Protocol Adherence 20/20 20.1
WA2 Tool Calling & Protocol Adherence 20/20 18.1
WA3 Tool Calling & Protocol Adherence 20/20 21.4
WA4 Tool Calling & Protocol Adherence 20/20 15.6
WA5 Tool Calling & Protocol Adherence 20/20 19.3
WB1 Agent Robustness & State Management 20/20 18.7
WB2 Agent Robustness & State Management 20/20 17.6
WB3 Agent Robustness & State Management 14/20 15.9

Category mean Tool Calling & Protocol Adherence 20Agent Robustness & State Management 18

Coding depth 1c 156 / 180 19.5 8/9 ● n=1 18.6 run 2026-09-15
Test Category Score tok/s
K1 Coding — Concurrency 20/20 21.1
K2 Coding — Algorithms 20/20 18.8
K3 Coding — Security Review 16/20 17.2
K4 Coding — Performance 20/20 17.1
K5 Coding — Refactoring Under Constraints 20/20 18.2
K6 Coding — Test Writing 20/20 19
K7 Coding — React Component (typed) 20/20 18.9
K8 Coding — TypeScript Type System 20/20 19.3

Category mean Coding — Concurrency 20Coding — Algorithms 20Coding — Security Review 16Coding — Performance 20Coding — Refactoring Under Constraints 20Coding — Test Writing 20Coding — React Component (typed) 20Coding — TypeScript Type System 20

Doc/OCR vision 1d 160 / 160 20 8/8 n=1 22.5 run 2026-08-15
Test Category Score tok/s
D1 Doc/OCR — Transcription 20/20 13
D2 Doc/OCR — Document QA 20/20 7.3
D3 Doc/OCR — Field Extraction 20/20 14.9
D4 Doc/OCR — Document QA 20/20 13.6
D5 Doc/OCR — Transcription 20/20 10
D6 Doc/OCR — Field Extraction 20/20 8.8
D7 Doc/OCR — Degraded Input 20/20 13.8
D8 Doc/OCR — Degraded Input 20/20 12.4

Category mean Doc/OCR — Transcription 20Doc/OCR — Document QA 20Doc/OCR — Field Extraction 20Doc/OCR — Degraded Input 20

Doc/OCR — real-degraded tier 1d2 64 / 80 16 4/4 n=1 20.5 run 2026-08-15
Test Category Score tok/s
D9 Doc/OCR — Real Degraded Transcription 16/20 14
D10 Doc/OCR — Handwriting Extraction 16/20 9.6
D11 Doc/OCR — Degraded Table Transcription 17/20 17.7
D12 Doc/OCR — Degraded Table QA 15/20 14.4

Category mean Doc/OCR — Real Degraded Transcription 16Doc/OCR — Handwriting Extraction 16Doc/OCR — Degraded Table Transcription 17Doc/OCR — Degraded Table QA 15

Content-production depth 1e 118.8 / 120 19.8 6/6 n=1 18 run 2026-08-15
Test Category Score tok/s
E1 SOW & Requirements Decomposition 20/20 18.1
E2 SOW & Requirements Decomposition 20/20 17.4
E3 Executive & Proposal Content 20/20 16.7
E4 Executive & Proposal Content 20/20 18.7
E5 Executive & Proposal Content 19/20 19.5
E6 Executive & Proposal Content 20/20 17.5

Category mean SOW & Requirements Decomposition 20Executive & Proposal Content 19.8

Long-context retrieval & synthesis 1g 120 / 120 20 6/6 n=1 20.6 run 2026-08-15
Test Category Score tok/s
L1 Long-Context — Needle Retrieval 20/20 5.5
L2 Long-Context — Grounded QA with Distractors 20/20 7.5
L3 Long-Context — Ops Log Reasoning 20/20 3.7
L4 Long-Context — Faithful Summarization 20/20 5.3
L5 Long-Context — Instruction Retention 20/20 3.2
L6 Long-Context — Full-Haystack QA 20/20 1.3

Category mean Long-Context — Needle Retrieval 20Long-Context — Grounded QA with Distractors 20Long-Context — Ops Log Reasoning 20Long-Context — Faithful Summarization 20Long-Context — Instruction Retention 20Long-Context — Full-Haystack QA 20

Long-context multi-needle (MRCR) 1g2 44.1 / 60 14.7 3/3 n=1 18.1 run 2026-08-15
Test Category Score tok/s
L7 Long-Context — Multi-Needle Retrieval (MRCR) 20/20 3.7
L8 Long-Context — Multi-Needle Retrieval (MRCR) 20/20 2.1
L9 Long-Context — Multi-Needle Retrieval (MRCR) 4/20 1.7
Live one-shot builds (runtime-verified) 1h 80 / 80 20 4/4 n=1 19.4 run 2026-08-15
Test Category Score tok/s
H1 Web-Estate One-Shot Builds 20/20 24
H2 Web-Estate One-Shot Builds 20/20 19.2
H3 Web-Estate One-Shot Builds 20/20 16.7
H4 Web-Estate One-Shot Builds 20/20 18.6
Production replay (real agent workload) 1i 177.6 / 240 14.8 12/12 n=1.1 15.1 run 2026-08-16
Test Category Score tok/s
I1 Nightly Fix Replay 9/20 13.5
I2 Nightly Fix Replay 11/20 13.9
I3 Nightly Fix Replay 8/20 8.9
I4 Nightly Fix Replay 14/20 13.5
I5 Nightly Fix Replay 12/20 14.5
I6 Nightly Fix Replay 10/20 13.6
I7 Audit Judgment 14/20 14.3
I8 Audit Judgment 20/20 14.5
I9 Cron Dependency Reasoning 20/20 12.9
I10 Cron Dependency Reasoning 20/20 13.8
I11 Curation Judgment 20/20 13.3
I12 Curation Judgment 20/20 14

Category mean Nightly Fix Replay 10.7Audit Judgment 17Cron Dependency Reasoning 20Curation Judgment 20

Brand-constrained static build (negative-constraint adherence) 1j 91 / 100 18.2 5/5 n=1 24.5 run 2026-09-10
Test Category Score tok/s
J1 Brand-Constrained Static Build 11/20 26.2
J3 Brand-Constrained Static Build 20/20 22.9
J4 Brand-Constrained Static Build 20/20 24.4
J5 Brand-Constrained Static Build 20/20 25.7
J6 Brand-Constrained Static Build 20/20 23.3

Each test is LLM-judged 0–20 against a fixed rubric; suite max = tests × 20. An ADJ chip marks a cell carrying a subjective adjustment from the maintainer; the judged score is the one shown, is unchanged, and remains what the axes, ranks and tie bands are computed from. Hover for the reason. Rows with a ▶ expand to the per-test breakdown. Where fewer tests ran than the suite total, unrun tests count as zero toward the axis score. See methodology.