← Leaderboard

Qwen3.8-27B Q4_K_M GGUF

Tested RAN THE GAUNTLET Rank #4 of 51 · 8/8 GAUNTLET progress
Generalist — 95 / 100 G Agentic — 100 / 100 A Understanding — 95.3 / 100 U Needle — 91.2 / 100 N Thinking — 95.7 / 100 T Live — 85.6 / 100 L Engineering — 94.5 / 100 E Throughput — 7.6 / 100 T
G
95
A
100
U
95.3
N
91.2
T
95.7
L
85.6
E
94.5
T
7.6

Specification

Parameters
27B dense
Architecture
qwen3_5
Size on disk
17.74 GB
Quantization
Q4_K_M
Format
GGUF
Runtime
LM Studio
Class
Standard
Reasoning (CoT)
Yes — emits reasoning tokens
Internal ID
M25a
Mean speed
26.7 tok/s across suites
Stall census
1 stall in 23 observed tests (4.3%)
Reasoning appetite
2,350 tokens mean · 16,383 max
Model card
huggingface.co

Suite results

Suite Score Avg / 20 Tests n tok/s Run
General capability (13-task real-workload suite) 1a 247 / 260 19 13/13 n=1.8 23.4 run 2026-08-15
Test Category Score tok/s
A1 Dev/Ops Scripting 20/20 26.3
A2 Dev/Ops Scripting 20/20 22.2
A3 Dev/Ops Scripting 10/20 20.9
A4 Dev/Ops Scripting 20/20 21.5
A5 Dev/Ops Scripting 20/20 20.7
A6 Dev/Ops Scripting 20/20 22.3
B1 Document Processing 20/20 24
B2 Document Processing 20/20 24.8
B3 Document Processing 20/20 22.4
C1 Content Production 20/20 18.2
C2 Content Production 18/20 20.5
C3 Content Production 20/20 21
C4 Content Production 20/20 21.7

Category mean Dev/Ops Scripting 18.3Document Processing 20Content Production 19.5

Agentic tool-calling & protocol adherence 1b 160 / 160 20 8/8 n=1 32 run 2026-08-15
Test Category Score tok/s
WA1 Tool Calling & Protocol Adherence 20/20 23.4
WA2 Tool Calling & Protocol Adherence 20/20 20.6
WA3 Tool Calling & Protocol Adherence 20/20 27.6
WA4 Tool Calling & Protocol Adherence 20/20 18.4
WA5 Tool Calling & Protocol Adherence 20/20 24.9
WB1 Agent Robustness & State Management 20/20 24.2
WB2 Agent Robustness & State Management 20/20 20.6
WB3 Agent Robustness & State Management 20/20 20.5

Category mean Tool Calling & Protocol Adherence 20Agent Robustness & State Management 20

Coding depth 1c 170.1 / 180 18.9 9/9 n=1 27.4 run 2026-09-17
Test Category Score tok/s
K1 Coding — Concurrency 20/20 29
K2 Coding — Algorithms 20/20 28.8
K3 Coding — Security Review 19/20 27.1
K4 Coding — Performance 19/20 25.9
K5 Coding — Refactoring Under Constraints 20/20 24.7
K6 Coding — Test Writing 20/20 29.2
K7 Coding — React Component (typed) 16/20 27.4
K8 Coding — TypeScript Type System 16/20 29.4
K9 Coding — React Debugging 20/20 25.1

Category mean Coding — Concurrency 20Coding — Algorithms 20Coding — Security Review 19Coding — Performance 19Coding — Refactoring Under Constraints 20Coding — Test Writing 20Coding — React Component (typed) 16Coding — TypeScript Type System 16Coding — React Debugging 20

Doc/OCR vision 1d 160 / 160 20 8/8 n=1 28.7 run 2026-09-24
Test Category Score tok/s
D1 Doc/OCR — Transcription 20/20 18.8
D2 Doc/OCR — Document QA 20/20 11.8
D3 Doc/OCR — Field Extraction 20/20 19.1
D4 Doc/OCR — Document QA 20/20 28.1
D5 Doc/OCR — Transcription 20/20 15.8
D6 Doc/OCR — Field Extraction 20/20 15.3
D7 Doc/OCR — Degraded Input 20/20 21.4
D8 Doc/OCR — Degraded Input 20/20 20.2

Category mean Doc/OCR — Transcription 20Doc/OCR — Document QA 20Doc/OCR — Field Extraction 20Doc/OCR — Degraded Input 20

Doc/OCR — real-degraded tier 1d2 68.8 / 80 17.2 4/4 n=1 32.6 run 2026-09-22
Test Category Score tok/s
D9 Doc/OCR — Real Degraded Transcription 15/20 24.8
D10 Doc/OCR — Handwriting Extraction 17/20 16
D11 Doc/OCR — Degraded Table Transcription 17/20 29.7
D12 Doc/OCR — Degraded Table QA 20/20 26.7

Category mean Doc/OCR — Real Degraded Transcription 15Doc/OCR — Handwriting Extraction 17Doc/OCR — Degraded Table Transcription 17Doc/OCR — Degraded Table QA 20

Content-production depth 1e 99 / 120 19.8 5/6 ● n=1 24.9 run 2026-09-25
Test Category Score tok/s
E1 SOW & Requirements Decomposition 20/20 26.9
E2 SOW & Requirements Decomposition 20/20 26.2
E4 Executive & Proposal Content 20/20 24.1
E5 Executive & Proposal Content 19/20 25.2
E6 Executive & Proposal Content 20/20 21.9

Category mean SOW & Requirements Decomposition 20Executive & Proposal Content 19.7

Long-context retrieval & synthesis 1g 120 / 120 20 6/6 n=1 25.8 run 2026-09-24
Test Category Score tok/s
L1 Long-Context — Needle Retrieval 20/20 12.2
L2 Long-Context — Grounded QA with Distractors 20/20 13.2
L3 Long-Context — Ops Log Reasoning 20/20 8.4
L4 Long-Context — Faithful Summarization 20/20 16.6
L5 Long-Context — Instruction Retention 20/20 4.9
L6 Long-Context — Full-Haystack QA 20/20 2.9

Category mean Long-Context — Needle Retrieval 20Long-Context — Grounded QA with Distractors 20Long-Context — Ops Log Reasoning 20Long-Context — Faithful Summarization 20Long-Context — Instruction Retention 20Long-Context — Full-Haystack QA 20

Long-context multi-needle (MRCR) 1g2 44.1 / 60 14.7 3/3 n=1.3 24 run 2026-09-21
Test Category Score tok/s
L7 Long-Context — Multi-Needle Retrieval (MRCR) 20/20 32.3
L8 Long-Context — Multi-Needle Retrieval (MRCR) 20/20 4.4
L9 Long-Context — Multi-Needle Retrieval (MRCR) 4/20 17.4
Live one-shot builds (runtime-verified) 1h 59.1 / 80 19.7 3/4 ● n=1 27.4 run 2026-09-18
Test Category Score tok/s
H1 Web-Estate One-Shot Builds 20/20 31.2
H2 Web-Estate One-Shot Builds 19/20 26.5
H4 Web-Estate One-Shot Builds 20/20 24.5
Brand-constrained static build (negative-constraint adherence) 1j 95 / 100 19 5/5 n=1 21.2 run 2026-09-23
Test Category Score tok/s
J1 Brand-Constrained Static Build 20/20 26.4
J3 Brand-Constrained Static Build 20/20 19.4
J4 Brand-Constrained Static Build 20/20 18.7
J5 Brand-Constrained Static Build 20/20 20.6
J6 Brand-Constrained Static Build 15/20 21.1

Each test is LLM-judged 0–20 against a fixed rubric; suite max = tests × 20. An ADJ chip marks a cell carrying a subjective adjustment from the maintainer; the judged score is the one shown, is unchanged, and remains what the axes, ranks and tie bands are computed from. Hover for the reason. Rows with a ▶ expand to the per-test breakdown. Where fewer tests ran than the suite total, unrun tests count as zero toward the axis score. See methodology.