← Leaderboard

Qwen3.8-27B Q6_K GGUF

Daily-driver trial RAN THE GAUNTLET Rank #1 of 33 · 8/8 GAUNTLET progress
Generalist — 99 / 100 G Agentic — 100 / 100 A Understanding — 90.1 / 100 U Needle — 89.3 / 100 N Thinking — 96.5 / 100 T Live — 75.7 / 100 L Engineering — 98.5 / 100 E Throughput — 8.1 / 100 T
G
99
A
100
U
90.1
N
89.3
T
96.5
L
75.7
E
98.5
T
8.1

Specification

Parameters
27B dense
Architecture
qwen3_5
Size on disk
23.36 GB
Quantization
Q6_K
Format
GGUF
Reasoning (CoT)
Yes — emits reasoning tokens
Internal ID
M25b
Mean speed
13.2 tok/s across suites
Stall census
3 stalls in 86 observed tests (3.5%)
Reasoning appetite
3,233 tokens mean · 31,137 max
Model card
huggingface.co

Suite results

Suite Score Avg / 20 Tests tok/s
General capability (13-task real-workload suite) 1a 257.4 / 260 19.8 13/13 18.2
Agentic tool-calling & protocol adherence 1b 160 / 160 20 8/8 18.3
Coding depth 1c 118.2 / 120 19.7 6/6 18.6
Doc/OCR vision 1d 155.2 / 160 19.4 8/8 11.7
Doc/OCR — real-degraded tier 1d2 61 / 80 15.3 4/4 13.9
Content-production depth 1e 118.8 / 120 19.8 6/6 18
Long-context retrieval & synthesis 1g 118.8 / 120 19.8 6/6 4.4
Long-context multi-needle (MRCR) 1g2 42 / 60 14 3/3 2.5
Live one-shot builds (runtime-verified) 1h 79 / 80 19.8 4/4
Production replay (real agent workload) 1i 163.2 / 240 13.6 12/12 13.4

Each test is LLM-judged 0–20 against a fixed rubric; suite max = tests × 20. Where fewer tests ran than the suite total, unrun tests count as zero toward the axis score. See methodology.