← Leaderboard

Qwen3.8-27B Q8_0 GGUF

tested-offloaded Rank #39 of 51 · 4/8 GAUNTLET progress
Generalist — 94 / 100 G Agentic — 95.5 / 100 A Understanding — pending U Needle — pending N Thinking — 100 / 100 T Live — pending L Engineering — pending E Throughput — 7.8 / 100 T

Hollow markers & dashed spokes: axis not yet scored

G
94
A
95.5
U
—
N
—
T
100
L
—
E
—
T
7.8

Specification

Parameters
27B dense
Architecture
qwen3_5
Size on disk
29.98 GB
Quantization
Q8_0
Format
GGUF
Runtime
LM Studio
Class
Standard
Reasoning (CoT)
Yes — emits reasoning tokens
Internal ID
M25c
Mean speed
27.6 tok/s across suites
Stall census
0 stalls in 21 observed tests (0.0%)
Reasoning appetite
1,614 tokens mean · 11,418 max
Model card
huggingface.co

Suite results

Suite Score Avg / 20 Tests n tok/s Run
General capability (13-task real-workload suite) 1a 244.4 / 260 18.8 13/13 n=1 24 run 2026-08-15
Test Category Score tok/s
A1 Dev/Ops Scripting 20/20 24.7
A2 Dev/Ops Scripting 20/20 22.8
A3 Dev/Ops Scripting 10/20 21.5
A4 Dev/Ops Scripting 19/20 22.6
A5 Dev/Ops Scripting 20/20 23.8
A6 Dev/Ops Scripting 20/20 23.8
B1 Document Processing 20/20 26.8
B2 Document Processing 20/20 28.3
B3 Document Processing 20/20 23.8
C1 Content Production 20/20 20.8
C2 Content Production 18/20 21.4
C3 Content Production 18/20 20
C4 Content Production 20/20 21.7

Category mean Dev/Ops Scripting 18.2Document Processing 20Content Production 19

Agentic tool-calling & protocol adherence 1b 152.8 / 160 19.1 8/8 n=1 31.1 run 2026-08-15
Test Category Score tok/s
WA1 Tool Calling & Protocol Adherence 20/20 24.7
WA2 Tool Calling & Protocol Adherence 20/20 21.9
WA3 Tool Calling & Protocol Adherence 20/20 24.5
WA4 Tool Calling & Protocol Adherence 20/20 16.9
WA5 Tool Calling & Protocol Adherence 20/20 26.3
WB1 Agent Robustness & State Management 20/20 23.7
WB2 Agent Robustness & State Management 20/20 23.1
WB3 Agent Robustness & State Management 13/20 20.3

Category mean Tool Calling & Protocol Adherence 20Agent Robustness & State Management 17.7

Each test is LLM-judged 0–20 against a fixed rubric; suite max = tests × 20. An ADJ chip marks a cell carrying a subjective adjustment from the maintainer; the judged score is the one shown, is unchanged, and remains what the axes, ranks and tie bands are computed from. Hover for the reason. Rows with a ▶ expand to the per-test breakdown. Where fewer tests ran than the suite total, unrun tests count as zero toward the axis score. See methodology.