← Leaderboard

Qwen3.8-27B MLX 5bit (lmstudio-community)

untested Rank #23 of 36 · 3/8 GAUNTLET progress
Generalist — 74.6 / 100 G Agentic — 100 / 100 A Understanding — pending U Needle — pending N Thinking — pending T Live — pending L Engineering — pending E Throughput — 6.6 / 100 T

Hollow markers & dashed spokes: axis not yet scored

G
74.6
A
100
U
N
T
n/d
L
E
T
6.6

Specification

Parameters
27B dense
Architecture
qwen3_5
Size on disk
19 GB
Quantization
5bit
Format
MLX
Class
Standard
Reasoning (CoT)
Yes — emits reasoning tokens
Internal ID
M25i
Mean speed
15.8 tok/s across suites
Model card
huggingface.co

Suite results

Suite Score Avg / 20 Tests tok/s Run
General capability (13-task real-workload suite) 1a 194 / 260 19.4 10/13 16.3 run 2026-09-02
Test Category Score tok/s
A1 Dev/Ops Scripting 20/20 13.1
A2 Dev/Ops Scripting 19/20 13.4
A3 Dev/Ops Scripting 20/20 13.2
A4 Dev/Ops Scripting 19/20 17.3
A5 Dev/Ops Scripting 20/20 19.1
A6 Dev/Ops Scripting 19/20 19.7
B1 Document Processing 20/20 18.9
B2 Document Processing 20/20 17.2
C2 Content Production 18/20 14.8
C4 Content Production 19/20 15.9

Category mean Dev/Ops Scripting 19.5Document Processing 20Content Production 18.5

Agentic tool-calling & protocol adherence 1b 160 / 160 20 8/8 15.3 run 2026-09-02
Test Category Score tok/s
WA1 Tool Calling & Protocol Adherence 20/20 14.9
WA2 Tool Calling & Protocol Adherence 20/20 16.8
WA3 Tool Calling & Protocol Adherence 20/20 16.5
WA4 Tool Calling & Protocol Adherence 20/20 19.8
WA5 Tool Calling & Protocol Adherence 20/20 14.5
WB1 Agent Robustness & State Management 20/20 12.8
WB2 Agent Robustness & State Management 20/20 13.8
WB3 Agent Robustness & State Management 20/20 13.5

Category mean Tool Calling & Protocol Adherence 20Agent Robustness & State Management 20

Each test is LLM-judged 0–20 against a fixed rubric; suite max = tests × 20. Rows with a expand to the per-test breakdown. Where fewer tests ran than the suite total, unrun tests count as zero toward the axis score. See methodology.