← Leaderboard

Qwen3.6-35B-A3B MLX 8bit Uniform (lmstudio-community)

Retired Rank #44 of 51 · 2/8 GAUNTLET progress
Generalist — 88 / 100 G Agentic — pending A Understanding — pending U Needle — pending N Thinking — pending T Live — pending L Engineering — pending E Throughput — 25.3 / 100 T

Hollow markers & dashed spokes: axis not yet scored

G
88
A
—
U
—
N
—
T
n/d
L
—
E
—
T
25.3

Specification

Parameters
35B total / 3B active per token
Architecture
qwen3_5_moe
Size on disk
37.75 GB
Quantization
8bit uniform
Format
MLX
Runtime
LM Studio
Class
Standard
Reasoning (CoT)
Yes — emits reasoning tokens
Internal ID
M9
Mean speed
88.8 tok/s across suites
Stall census
1 stall in 13 observed tests (7.7%)
Reasoning appetite
3,254 tokens mean · 7,730 max
Model card
lmstudio.ai

Suite results

Suite Score Avg / 20 Tests n tok/s Run
General capability (13-task real-workload suite) 1a 228.8 / 260 17.6 13/13 n=2 88.8 run 2026-09-02
Test Category Score tok/s
A1 Dev/Ops Scripting 14/20 89.5
A2 Dev/Ops Scripting 17/20 89.4
A3 Dev/Ops Scripting 20/20 89.3
A4 Dev/Ops Scripting 19/20 91.2
A5 Dev/Ops Scripting 20/20 91.3
A6 Dev/Ops Scripting 20/20 91.1
B1 Document Processing 20/20 88.9
B2 Document Processing 20/20 87.3
B3 Document Processing 0/20 82.9
C1 Content Production 20/20 83.9
C2 Content Production 19/20 82.6
C3 Content Production 20/20 84.2
C4 Content Production 20/20 85.9

Category mean Dev/Ops Scripting 18.3Document Processing 13.3Content Production 19.8

Each test is LLM-judged 0–20 against a fixed rubric; suite max = tests × 20. An ADJ chip marks a cell carrying a subjective adjustment from the maintainer; the judged score is the one shown, is unchanged, and remains what the axes, ranks and tie bands are computed from. Hover for the reason. Rows with a ▶ expand to the per-test breakdown. Where fewer tests ran than the suite total, unrun tests count as zero toward the axis score. See methodology.