← Leaderboard

Qwen3.6-35B-A3B Unsloth Dynamic UD-Q8_K_XL MLX (Brooooooklyn)

Keeper RAN THE GAUNTLET Rank #6 of 51 · 8/8 GAUNTLET progress
Generalist — 93.5 / 100 G Agentic — 96 / 100 A Understanding — 79.7 / 100 U Needle — 91.2 / 100 N Thinking — 100 / 100 T Live — 83.3 / 100 L Engineering — 98 / 100 E Throughput — 23.7 / 100 T
G
93.5
A
96
U
79.7
N
91.2
T
100
L
83.3
E
98
T
23.7

Specification

Parameters
35B total / 3B active per token
Architecture
qwen3_5_moe
Size on disk
38.06 GB
Quantization
UD-Q8_K_XL mixed-precision
Format
MLX
Runtime
LM Studio
Class
Standard
Reasoning (CoT)
Yes — emits reasoning tokens
Internal ID
M9b
Mean speed
83.2 tok/s across suites
Stall census
0 stalls in 30 observed tests (0.0%)
Reasoning appetite
2,240 tokens mean · 6,384 max
Model card
huggingface.co

Suite results

Suite Score Avg / 20 Tests n tok/s Run
General capability (13-task real-workload suite) 1a 243.1 / 260 18.7 13/13 n=5.4 83.6 run 2026-09-02
Test Category Score tok/s
A1 Dev/Ops Scripting 12.5/20 81
A2 Dev/Ops Scripting 17.6/20 83
A3 Dev/Ops Scripting 20/20 83.7
A4 Dev/Ops Scripting 19/20 84.4
A5 Dev/Ops Scripting 16.8/20 86.5
A6 Dev/Ops Scripting 20/20 87.1
B1 Document Processing 20/20 83.6
B2 Document Processing 20/20 81.8
B3 Document Processing 19.8/20 73.3
C1 Content Production 20/20 71.3
C2 Content Production 17.6/20 76.9
C3 Content Production 20/20 71.8
C4 Content Production 19.8/20 79.8

Category mean Dev/Ops Scripting 17.7Document Processing 19.9Content Production 19.4

Agentic tool-calling & protocol adherence 1b 153.6 / 160 19.2 8/8 n=1 94.9 run 2026-07-28
Test Category Score tok/s
WA1 Tool Calling & Protocol Adherence 20/20 49.1
WA2 Tool Calling & Protocol Adherence 20/20 84.4
WA3 Tool Calling & Protocol Adherence 20/20 73.1
WA4 Tool Calling & Protocol Adherence 20/20 76.9
WA5 Tool Calling & Protocol Adherence 20/20 90.9
WB1 Agent Robustness & State Management 20/20 91.3
WB2 Agent Robustness & State Management 20/20 91.1
WB3 Agent Robustness & State Management 14/20 93.4

Category mean Tool Calling & Protocol Adherence 20Agent Robustness & State Management 18

Coding depth 1c 176.4 / 180 19.6 9/9 n=1 85.8 run 2026-09-16
Test Category Score tok/s
K1 Coding — Concurrency 20/20 86.1
K2 Coding — Algorithms 20/20 88.2
K3 Coding — Security Review 19/20 82.9
K4 Coding — Performance 18/20 85.3
K5 Coding — Refactoring Under Constraints 20/20 82.7
K6 Coding — Test Writing 20/20 81.8
K7 Coding — React Component (typed) 20/20 85.1
K8 Coding — TypeScript Type System 19/20 84.5
K9 Coding — React Debugging 20/20 86.3

Category mean Coding — Concurrency 20Coding — Algorithms 20Coding — Security Review 19Coding — Performance 18Coding — Refactoring Under Constraints 20Coding — Test Writing 20Coding — React Component (typed) 20Coding — TypeScript Type System 19Coding — React Debugging 20

Doc/OCR vision 1d 160 / 160 20 8/8 n=1 87.1 run 2026-09-20
Test Category Score tok/s
D1 Doc/OCR — Transcription 20/20 63.6
D2 Doc/OCR — Document QA 20/20 64.5
D3 Doc/OCR — Field Extraction 20/20 63.5
D4 Doc/OCR — Document QA 20/20 74.7
D5 Doc/OCR — Transcription 20/20 60.5
D6 Doc/OCR — Field Extraction 20/20 61.8
D7 Doc/OCR — Degraded Input 20/20 77.4
D8 Doc/OCR — Degraded Input 20/20 78.2

Category mean Doc/OCR — Transcription 20Doc/OCR — Document QA 20Doc/OCR — Field Extraction 20Doc/OCR — Degraded Input 20

Doc/OCR — real-degraded tier 1d2 31.2 / 80 7.8 4/4 n=1 83.8 run 2026-09-20
Test Category Score tok/s
D9 Doc/OCR — Real Degraded Transcription 0/20 —
D10 Doc/OCR — Handwriting Extraction 19/20 64.2
D11 Doc/OCR — Degraded Table Transcription 0/20 80.5
D12 Doc/OCR — Degraded Table QA 12/20 75.1

Category mean Doc/OCR — Real Degraded Transcription 0Doc/OCR — Handwriting Extraction 19Doc/OCR — Degraded Table Transcription 0Doc/OCR — Degraded Table QA 12

Content-production depth 1e 114 / 120 19 6/6 n=1 87.2 run 2026-09-16
Test Category Score tok/s
E1 SOW & Requirements Decomposition 20/20 92.8
E2 SOW & Requirements Decomposition 20/20 91
E3 Executive & Proposal Content 20/20 79.3
E4 Executive & Proposal Content 20/20 82.6
E5 Executive & Proposal Content 19/20 85.5
E6 Executive & Proposal Content 15/20 84.7

Category mean SOW & Requirements Decomposition 20Executive & Proposal Content 18.5

Long-context retrieval & synthesis 1g 120 / 120 20 6/6 n=1 85.8 run 2026-08-06
Test Category Score tok/s
L1 Long-Context — Needle Retrieval 20/20 56.2
L2 Long-Context — Grounded QA with Distractors 20/20 70.1
L3 Long-Context — Ops Log Reasoning 20/20 63.3
L4 Long-Context — Faithful Summarization 20/20 69.9
L5 Long-Context — Instruction Retention 20/20 65.9
L6 Long-Context — Full-Haystack QA 20/20 33.3

Category mean Long-Context — Needle Retrieval 20Long-Context — Grounded QA with Distractors 20Long-Context — Ops Log Reasoning 20Long-Context — Faithful Summarization 20Long-Context — Instruction Retention 20Long-Context — Full-Haystack QA 20

Long-context multi-needle (MRCR) 1g2 44.1 / 60 14.7 3/3 n=1 63.9 run 2026-08-14
Test Category Score tok/s
L7 Long-Context — Multi-Needle Retrieval (MRCR) 20/20 60.5
L8 Long-Context — Multi-Needle Retrieval (MRCR) 20/20 44.4
L9 Long-Context — Multi-Needle Retrieval (MRCR) 4/20 43.7
Live one-shot builds (runtime-verified) 1h 70 / 80 17.5 4/4 n=1 73.3 run 2026-09-09
Test Category Score tok/s
H1 Web-Estate One-Shot Builds 16/20 76.4
H2 Web-Estate One-Shot Builds 20/20 75
H3 Web-Estate One-Shot Builds 17/20 71
H4 Web-Estate One-Shot Builds 17/20 70.8
Brand-constrained static build (negative-constraint adherence) 1j 80 / 100 16 5/5 n=1 86.8 run 2026-09-16
Test Category Score tok/s
J1 Brand-Constrained Static Build 20/20 86.1
J3 Brand-Constrained Static Build 20/20 88.3
J4 Brand-Constrained Static Build 20/20 89.2
J5 Brand-Constrained Static Build 17/20 85.9
J6 Brand-Constrained Static Build 3/20 80.3

Each test is LLM-judged 0–20 against a fixed rubric; suite max = tests × 20. An ADJ chip marks a cell carrying a subjective adjustment from the maintainer; the judged score is the one shown, is unchanged, and remains what the axes, ranks and tie bands are computed from. Hover for the reason. Rows with a ▶ expand to the per-test breakdown. Where fewer tests ran than the suite total, unrun tests count as zero toward the axis score. See methodology.