← Leaderboard

Qwen3.6-35B-A3B 4bit MLX

Daily-driver trial RAN THE GAUNTLET Rank #5 of 51 · 8/8 GAUNTLET progress
Generalist — 94 / 100 G Agentic — 78 / 100 A Understanding — 93.7 / 100 U Needle — 91.2 / 100 N Thinking — 97.1 / 100 T Live — 76.7 / 100 L Engineering — 96.5 / 100 E Throughput — 30.7 / 100 T
G
94
A
78
U
93.7
N
91.2
T
97.1
L
76.7
E
96.5
T
30.7

Specification

Parameters
35B total / 3B active per token
Architecture
qwen3_5_moe
Size on disk
20.43 GB
Quantization
4bit
Format
MLX
Runtime
LM Studio
Class
Standard
Reasoning (CoT)
Yes — emits reasoning tokens
Internal ID
M8
Mean speed
108 tok/s across suites
Stall census
1 stall in 35 observed tests (2.9%)
Reasoning appetite
2,471 tokens mean · 16,377 max
Model card
lmstudio.ai

Suite results

Suite Score Avg / 20 Tests n tok/s Run
General capability (13-task real-workload suite) 1a 244.4 / 260 18.8 13/13 n=5.3 114.8 run 2026-09-02
Test Category Score tok/s
A1 Dev/Ops Scripting 10/20 110.9
A2 Dev/Ops Scripting 18/20 114.9
A3 Dev/Ops Scripting 20/20 115.3
A4 Dev/Ops Scripting 19/20 116.8
A5 Dev/Ops Scripting 20/20 117.8
A6 Dev/Ops Scripting 20/20 118.3
B1 Document Processing 20/20 115.6
B2 Document Processing 20/20 111.1
B3 Document Processing 20/20 104.3
C1 Content Production 20/20 109.5
C2 Content Production 18/20 101.7
C3 Content Production 18/20 102.1
C4 Content Production 20/20 113

Category mean Dev/Ops Scripting 17.8Document Processing 20Content Production 19

Agentic tool-calling & protocol adherence 1b 124.8 / 160 15.6 8/8 n=4 122.5 run 2026-09-12
Test Category Score tok/s
WA1 Tool Calling & Protocol Adherence 20/20 95.2
WA2 Tool Calling & Protocol Adherence 20/20 106.7
WA3 Tool Calling & Protocol Adherence 20/20 111.1
WA4 Tool Calling & Protocol Adherence 20/20 109
WA5 Tool Calling & Protocol Adherence 5/20 116.5
WB1 Agent Robustness & State Management 20/20 116.6
WB2 Agent Robustness & State Management 20/20 125.9
WB3 Agent Robustness & State Management 0/20 105.6

Category mean Tool Calling & Protocol Adherence 17Agent Robustness & State Management 13.3

Coding depth 1c 173.7 / 180 19.3 9/9 n=1 117.8 run 2026-09-15
Test Category Score tok/s
K1 Coding — Concurrency 20/20 128.7
K2 Coding — Algorithms 20/20 123.9
K3 Coding — Security Review 20/20 120.7
K4 Coding — Performance 18/20 121.6
K5 Coding — Refactoring Under Constraints 20/20 102
K6 Coding — Test Writing 20/20 97.7
K7 Coding — React Component (typed) 17/20 116.8
K8 Coding — TypeScript Type System 20/20 117.2
K9 Coding — React Debugging 19/20 120.4

Category mean Coding — Concurrency 20Coding — Algorithms 20Coding — Security Review 20Coding — Performance 18Coding — Refactoring Under Constraints 20Coding — Test Writing 20Coding — React Component (typed) 17Coding — TypeScript Type System 20Coding — React Debugging 19

Doc/OCR vision 1d 160 / 160 20 8/8 n=1 112.2 run 2026-09-20
Test Category Score tok/s
D1 Doc/OCR — Transcription 20/20 73.3
D2 Doc/OCR — Document QA 20/20 90.8
D3 Doc/OCR — Field Extraction 20/20 74.2
D4 Doc/OCR — Document QA 20/20 105.8
D5 Doc/OCR — Transcription 20/20 71.3
D6 Doc/OCR — Field Extraction 20/20 66.9
D7 Doc/OCR — Degraded Input 20/20 107.3
D8 Doc/OCR — Degraded Input 20/20 89.4

Category mean Doc/OCR — Transcription 20Doc/OCR — Document QA 20Doc/OCR — Field Extraction 20Doc/OCR — Degraded Input 20

Doc/OCR — real-degraded tier 1d2 64.8 / 80 16.2 4/4 n=1 73.9 run 2026-09-20
Test Category Score tok/s
D9 Doc/OCR — Real Degraded Transcription 13/20 46.3
D10 Doc/OCR — Handwriting Extraction 18/20 25.6
D11 Doc/OCR — Degraded Table Transcription 14/20 56
D12 Doc/OCR — Degraded Table QA 20/20 104.3

Category mean Doc/OCR — Real Degraded Transcription 13Doc/OCR — Handwriting Extraction 18Doc/OCR — Degraded Table Transcription 14Doc/OCR — Degraded Table QA 20

Content-production depth 1e 115.2 / 120 19.2 6/6 n=1 127 run 2026-09-16
Test Category Score tok/s
E1 SOW & Requirements Decomposition 20/20 127.9
E2 SOW & Requirements Decomposition 20/20 131.6
E3 Executive & Proposal Content 20/20 125.6
E4 Executive & Proposal Content 19/20 130.7
E5 Executive & Proposal Content 16/20 116.3
E6 Executive & Proposal Content 20/20 121.6

Category mean SOW & Requirements Decomposition 20Executive & Proposal Content 18.8

Long-context retrieval & synthesis 1g 120 / 120 20 6/6 n=1 113.7 run 2026-08-06
Test Category Score tok/s
L1 Long-Context — Needle Retrieval 20/20 88.7
L2 Long-Context — Grounded QA with Distractors 20/20 88.3
L3 Long-Context — Ops Log Reasoning 20/20 70
L4 Long-Context — Faithful Summarization 20/20 90.3
L5 Long-Context — Instruction Retention 20/20 80.8
L6 Long-Context — Full-Haystack QA 20/20 31.4

Category mean Long-Context — Needle Retrieval 20Long-Context — Grounded QA with Distractors 20Long-Context — Ops Log Reasoning 20Long-Context — Faithful Summarization 20Long-Context — Instruction Retention 20Long-Context — Full-Haystack QA 20

Long-context multi-needle (MRCR) 1g2 44.1 / 60 14.7 3/3 n=1 88.3 run 2026-08-14
Test Category Score tok/s
L7 Long-Context — Multi-Needle Retrieval (MRCR) 20/20 65
L8 Long-Context — Multi-Needle Retrieval (MRCR) 20/20 51.3
L9 Long-Context — Multi-Needle Retrieval (MRCR) 4/20 45.2
Live one-shot builds (runtime-verified) 1h 75.2 / 80 18.8 4/4 n=1 115.4 run 2026-08-15
Test Category Score tok/s
H1 Web-Estate One-Shot Builds 18/20 101.1
H2 Web-Estate One-Shot Builds 19/20 121.1
H3 Web-Estate One-Shot Builds 18/20 108.7
H4 Web-Estate One-Shot Builds 20/20 109.8
Production replay (real agent workload) 1i 153.6 / 240 12.8 12/12 n=1 76 run 2026-09-26
Test Category Score tok/s
I1 Nightly Fix Replay 0/20 87.5
I2 Nightly Fix Replay 0/20 67.7
I3 Nightly Fix Replay 8/20 63.3
I4 Nightly Fix Replay 16/20 70.9
I5 Nightly Fix Replay 14/20 58
I6 Nightly Fix Replay 13/20 61.1
I7 Audit Judgment 15/20 82.4
I8 Audit Judgment 20/20 89.7
I9 Cron Dependency Reasoning 19/20 75.5
I10 Cron Dependency Reasoning 17/20 71.5
I11 Curation Judgment 13/20 76.2
I12 Curation Judgment 19/20 87.6

Category mean Nightly Fix Replay 8.5Audit Judgment 17.5Cron Dependency Reasoning 18Curation Judgment 16

Brand-constrained static build (negative-constraint adherence) 1j 93.5 / 100 18.7 5/5 n=3.6 125.9 run 2026-09-13
Test Category Score tok/s
J1 Brand-Constrained Static Build 20/20 122.5
J3 Brand-Constrained Static Build 17/20 122.9
J4 Brand-Constrained Static Build 20/20 130.2
J5 Brand-Constrained Static Build 16/20 124.3
J6 Brand-Constrained Static Build 20/20 122.6

Each test is LLM-judged 0–20 against a fixed rubric; suite max = tests × 20. An ADJ chip marks a cell carrying a subjective adjustment from the maintainer; the judged score is the one shown, is unchanged, and remains what the axes, ranks and tie bands are computed from. Hover for the reason. Rows with a ▶ expand to the per-test breakdown. Where fewer tests ran than the suite total, unrun tests count as zero toward the axis score. See methodology.