← Leaderboard

Qwen3.5-122B-A10B 4bit MLX

Keeper RAN THE GAUNTLET Rank #8 of 51 · 8/8 GAUNTLET progress
Generalist — 90 / 100 G Agentic — 94.5 / 100 A Understanding — 100 / 100 U Needle — 91.2 / 100 N Thinking — 100 / 100 T Live — 84.4 / 100 L Engineering — 98.5 / 100 E Throughput — 14.2 / 100 T
G
90
A
94.5
U
100
N
91.2
T
100
L
84.4
E
98.5
T
14.2

Specification

Parameters
122B total / 10B active per token
Architecture
qwen3_5_moe
Size on disk
69.62 GB
Quantization
4bit
Format
MLX
Runtime
LM Studio
Class
Standard
Reasoning (CoT)
Yes — emits reasoning tokens
Internal ID
M10
Mean speed
50 tok/s across suites
Stall census
0 stalls in 30 observed tests (0.0%)
Reasoning appetite
2,043 tokens mean · 9,288 max
Model card
huggingface.co

Suite results

Suite Score Avg / 20 Tests n tok/s Run
General capability (13-task real-workload suite) 1a 234 / 260 18 13/13 n=4 54.9 run 2026-09-02
Test Category Score tok/s
A1 Dev/Ops Scripting 14/20 55.8
A2 Dev/Ops Scripting 18/20 54.2
A3 Dev/Ops Scripting 20/20 53
A4 Dev/Ops Scripting 11/20 56.9
A5 Dev/Ops Scripting 14/20 51.6
A6 Dev/Ops Scripting 20/20 51.3
B1 Document Processing 20/20 54.9
B2 Document Processing 20/20 54.4
B3 Document Processing 20/20 50.2
C1 Content Production 20/20 50.9
C2 Content Production 17/20 52.6
C3 Content Production 20/20 54.8
C4 Content Production 20/20 55.6

Category mean Dev/Ops Scripting 16.2Document Processing 20Content Production 19.3

Agentic tool-calling & protocol adherence 1b 151.2 / 160 18.9 8/8 n=1 65.2 run 2026-07-28
Test Category Score tok/s
WA1 Tool Calling & Protocol Adherence 20/20 15.6
WA2 Tool Calling & Protocol Adherence 20/20 41.7
WA3 Tool Calling & Protocol Adherence 20/20 41.2
WA4 Tool Calling & Protocol Adherence 20/20 32.7
WA5 Tool Calling & Protocol Adherence 20/20 50.7
WB1 Agent Robustness & State Management 20/20 60.4
WB2 Agent Robustness & State Management 20/20 58
WB3 Agent Robustness & State Management 11/20 49.7

Category mean Tool Calling & Protocol Adherence 20Agent Robustness & State Management 17

Coding depth 1c 177.3 / 180 19.7 9/9 n=1 52.6 run 2026-09-17
Test Category Score tok/s
K1 Coding — Concurrency 20/20 55.6
K2 Coding — Algorithms 20/20 54.3
K3 Coding — Security Review 20/20 52.1
K4 Coding — Performance 18/20 51.5
K5 Coding — Refactoring Under Constraints 20/20 51.6
K6 Coding — Test Writing 20/20 48.4
K7 Coding — React Component (typed) 19/20 50.5
K8 Coding — TypeScript Type System 20/20 48.5
K9 Coding — React Debugging 20/20 50.9

Category mean Coding — Concurrency 20Coding — Algorithms 20Coding — Security Review 20Coding — Performance 18Coding — Refactoring Under Constraints 20Coding — Test Writing 20Coding — React Component (typed) 19Coding — TypeScript Type System 20Coding — React Debugging 20

Doc/OCR vision 1d 160 / 160 20 8/8 n=1 22.1 run 2026-09-20
Test Category Score tok/s
D1 Doc/OCR — Transcription 20/20 17.3
D2 Doc/OCR — Document QA 20/20 19.6
D3 Doc/OCR — Field Extraction 20/20 19.4
D4 Doc/OCR — Document QA 20/20 21.1
D5 Doc/OCR — Transcription 20/20 18
D6 Doc/OCR — Field Extraction 20/20 19.5
D7 Doc/OCR — Degraded Input 20/20 22.4
D8 Doc/OCR — Degraded Input 20/20 13.3

Category mean Doc/OCR — Transcription 20Doc/OCR — Document QA 20Doc/OCR — Field Extraction 20Doc/OCR — Degraded Input 20

Content-production depth 1e 114 / 120 19 6/6 n=1 52.4 run 2026-09-16
Test Category Score tok/s
E1 SOW & Requirements Decomposition 20/20 49.2
E2 SOW & Requirements Decomposition 18/20 52.2
E3 Executive & Proposal Content 18/20 52.5
E4 Executive & Proposal Content 18/20 52.8
E5 Executive & Proposal Content 20/20 51.8
E6 Executive & Proposal Content 20/20 49.6

Category mean SOW & Requirements Decomposition 19Executive & Proposal Content 19

Long-context retrieval & synthesis 1g 120 / 120 20 6/6 n=1 52.3 run 2026-08-06
Test Category Score tok/s
L1 Long-Context — Needle Retrieval 20/20 53.5
L2 Long-Context — Grounded QA with Distractors 20/20 41.6
L3 Long-Context — Ops Log Reasoning 20/20 33.5
L4 Long-Context — Faithful Summarization 20/20 36.9
L5 Long-Context — Instruction Retention 20/20 31.9
L6 Long-Context — Full-Haystack QA 20/20 17.8

Category mean Long-Context — Needle Retrieval 20Long-Context — Grounded QA with Distractors 20Long-Context — Ops Log Reasoning 20Long-Context — Faithful Summarization 20Long-Context — Instruction Retention 20Long-Context — Full-Haystack QA 20

Long-context multi-needle (MRCR) 1g2 44.1 / 60 14.7 3/3 n=1 38.2 run 2026-08-14
Test Category Score tok/s
L7 Long-Context — Multi-Needle Retrieval (MRCR) 20/20 31.1
L8 Long-Context — Multi-Needle Retrieval (MRCR) 20/20 17.1
L9 Long-Context — Multi-Needle Retrieval (MRCR) 4/20 24.7
Live one-shot builds (runtime-verified) 1h 66 / 80 16.5 4/4 n=1 57.9 run 2026-09-03
Test Category Score tok/s
H1 Web-Estate One-Shot Builds 19/20 60.8
H2 Web-Estate One-Shot Builds 16/20 58.7
H3 Web-Estate One-Shot Builds 14/20 53.9
H4 Web-Estate One-Shot Builds 17/20 53.8
Brand-constrained static build (negative-constraint adherence) 1j 86 / 100 17.2 5/5 n=1 54.2 run 2026-09-16
Test Category Score tok/s
J1 Brand-Constrained Static Build 19/20 57
J3 Brand-Constrained Static Build 17/20 56.8
J4 Brand-Constrained Static Build 20/20 50.9
J5 Brand-Constrained Static Build 17/20 51.5
J6 Brand-Constrained Static Build 13/20 49.1

Each test is LLM-judged 0–20 against a fixed rubric; suite max = tests × 20. An ADJ chip marks a cell carrying a subjective adjustment from the maintainer; the judged score is the one shown, is unchanged, and remains what the axes, ranks and tie bands are computed from. Hover for the reason. Rows with a ▶ expand to the per-test breakdown. Where fewer tests ran than the suite total, unrun tests count as zero toward the axis score. See methodology.