← Leaderboard

Qwen3.5-122B-A10B 4bit MLX

Keeper Rank #9 of 33 · 5/8 GAUNTLET progress
Generalist — 92.5 / 100 G Agentic — 95 / 100 A Understanding — pending U Needle — 90.2 / 100 N Thinking — 100 / 100 T Live — pending L Engineering — pending E Throughput — 24.1 / 100 T

Hollow markers & dashed spokes: axis not yet scored

G
92.5
A
95
U
N
90.2
T
100
L
E
T
24.1

Specification

Parameters
122B total / 10B active per token
Architecture
qwen3_5_moe
Size on disk
69.62 GB
Quantization
4bit
Format
MLX
Reasoning (CoT)
Yes — emits reasoning tokens
Internal ID
M10
Mean speed
39.4 tok/s across suites
Stall census
0 stalls in 30 observed tests (0.0%)
Reasoning appetite
2,043 tokens mean · 9,288 max
Model card
huggingface.co

Suite results

Suite Score Avg / 20 Tests tok/s
General capability (13-task real-workload suite) 1a 240.5 / 260 18.5 13/13 53.6
Agentic tool-calling & protocol adherence 1b 152 / 160 19 8/8 43.8
Long-context retrieval & synthesis 1g 118.2 / 120 19.7 6/6 35.9
Long-context multi-needle (MRCR) 1g2 44.1 / 60 14.7 3/3 24.3

Each test is LLM-judged 0–20 against a fixed rubric; suite max = tests × 20. Where fewer tests ran than the suite total, unrun tests count as zero toward the axis score. See methodology.