← Leaderboard

Qwen3.6-35B-A3B 4bit MLX

Keeper Rank #6 of 33 · 6/8 GAUNTLET progress
Generalist — 93 / 100 G Agentic — 87.5 / 100 A Understanding — pending U Needle — 91.2 / 100 N Thinking — 97.1 / 100 T Live — 93.8 / 100 L Engineering — pending E Throughput — 54.7 / 100 T

Hollow markers & dashed spokes: axis not yet scored

G
93
A
87.5
U
N
91.2
T
97.1
L
93.8
E
T
54.7

Specification

Parameters
35B total / 3B active per token
Architecture
qwen3_5_moe
Size on disk
20.43 GB
Quantization
4bit
Format
MLX
Reasoning (CoT)
Yes — emits reasoning tokens
Internal ID
M8
Mean speed
89.4 tok/s across suites
Stall census
1 stall in 35 observed tests (2.9%)
Reasoning appetite
2,471 tokens mean · 16,377 max
Model card
lmstudio.ai

Suite results

Suite Score Avg / 20 Tests tok/s
General capability (13-task real-workload suite) 1a 241.8 / 260 18.6 13/13 119.9
Agentic tool-calling & protocol adherence 1b 140 / 160 17.5 8/8 108.9
Long-context retrieval & synthesis 1g 120 / 120 20 6/6 74.9
Long-context multi-needle (MRCR) 1g2 44.1 / 60 14.7 3/3 53.8
Live one-shot builds (runtime-verified) 1h 75 / 80 18.8 4/4

Each test is LLM-judged 0–20 against a fixed rubric; suite max = tests × 20. Where fewer tests ran than the suite total, unrun tests count as zero toward the axis score. See methodology.