← Leaderboard

Qwen3-4B-2507 (non-thinking) 4bit MLXSMALL

Tested Rank #32 of 51 · 6/8 GAUNTLET progress
Generalist — 68 / 100 G Agentic — 93 / 100 A Understanding — pending U Needle — 80 / 100 N Thinking — pending T Live — 46.6 / 100 L Engineering — 56 / 100 E Throughput — 38 / 100 T

Hollow markers & dashed spokes: axis not yet scored

G
68
A
93
U
—
N
80
T
n/d
L
46.6
E
56
T
38

Specification

Parameters
4B
Architecture
qwen3
Size on disk
2.28 GB
Quantization
4bit
Format
MLX
Runtime
LM Studio
Class
Small (under 8B total parameters — phone/edge class)
Reasoning (CoT)
No
Internal ID
M18
Mean speed
133.6 tok/s across suites
Stall census
0 stalls in 14 observed tests (0.0%)
Model card
lmstudio.ai

Suite results

Suite Score Avg / 20 Tests n tok/s Run
General capability (13-task real-workload suite) 1a 176.8 / 260 13.6 13/13 n=1.1 144.3 run 2026-08-04
Test Category Score tok/s
A1 Dev/Ops Scripting 3/20 105.9
A2 Dev/Ops Scripting 12/20 158.6
A3 Dev/Ops Scripting 9/20 121.4
A4 Dev/Ops Scripting 6/20 141.2
A5 Dev/Ops Scripting 15/20 135.3
A6 Dev/Ops Scripting 13/20 141.6
B1 Document Processing 20/20 148.6
B2 Document Processing 20/20 153.8
B3 Document Processing 8/20 99.2
C1 Content Production 17/20 145.3
C2 Content Production 18/20 143.4
C3 Content Production 18/20 145.1
C4 Content Production 17/20 148.9

Category mean Dev/Ops Scripting 9.7Document Processing 16Content Production 17.5

Agentic tool-calling & protocol adherence 1b 148.8 / 160 18.6 8/8 n=1 177.8 run 2026-09-21
Test Category Score tok/s
WA1 Tool Calling & Protocol Adherence 20/20 88
WA2 Tool Calling & Protocol Adherence 20/20 86.6
WA3 Tool Calling & Protocol Adherence 20/20 104.3
WA4 Tool Calling & Protocol Adherence 20/20 79.1
WA5 Tool Calling & Protocol Adherence 20/20 143.3
WB1 Agent Robustness & State Management 18/20 165
WB2 Agent Robustness & State Management 20/20 167.3
WB3 Agent Robustness & State Management 11/20 158.9

Category mean Tool Calling & Protocol Adherence 20Agent Robustness & State Management 16.3

Coding depth 1c 100.8 / 180 11.2 9/9 n=1 155.4 run 2026-09-30
Test Category Score tok/s
K1 Coding — Concurrency 16/20 149.5
K2 Coding — Algorithms 13/20 166.7
K3 Coding — Security Review 9/20 167
K4 Coding — Performance 16/20 162
K5 Coding — Refactoring Under Constraints 15/20 154.6
K6 Coding — Test Writing 16/20 162.5
K7 Coding — React Component (typed) 6/20 158.4
K8 Coding — TypeScript Type System 9/20 153.4
K9 Coding — React Debugging 1/20 81.2

Category mean Coding — Concurrency 16Coding — Algorithms 13Coding — Security Review 9Coding — Performance 16Coding — Refactoring Under Constraints 15Coding — Test Writing 16Coding — React Component (typed) 6Coding — TypeScript Type System 9Coding — React Debugging 1

Content-production depth 1e 94.8 / 120 15.8 6/6 n=1 168 run 2026-09-25
Test Category Score tok/s
E1 SOW & Requirements Decomposition 20/20 160.8
E2 SOW & Requirements Decomposition 20/20 154.7
E3 Executive & Proposal Content 17/20 152.6
E4 Executive & Proposal Content 15/20 166
E5 Executive & Proposal Content 15/20 155.5
E6 Executive & Proposal Content 8/20 156.5

Category mean SOW & Requirements Decomposition 20Executive & Proposal Content 13.8

Long-context retrieval & synthesis 1g 120 / 120 20 6/6 n=1 73.8 run 2026-09-23
Test Category Score tok/s
L1 Long-Context — Needle Retrieval 20/20 20.9
L2 Long-Context — Grounded QA with Distractors 20/20 7.2
L3 Long-Context — Ops Log Reasoning 20/20 15.6
L4 Long-Context — Faithful Summarization 20/20 32.9
L5 Long-Context — Instruction Retention 20/20 7.5
L6 Long-Context — Full-Haystack QA 20/20 1.8

Category mean Long-Context — Needle Retrieval 20Long-Context — Grounded QA with Distractors 20Long-Context — Ops Log Reasoning 20Long-Context — Faithful Summarization 20Long-Context — Instruction Retention 20Long-Context — Full-Haystack QA 20

Long-context multi-needle (MRCR) 1g2 24 / 60 8 3/3 n=1 48.8 run 2026-09-19
Test Category Score tok/s
L7 Long-Context — Multi-Needle Retrieval (MRCR) 10/20 30.2
L8 Long-Context — Multi-Needle Retrieval (MRCR) 10/20 11.6
L9 Long-Context — Multi-Needle Retrieval (MRCR) 4/20 5
Live one-shot builds (runtime-verified) 1h 32.8 / 80 8.2 4/4 n=1 139 run 2026-09-19
Test Category Score tok/s
H1 Web-Estate One-Shot Builds 12/20 119.3
H2 Web-Estate One-Shot Builds 12/20 143.1
H3 Web-Estate One-Shot Builds 7/20 143.7
H4 Web-Estate One-Shot Builds 2/20 142.9
Brand-constrained static build (negative-constraint adherence) 1j 51 / 100 10.2 5/5 n=1 161.4 run 2026-09-22
Test Category Score tok/s
J1 Brand-Constrained Static Build 10/20 153.4
J3 Brand-Constrained Static Build 8/20 155.9
J4 Brand-Constrained Static Build 12/20 147.6
J5 Brand-Constrained Static Build 13/20 157
J6 Brand-Constrained Static Build 8/20 154.9

Each test is LLM-judged 0–20 against a fixed rubric; suite max = tests × 20. An ADJ chip marks a cell carrying a subjective adjustment from the maintainer; the judged score is the one shown, is unchanged, and remains what the axes, ranks and tie bands are computed from. Hover for the reason. Rows with a ▶ expand to the per-test breakdown. Where fewer tests ran than the suite total, unrun tests count as zero toward the axis score. See methodology.