← Leaderboard

Qwen3-Coder-Next 80B 4bit MLX

Keeper Rank #19 of 51 · 7/8 GAUNTLET progress
Generalist — 88.5 / 100 G Agentic — 94.5 / 100 A Understanding — pending U Needle — 64.4 / 100 N Thinking — 100 / 100 T Live — 87.8 / 100 L Engineering — 89.5 / 100 E Throughput — 23.3 / 100 T

Hollow markers & dashed spokes: axis not yet scored

G
88.5
A
94.5
U
—
N
64.4
T
100
L
87.8
E
89.5
T
23.3

Specification

Parameters
80B
Architecture
qwen3_next
Size on disk
44.86 GB
Quantization
4bit
Format
MLX
Runtime
LM Studio
Class
Standard
Reasoning (CoT)
No
Internal ID
M6
Mean speed
81.8 tok/s across suites
Stall census
0 stalls in 46 observed tests (0.0%)
Model card
lmstudio.ai

Suite results

Suite Score Avg / 20 Tests n tok/s Run
General capability (13-task real-workload suite) 1a 230.1 / 260 17.7 13/13 n=5 86.5 run 2026-09-02
Test Category Score tok/s
A1 Dev/Ops Scripting 8.6/20 76.3
A2 Dev/Ops Scripting 16.2/20 76.2
A3 Dev/Ops Scripting 18.4/20 74.6
A4 Dev/Ops Scripting 19/20 70.7
A5 Dev/Ops Scripting 20/20 75.6
A6 Dev/Ops Scripting 20/20 75.4
B1 Document Processing 20/20 70.8
B2 Document Processing 19.2/20 71.8
B3 Document Processing 18.8/20 76.5
C1 Content Production 18.4/20 77.4
C2 Content Production 17.6/20 73.8
C3 Content Production 17/20 77.5
C4 Content Production 16.6/20 74.1

Category mean Dev/Ops Scripting 17Document Processing 19.3Content Production 17.4

Agentic tool-calling & protocol adherence 1b 151.2 / 160 18.9 8/8 n=1 89.6 run 2026-07-28
Test Category Score tok/s
WA1 Tool Calling & Protocol Adherence 20/20 4.6
WA2 Tool Calling & Protocol Adherence 20/20 31.4
WA3 Tool Calling & Protocol Adherence 20/20 35.8
WA4 Tool Calling & Protocol Adherence 20/20 26
WA5 Tool Calling & Protocol Adherence 20/20 61.6
WB1 Agent Robustness & State Management 20/20 81.7
WB2 Agent Robustness & State Management 20/20 79
WB3 Agent Robustness & State Management 11/20 49.1

Category mean Tool Calling & Protocol Adherence 20Agent Robustness & State Management 17

Coding depth 1c 161.1 / 180 17.9 9/9 n=1 78 run 2026-09-15
Test Category Score tok/s
K1 Coding — Concurrency 19/20 62.5
K2 Coding — Algorithms 20/20 67
K3 Coding — Security Review 13/20 70.8
K4 Coding — Performance 20/20 71.4
K5 Coding — Refactoring Under Constraints 20/20 74.9
K6 Coding — Test Writing 18/20 69.2
K7 Coding — React Component (typed) 16/20 79.3
K8 Coding — TypeScript Type System 19/20 68.6
K9 Coding — React Debugging 16/20 80.7

Category mean Coding — Concurrency 19Coding — Algorithms 20Coding — Security Review 13Coding — Performance 20Coding — Refactoring Under Constraints 20Coding — Test Writing 18Coding — React Component (typed) 16Coding — TypeScript Type System 19Coding — React Debugging 16

Content-production depth 1e 100.2 / 120 16.7 6/6 n=1 87.3 run 2026-08-02
Test Category Score tok/s
E1 SOW & Requirements Decomposition 20/20 81.9
E2 SOW & Requirements Decomposition 16/20 67.2
E3 Executive & Proposal Content 17/20 84.4
E4 Executive & Proposal Content 15/20 78.5
E5 Executive & Proposal Content 17/20 73.1
E6 Executive & Proposal Content 15/20 76.5

Category mean SOW & Requirements Decomposition 18Executive & Proposal Content 16

Long-context retrieval & synthesis 1g 100 / 120 20 5/6 ● n=1 81.6 run 2026-09-17
Test Category Score tok/s
L1 Long-Context — Needle Retrieval 20/20 7.8
L2 Long-Context — Grounded QA with Distractors 20/20 4.7
L3 Long-Context — Ops Log Reasoning 20/20 16.2
L4 Long-Context — Faithful Summarization 20/20 29.8
L5 Long-Context — Instruction Retention 20/20 7.3

Category mean Long-Context — Needle Retrieval 20Long-Context — Grounded QA with Distractors 20Long-Context — Ops Log Reasoning 20Long-Context — Faithful Summarization 20Long-Context — Instruction Retention 20

Long-context multi-needle (MRCR) 1g2 15.9 / 60 5.3 3/3 n=1 69.3 run 2026-09-02
Test Category Score tok/s
L7 Long-Context — Multi-Needle Retrieval (MRCR) 7/20 28.8
L8 Long-Context — Multi-Needle Retrieval (MRCR) 5/20 16.8
L9 Long-Context — Multi-Needle Retrieval (MRCR) 4/20 9.5
Live one-shot builds (runtime-verified) 1h 68 / 80 17 4/4 n=1 74.3 run 2026-09-04
Test Category Score tok/s
H1 Web-Estate One-Shot Builds 19/20 62
H2 Web-Estate One-Shot Builds 19/20 83.8
H3 Web-Estate One-Shot Builds 17/20 64.7
H4 Web-Estate One-Shot Builds 13/20 77.2
Brand-constrained static build (negative-constraint adherence) 1j 90 / 100 18 5/5 n=1 88.2 run 2026-09-16
Test Category Score tok/s
J1 Brand-Constrained Static Build 20/20 85.1
J3 Brand-Constrained Static Build 20/20 83.7
J4 Brand-Constrained Static Build 17/20 73.9
J5 Brand-Constrained Static Build 17/20 86.4
J6 Brand-Constrained Static Build 16/20 84.2

Each test is LLM-judged 0–20 against a fixed rubric; suite max = tests × 20. An ADJ chip marks a cell carrying a subjective adjustment from the maintainer; the judged score is the one shown, is unchanged, and remains what the axes, ranks and tie bands are computed from. Hover for the reason. Rows with a ▶ expand to the per-test breakdown. Where fewer tests ran than the suite total, unrun tests count as zero toward the axis score. See methodology.