← Leaderboard

Qwen3.6-27B Dense 8bit MLX

Keeper RAN THE GAUNTLET Rank #7 of 51 · 8/8 GAUNTLET progress
Generalist — 90.5 / 100 G Agentic — 100 / 100 A Understanding — 96.3 / 100 U Needle — 100 / 100 N Thinking — 97.2 / 100 T Live — 79 / 100 L Engineering — 77.8 / 100 E Throughput — 4.2 / 100 T
G
90.5
A
100
U
96.3
N
100
T
97.2
L
79
E
77.8
T
4.2

Specification

Parameters
27B dense
Architecture
qwen3_5
Size on disk
29.53 GB
Quantization
8bit
Format
MLX
Runtime
LM Studio
Class
Standard
Reasoning (CoT)
Yes — emits reasoning tokens
Internal ID
M7
Mean speed
14.6 tok/s across suites
Stall census
2 stalls in 71 observed tests (2.8%)
Reasoning appetite
2,881 tokens mean · 16,310 max
Model card
lmstudio.ai

Suite results

Suite Score Avg / 20 Tests n tok/s Run
General capability (13-task real-workload suite) 1a 235.3 / 260 18.1 13/13 n=2.8 15.2 run 2026-09-02
Test Category Score tok/s
A1 Dev/Ops Scripting 20/20 16.7
A2 Dev/Ops Scripting 19/20 14.6
A3 Dev/Ops Scripting 0/20 15
A4 Dev/Ops Scripting 18/20 15.5
A5 Dev/Ops Scripting 19/20 15.6
A6 Dev/Ops Scripting 20/20 15.7
B1 Document Processing 20/20 15.2
B2 Document Processing 20/20 14.6
B3 Document Processing 20/20 14.5
C1 Content Production 20/20 15.2
C2 Content Production 20/20 15
C3 Content Production 20/20 14.9
C4 Content Production 20/20 15.3

Category mean Dev/Ops Scripting 16Document Processing 20Content Production 20

Agentic tool-calling & protocol adherence 1b 160 / 160 20 8/8 n=1 17.6 run 2026-07-28
Test Category Score tok/s
WA1 Tool Calling & Protocol Adherence 20/20 14.6
WA2 Tool Calling & Protocol Adherence 20/20 16.6
WA3 Tool Calling & Protocol Adherence 20/20 16.8
WA4 Tool Calling & Protocol Adherence 20/20 16.2
WA5 Tool Calling & Protocol Adherence 20/20 17.5
WB1 Agent Robustness & State Management 20/20 16.9
WB2 Agent Robustness & State Management 20/20 16.9
WB3 Agent Robustness & State Management 20/20 17.2

Category mean Tool Calling & Protocol Adherence 20Agent Robustness & State Management 20

Coding depth 1c 140 / 180 20 7/9 ● n=1 13.9 run 2026-09-15
Test Category Score tok/s
K1 Coding — Concurrency 20/20 15
K2 Coding — Algorithms 20/20 14.9
K3 Coding — Security Review 20/20 13.2
K4 Coding — Performance 20/20 12.6
K5 Coding — Refactoring Under Constraints 20/20 13.8
K6 Coding — Test Writing 20/20 14
K8 Coding — TypeScript Type System 20/20 16.7

Category mean Coding — Concurrency 20Coding — Algorithms 20Coding — Security Review 20Coding — Performance 20Coding — Refactoring Under Constraints 20Coding — Test Writing 20Coding — TypeScript Type System 20

Doc/OCR vision 1d 160 / 160 20 8/8 n=1 16.2 run 2026-08-06
Test Category Score tok/s
D1 Doc/OCR — Transcription 20/20 15
D2 Doc/OCR — Document QA 20/20 13.6
D3 Doc/OCR — Field Extraction 20/20 13.5
D4 Doc/OCR — Document QA 20/20 14.4
D5 Doc/OCR — Transcription 20/20 13.8
D6 Doc/OCR — Field Extraction 20/20 12.3
D7 Doc/OCR — Degraded Input 20/20 15.2
D8 Doc/OCR — Degraded Input 20/20 15.1

Category mean Doc/OCR — Transcription 20Doc/OCR — Document QA 20Doc/OCR — Field Extraction 20Doc/OCR — Degraded Input 20

Doc/OCR — real-degraded tier 1d2 71.2 / 80 17.8 4/4 n=1 14.3 run 2026-08-14
Test Category Score tok/s
D9 Doc/OCR — Real Degraded Transcription 17/20 15.7
D10 Doc/OCR — Handwriting Extraction 17/20 10.9
D11 Doc/OCR — Degraded Table Transcription 17/20 13.3
D12 Doc/OCR — Degraded Table QA 20/20 12.7

Category mean Doc/OCR — Real Degraded Transcription 17Doc/OCR — Handwriting Extraction 17Doc/OCR — Degraded Table Transcription 17Doc/OCR — Degraded Table QA 20

Content-production depth 1e 115.2 / 120 19.2 6/6 n=1 16.4 run 2026-08-02
Test Category Score tok/s
E1 SOW & Requirements Decomposition 20/20 16.9
E2 SOW & Requirements Decomposition 20/20 16.3
E3 Executive & Proposal Content 18/20 16.2
E4 Executive & Proposal Content 18/20 16.8
E5 Executive & Proposal Content 19/20 16.4
E6 Executive & Proposal Content 20/20 15.7

Category mean SOW & Requirements Decomposition 20Executive & Proposal Content 18.8

Long-context retrieval & synthesis 1g 120 / 120 20 6/6 n=1 14.6 run 2026-08-06
Test Category Score tok/s
L1 Long-Context — Needle Retrieval 20/20 10.5
L2 Long-Context — Grounded QA with Distractors 20/20 10.3
L3 Long-Context — Ops Log Reasoning 20/20 11.2
L4 Long-Context — Faithful Summarization 20/20 12.6
L5 Long-Context — Instruction Retention 20/20 11.8
L6 Long-Context — Full-Haystack QA 20/20 6.6

Category mean Long-Context — Needle Retrieval 20Long-Context — Grounded QA with Distractors 20Long-Context — Ops Log Reasoning 20Long-Context — Faithful Summarization 20Long-Context — Instruction Retention 20Long-Context — Full-Haystack QA 20

Long-context multi-needle (MRCR) 1g2 60 / 60 20 3/3 n=1 12.4 run 2026-08-14
Test Category Score tok/s
L7 Long-Context — Multi-Needle Retrieval (MRCR) 20/20 10.9
L8 Long-Context — Multi-Needle Retrieval (MRCR) 20/20 8.2
L9 Long-Context — Multi-Needle Retrieval (MRCR) 20/20 8.4
Live one-shot builds (runtime-verified) 1h 80 / 80 20 4/4 n=1 13.8 run 2026-08-15
Test Category Score tok/s
H1 Web-Estate One-Shot Builds 20/20 14.1
H2 Web-Estate One-Shot Builds 20/20 13.5
H3 Web-Estate One-Shot Builds 20/20 13.6
H4 Web-Estate One-Shot Builds 20/20 13.6
Production replay (real agent workload) 1i 172.8 / 240 14.4 12/12 n=1 11.5 run 2026-08-16
Test Category Score tok/s
I1 Nightly Fix Replay 8/20 12
I2 Nightly Fix Replay 13/20 10.1
I3 Nightly Fix Replay 0/20 9.8
I4 Nightly Fix Replay 16/20 10.4
I5 Nightly Fix Replay 10/20 10.6
I6 Nightly Fix Replay 9/20 10.9
I7 Audit Judgment 19/20 11.7
I8 Audit Judgment 20/20 12.3
I9 Cron Dependency Reasoning 20/20 12
I10 Cron Dependency Reasoning 20/20 12.2
I11 Curation Judgment 19/20 12.8
I12 Curation Judgment 19/20 13

Category mean Nightly Fix Replay 9.3Audit Judgment 19.5Cron Dependency Reasoning 20Curation Judgment 19

Each test is LLM-judged 0–20 against a fixed rubric; suite max = tests × 20. An ADJ chip marks a cell carrying a subjective adjustment from the maintainer; the judged score is the one shown, is unchanged, and remains what the axes, ranks and tie bands are computed from. Hover for the reason. Rows with a ▶ expand to the per-test breakdown. Where fewer tests ran than the suite total, unrun tests count as zero toward the axis score. See methodology.