← Leaderboard

Qwen3.6-27B Fable-Fusion-711 Uncensored Heretic NEO-MAX MTPLX 8bit MLX

Tested RAN THE GAUNTLET Rank #3 of 51 · 8/8 GAUNTLET progress
Generalist — 95.5 / 100 G Agentic — 99 / 100 A Understanding — 90.8 / 100 U Needle — 73.5 / 100 N Thinking — 91.5 / 100 T Live — 53 / 100 L Engineering — 88.4 / 100 E Throughput — 4.5 / 100 T
G
95.5
A
99
U
90.8
N
73.5
T
91.5
L
53
E
88.4
T
4.5

Specification

Parameters
27B
Architecture
qwen3_5 (same fine-tune as M15, MLX repack by a different, unproven quantizer)
Size on disk
28 GB
Quantization
8bit
Format
MLX
Runtime
LM Studio
Class
Standard
Reasoning (CoT)
Yes — emits reasoning tokens
Internal ID
M19
Mean speed
15.7 tok/s across suites
Stall census
5 stalls in 59 observed tests (8.5%)
Reasoning appetite
2,767 tokens mean · 16,277 max
Model card
huggingface.co

Suite results

Suite Score Avg / 20 Tests n tok/s Run
General capability (13-task real-workload suite) 1a 248.3 / 260 19.1 13/13 n=1.1 16.5 run 2026-08-03
Test Category Score tok/s
A1 Dev/Ops Scripting 14/20 17.5
A2 Dev/Ops Scripting 17/20 16.7
A3 Dev/Ops Scripting 20/20 16
A4 Dev/Ops Scripting 17/20 15.5
A5 Dev/Ops Scripting 20/20 16.3
A6 Dev/Ops Scripting 20/20 16.8
B1 Document Processing 20/20 16.7
B2 Document Processing 20/20 16.3
B3 Document Processing 20/20 15.7
C1 Content Production 20/20 16.4
C2 Content Production 20/20 16.6
C3 Content Production 20/20 17
C4 Content Production 20/20 17

Category mean Dev/Ops Scripting 18Document Processing 20Content Production 20

Agentic tool-calling & protocol adherence 1b 158.4 / 160 19.8 8/8 n=1 17.8 run 2026-08-04
Test Category Score tok/s
WA1 Tool Calling & Protocol Adherence 20/20 15.9
WA2 Tool Calling & Protocol Adherence 20/20 16.4
WA3 Tool Calling & Protocol Adherence 20/20 16.7
WA4 Tool Calling & Protocol Adherence 20/20 15.9
WA5 Tool Calling & Protocol Adherence 20/20 17.9
WB1 Agent Robustness & State Management 20/20 17.7
WB2 Agent Robustness & State Management 20/20 17.1
WB3 Agent Robustness & State Management 18/20 16.9

Category mean Tool Calling & Protocol Adherence 20Agent Robustness & State Management 19.3

Coding depth 1c 159.2 / 180 19.9 8/9 ● n=1 16.8 run 2026-09-15
Test Category Score tok/s
K1 Coding — Concurrency 20/20 16.8
K2 Coding — Algorithms 20/20 16.7
K3 Coding — Security Review 20/20 17
K4 Coding — Performance 20/20 16.9
K5 Coding — Refactoring Under Constraints 20/20 16.7
K6 Coding — Test Writing 20/20 17
K7 Coding — React Component (typed) 19/20 17.9
K8 Coding — TypeScript Type System 20/20 17.4

Category mean Coding — Concurrency 20Coding — Algorithms 20Coding — Security Review 20Coding — Performance 20Coding — Refactoring Under Constraints 20Coding — Test Writing 20Coding — React Component (typed) 19Coding — TypeScript Type System 20

Doc/OCR vision 1d 160 / 160 20 8/8 n=1 16.9 run 2026-08-06
Test Category Score tok/s
D1 Doc/OCR — Transcription 20/20 15.1
D2 Doc/OCR — Document QA 20/20 13.5
D3 Doc/OCR — Field Extraction 20/20 14.2
D4 Doc/OCR — Document QA 20/20 15
D5 Doc/OCR — Transcription 20/20 14.8
D6 Doc/OCR — Field Extraction 20/20 14.8
D7 Doc/OCR — Degraded Input 20/20 16
D8 Doc/OCR — Degraded Input 20/20 15.4

Category mean Doc/OCR — Transcription 20Doc/OCR — Document QA 20Doc/OCR — Field Extraction 20Doc/OCR — Degraded Input 20

Doc/OCR — real-degraded tier 1d2 58 / 80 14.5 4/4 n=1 14.8 run 2026-08-14
Test Category Score tok/s
D9 Doc/OCR — Real Degraded Transcription 15/20 14.4
D10 Doc/OCR — Handwriting Extraction 19/20 10.6
D11 Doc/OCR — Degraded Table Transcription 14/20 13.9
D12 Doc/OCR — Degraded Table QA 10/20 14.2

Category mean Doc/OCR — Real Degraded Transcription 15Doc/OCR — Handwriting Extraction 19Doc/OCR — Degraded Table Transcription 14Doc/OCR — Degraded Table QA 10

Content-production depth 1e 115.2 / 120 19.2 6/6 n=1 16.5 run 2026-08-04
Test Category Score tok/s
E1 SOW & Requirements Decomposition 20/20 16.5
E2 SOW & Requirements Decomposition 20/20 16.4
E3 Executive & Proposal Content 18/20 16.4
E4 Executive & Proposal Content 20/20 16.7
E5 Executive & Proposal Content 17/20 16.7
E6 Executive & Proposal Content 20/20 16.1

Category mean SOW & Requirements Decomposition 20Executive & Proposal Content 18.8

Long-context multi-needle (MRCR) 1g2 44.1 / 60 14.7 3/3 n=1 13.8 run 2026-09-19
Test Category Score tok/s
L7 Long-Context — Multi-Needle Retrieval (MRCR) 20/20 11.5
L8 Long-Context — Multi-Needle Retrieval (MRCR) 20/20 9
L9 Long-Context — Multi-Needle Retrieval (MRCR) 4/20 12.5
Production replay (real agent workload) 1i 127.2 / 240 10.6 12/12 n=1 12.2 run 2026-08-16
Test Category Score tok/s
I1 Nightly Fix Replay 0/20 12.2
I2 Nightly Fix Replay 9/20 11.9
I3 Nightly Fix Replay 7/20 9
I4 Nightly Fix Replay 0/20 12.3
I5 Nightly Fix Replay 0/20 12.3
I6 Nightly Fix Replay 0/20 12.1
I7 Audit Judgment 15/20 13
I8 Audit Judgment 20/20 13
I9 Cron Dependency Reasoning 20/20 12.2
I10 Cron Dependency Reasoning 18/20 12.2
I11 Curation Judgment 20/20 12.9
I12 Curation Judgment 18/20 13

Category mean Nightly Fix Replay 2.7Audit Judgment 17.5Cron Dependency Reasoning 19Curation Judgment 19

Each test is LLM-judged 0–20 against a fixed rubric; suite max = tests × 20. An ADJ chip marks a cell carrying a subjective adjustment from the maintainer; the judged score is the one shown, is unchanged, and remains what the axes, ranks and tie bands are computed from. Hover for the reason. Rows with a ▶ expand to the per-test breakdown. Where fewer tests ran than the suite total, unrun tests count as zero toward the axis score. See methodology.