← Leaderboard

Qwen3.6-27B Fable-Fusion-711 Uncensored Heretic NEO-MAX MTPLX 8bit MLX

Tested Rank #4 of 33 · 7/8 GAUNTLET progress
Generalist — 95.5 / 100 G Agentic — 97.5 / 100 A Understanding — 88 / 100 U Needle — pending N Thinking — 91.5 / 100 T Live — 45.5 / 100 L Engineering — 100 / 100 E Throughput — 9.4 / 100 T

Hollow markers & dashed spokes: axis not yet scored

G
95.5
A
97.5
U
88
N
T
91.5
L
45.5
E
100
T
9.4

Specification

Parameters
27B
Architecture
qwen3_5 (same fine-tune as M15, MLX repack by a different, unproven quantizer)
Size on disk
28 GB
Quantization
8bit
Format
MLX
Reasoning (CoT)
Yes — emits reasoning tokens
Internal ID
M19
Mean speed
15.3 tok/s across suites
Stall census
5 stalls in 59 observed tests (8.5%)
Reasoning appetite
2,767 tokens mean · 16,277 max
Model card
huggingface.co

Suite results

Suite Score Avg / 20 Tests tok/s
General capability (13-task real-workload suite) 1a 248.3 / 260 19.1 13/13 16.5
Agentic tool-calling & protocol adherence 1b 156 / 160 19.5 8/8 16.8
Coding depth 1c 120 / 120 20 6/6 16.8
Doc/OCR vision 1d 152 / 160 19 8/8 14.8
Doc/OCR — real-degraded tier 1d2 59.2 / 80 14.8 4/4 13.3
Content-production depth 1e 117 / 120 19.5 6/6 16.5
Production replay (real agent workload) 1i 109.2 / 240 9.1 12/12 12.2

Each test is LLM-judged 0–20 against a fixed rubric; suite max = tests × 20. Where fewer tests ran than the suite total, unrun tests count as zero toward the axis score. See methodology.