← Leaderboard
Qwen3.6-27B Fable-Fusion-711 Uncensored Heretic NEO-MAX MTPLX 8bit MLX
Tested Rank #4 of 33 · 7/8 GAUNTLET progress
Hollow markers & dashed spokes: axis not yet scored
G
95.5
A
97.5
U
88
N
—
T
91.5
L
45.5
E
100
T
9.4
Specification
- Parameters
- 27B
- Architecture
- qwen3_5 (same fine-tune as M15, MLX repack by a different, unproven quantizer)
- Size on disk
- 28 GB
- Quantization
- 8bit
- Format
- MLX
- Reasoning (CoT)
- Yes — emits reasoning tokens
- Internal ID
- M19
- Mean speed
- 15.3 tok/s across suites
- Stall census
- 5 stalls in 59 observed tests (8.5%)
- Reasoning appetite
- 2,767 tokens mean · 16,277 max
- Model card
- huggingface.co
Suite results
| Suite | Score | Avg / 20 | Tests | tok/s |
|---|---|---|---|---|
| General capability (13-task real-workload suite) 1a | 248.3 / 260 | 19.1 | 13/13 | 16.5 |
| Agentic tool-calling & protocol adherence 1b | 156 / 160 | 19.5 | 8/8 | 16.8 |
| Coding depth 1c | 120 / 120 | 20 | 6/6 | 16.8 |
| Doc/OCR vision 1d | 152 / 160 | 19 | 8/8 | 14.8 |
| Doc/OCR — real-degraded tier 1d2 | 59.2 / 80 | 14.8 | 4/4 | 13.3 |
| Content-production depth 1e | 117 / 120 | 19.5 | 6/6 | 16.5 |
| Production replay (real agent workload) 1i | 109.2 / 240 | 9.1 | 12/12 | 12.2 |
Each test is LLM-judged 0–20 against a fixed rubric; suite max = tests × 20. Where fewer tests ran than the suite total, unrun tests count as zero toward the axis score. See methodology.