← Leaderboard
Qwen3.6-27B Dense 8bit MLX
Keeper RAN THE GAUNTLET Rank #3 of 33 · 8/8 GAUNTLET progress
G
91
A
100
U
93.2
N
99
T
97.2
L
68.3
E
91.5
T
8.2
Specification
- Parameters
- 27B dense
- Architecture
- qwen3_5
- Size on disk
- 29.53 GB
- Quantization
- 8bit
- Format
- MLX
- Reasoning (CoT)
- Yes — emits reasoning tokens
- Internal ID
- M7
- Mean speed
- 13.4 tok/s across suites
- Stall census
- 2 stalls in 71 observed tests (2.8%)
- Reasoning appetite
- 2,881 tokens mean · 16,310 max
- Model card
- lmstudio.ai
Suite results
| Suite | Score | Avg / 20 | Tests | tok/s |
|---|---|---|---|---|
| General capability (13-task real-workload suite) 1a | 236.6 / 260 | 18.2 | 13/13 | 15.2 |
| Agentic tool-calling & protocol adherence 1b | 160 / 160 | 20 | 8/8 | 16.6 |
| Coding depth 1c | 109.8 / 120 | 18.3 | 6/6 | 13.9 |
| Doc/OCR vision 1d | 153.6 / 160 | 19.2 | 8/8 | 14.1 |
| Doc/OCR — real-degraded tier 1d2 | 70 / 80 | 17.5 | 4/4 | 13.2 |
| Content-production depth 1e | 115.8 / 120 | 19.3 | 6/6 | 16.4 |
| Long-context retrieval & synthesis 1g | 118.2 / 120 | 19.7 | 6/6 | 10.5 |
| Long-context multi-needle (MRCR) 1g2 | 60 / 60 | 20 | 3/3 | 9.2 |
| Live one-shot builds (runtime-verified) 1h | 77 / 80 | 19.3 | 4/4 | — |
| Production replay (real agent workload) 1i | 141.6 / 240 | 11.8 | 12/12 | 11.5 |
Each test is LLM-judged 0–20 against a fixed rubric; suite max = tests × 20. Where fewer tests ran than the suite total, unrun tests count as zero toward the axis score. See methodology.