← Leaderboard

Qwen3.6-27B Fable-Fusion-711 Uncensored Heretic NEO-MAX Q8_0 GGUF

Keeper Rank #18 of 51 · 7/8 GAUNTLET progress
Generalist — 89.5 / 100 G Agentic — 100 / 100 A Understanding — 84.2 / 100 U Needle — 91.2 / 100 N Thinking — 97.7 / 100 T Live — pending L Engineering — 97 / 100 E Throughput — 3.7 / 100 T

Hollow markers & dashed spokes: axis not yet scored

G
89.5
A
100
U
84.2
N
91.2
T
97.7
L
—
E
97
T
3.7

Specification

Parameters
27B
Architecture
qwen3.6 hybrid Gated DeltaNet/Attention (same base architecture as M7)
Size on disk
29.79 GB
Quantization
Q8_0
Format
GGUF
Runtime
LM Studio
Class
Standard
Reasoning (CoT)
Yes — emits reasoning tokens
Internal ID
M15
Mean speed
13.1 tok/s across suites
Stall census
1 stall in 43 observed tests (2.3%)
Reasoning appetite
2,046 tokens mean · 16,383 max
Model card
huggingface.co

Suite results

Suite Score Avg / 20 Tests n tok/s Run
General capability (13-task real-workload suite) 1a 232.7 / 260 17.9 13/13 n=2.2 12.8 run 2026-07-29
Test Category Score tok/s
A1 Dev/Ops Scripting 20/20 14.9
A2 Dev/Ops Scripting 20/20 12.2
A3 Dev/Ops Scripting 0/20 12.3
A4 Dev/Ops Scripting 19/20 12.5
A5 Dev/Ops Scripting 19/20 13.6
A6 Dev/Ops Scripting 15/20 13.4
B1 Document Processing 20/20 12.6
B2 Document Processing 20/20 12.2
B3 Document Processing 20/20 12.5
C1 Content Production 20/20 12.9
C2 Content Production 20/20 12.8
C3 Content Production 20/20 12.2
C4 Content Production 19/20 12

Category mean Dev/Ops Scripting 15.5Document Processing 20Content Production 19.8

Agentic tool-calling & protocol adherence 1b 160 / 160 20 8/8 n=1 15.1 run 2026-07-30
Test Category Score tok/s
WA1 Tool Calling & Protocol Adherence 20/20 11.7
WA2 Tool Calling & Protocol Adherence 20/20 12.9
WA3 Tool Calling & Protocol Adherence 20/20 13.2
WA4 Tool Calling & Protocol Adherence 20/20 11.6
WA5 Tool Calling & Protocol Adherence 20/20 14.6
WB1 Agent Robustness & State Management 20/20 12.5
WB2 Agent Robustness & State Management 20/20 11
WB3 Agent Robustness & State Management 20/20 11.4

Category mean Tool Calling & Protocol Adherence 20Agent Robustness & State Management 20

Coding depth 1c 174.6 / 180 19.4 9/9 n=1 14 run 2026-09-15
Test Category Score tok/s
K1 Coding — Concurrency 20/20 13.5
K2 Coding — Algorithms 20/20 13.4
K3 Coding — Security Review 20/20 14.1
K4 Coding — Performance 19/20 14.3
K5 Coding — Refactoring Under Constraints 20/20 14.3
K6 Coding — Test Writing 19/20 14.2
K7 Coding — React Component (typed) 17/20 13.1
K8 Coding — TypeScript Type System 20/20 13.7
K9 Coding — React Debugging 20/20 13.4

Category mean Coding — Concurrency 20Coding — Algorithms 20Coding — Security Review 20Coding — Performance 19Coding — Refactoring Under Constraints 20Coding — Test Writing 19Coding — React Component (typed) 17Coding — TypeScript Type System 20Coding — React Debugging 20

Doc/OCR vision 1d 160 / 160 20 8/8 n=1 11.4 run 2026-09-20
Test Category Score tok/s
D1 Doc/OCR — Transcription 20/20 10
D2 Doc/OCR — Document QA 20/20 8.7
D3 Doc/OCR — Field Extraction 20/20 9.8
D4 Doc/OCR — Document QA 20/20 9.4
D5 Doc/OCR — Transcription 20/20 11
D6 Doc/OCR — Field Extraction 20/20 9.9
D7 Doc/OCR — Degraded Input 20/20 11.9
D8 Doc/OCR — Degraded Input 20/20 11.6

Category mean Doc/OCR — Transcription 20Doc/OCR — Document QA 20Doc/OCR — Field Extraction 20Doc/OCR — Degraded Input 20

Doc/OCR — real-degraded tier 1d2 42 / 80 14 3/4 ● n=1 11.2 run 2026-09-20
Test Category Score tok/s
D9 Doc/OCR — Real Degraded Transcription 13/20 11
D10 Doc/OCR — Handwriting Extraction 17/20 11.7
D12 Doc/OCR — Degraded Table QA 12/20 11

Category mean Doc/OCR — Real Degraded Transcription 13Doc/OCR — Handwriting Extraction 17Doc/OCR — Degraded Table QA 12

Content-production depth 1e 115.2 / 120 19.2 6/6 n=1 15.2 run 2026-08-02
Test Category Score tok/s
E1 SOW & Requirements Decomposition 20/20 15.7
E2 SOW & Requirements Decomposition 20/20 15.6
E3 Executive & Proposal Content 18/20 15.2
E4 Executive & Proposal Content 19/20 14.8
E5 Executive & Proposal Content 20/20 14.8
E6 Executive & Proposal Content 18/20 14.8

Category mean SOW & Requirements Decomposition 20Executive & Proposal Content 18.8

Long-context retrieval & synthesis 1g 120 / 120 20 6/6 n=1 14.2 run 2026-08-06
Test Category Score tok/s
L1 Long-Context — Needle Retrieval 20/20 11.4
L2 Long-Context — Grounded QA with Distractors 20/20 7.7
L3 Long-Context — Ops Log Reasoning 20/20 6.6
L4 Long-Context — Faithful Summarization 20/20 8.7
L5 Long-Context — Instruction Retention 20/20 5.3
L6 Long-Context — Full-Haystack QA 20/20 2.1

Category mean Long-Context — Needle Retrieval 20Long-Context — Grounded QA with Distractors 20Long-Context — Ops Log Reasoning 20Long-Context — Faithful Summarization 20Long-Context — Instruction Retention 20Long-Context — Full-Haystack QA 20

Long-context multi-needle (MRCR) 1g2 44.1 / 60 14.7 3/3 n=1 10.9 run 2026-08-14
Test Category Score tok/s
L7 Long-Context — Multi-Needle Retrieval (MRCR) 20/20 5.6
L8 Long-Context — Multi-Needle Retrieval (MRCR) 20/20 4.9
L9 Long-Context — Multi-Needle Retrieval (MRCR) 4/20 3

Each test is LLM-judged 0–20 against a fixed rubric; suite max = tests × 20. An ADJ chip marks a cell carrying a subjective adjustment from the maintainer; the judged score is the one shown, is unchanged, and remains what the axes, ranks and tie bands are computed from. Hover for the reason. Rows with a ▶ expand to the per-test breakdown. Where fewer tests ran than the suite total, unrun tests count as zero toward the axis score. See methodology.