← Leaderboard

Bonsai 27B Ternary 2bit MLX

Tested RAN THE GAUNTLET Rank #10 of 51 · 8/8 GAUNTLET progress
Generalist — 81 / 100 G Agentic — 92 / 100 A Understanding — 90.3 / 100 U Needle — 80 / 100 N Thinking — 93.1 / 100 T Live — 75.6 / 100 L Engineering — 96 / 100 E Throughput — 10.4 / 100 T
G
81
A
92
U
90.3
N
80
T
93.1
L
75.6
E
96
T
10.4

Specification

Parameters
27B
Architecture
qwen3_5
Size on disk
7.9 GB
Quantization
2bit (ternary, PrismML extreme compression)
Format
MLX
Runtime
LM Studio
Class
Standard
Reasoning (CoT)
Yes — emits reasoning tokens
Internal ID
M21
Mean speed
36.7 tok/s across suites
Stall census
2 stalls in 29 observed tests (6.9%)
Reasoning appetite
2,960 tokens mean · 15,892 max
Model card
lmstudio.ai

Suite results

Suite Score Avg / 20 Tests n tok/s Run
General capability (13-task real-workload suite) 1a 210.6 / 260 16.2 13/13 n=1.1 40.7 run 2026-08-05
Test Category Score tok/s
A1 Dev/Ops Scripting 5/20 48.8
A2 Dev/Ops Scripting 13/20 45.5
A3 Dev/Ops Scripting 0/20 37.1
A4 Dev/Ops Scripting 17/20 38.3
A5 Dev/Ops Scripting 20/20 41
A6 Dev/Ops Scripting 20/20 42.8
B1 Document Processing 20/20 43.1
B2 Document Processing 20/20 40.2
B3 Document Processing 20/20 36.1
C1 Content Production 20/20 38.5
C2 Content Production 18/20 39.5
C3 Content Production 18/20 38.9
C4 Content Production 19/20 39.3

Category mean Dev/Ops Scripting 12.5Document Processing 20Content Production 18.8

Agentic tool-calling & protocol adherence 1b 147.2 / 160 18.4 8/8 n=1 55.2 run 2026-08-05
Test Category Score tok/s
WA1 Tool Calling & Protocol Adherence 20/20 50.1
WA2 Tool Calling & Protocol Adherence 16/20 45.7
WA3 Tool Calling & Protocol Adherence 20/20 50.7
WA4 Tool Calling & Protocol Adherence 20/20 43.2
WA5 Tool Calling & Protocol Adherence 20/20 53.6
WB1 Agent Robustness & State Management 20/20 48.4
WB2 Agent Robustness & State Management 20/20 46.5
WB3 Agent Robustness & State Management 11/20 53.1

Category mean Tool Calling & Protocol Adherence 19.2Agent Robustness & State Management 17

Coding depth 1c 172.8 / 180 19.2 9/9 n=1 35.1 run 2026-09-15
Test Category Score tok/s
K1 Coding — Concurrency 20/20 43.8
K2 Coding — Algorithms 20/20 36.2
K3 Coding — Security Review 20/20 32
K4 Coding — Performance 19/20 30.8
K5 Coding — Refactoring Under Constraints 20/20 29.7
K6 Coding — Test Writing 20/20 30
K7 Coding — React Component (typed) 14/20 39.4
K8 Coding — TypeScript Type System 20/20 37.5
K9 Coding — React Debugging 20/20 36.5

Category mean Coding — Concurrency 20Coding — Algorithms 20Coding — Security Review 20Coding — Performance 19Coding — Refactoring Under Constraints 20Coding — Test Writing 20Coding — React Component (typed) 14Coding — TypeScript Type System 20Coding — React Debugging 20

Doc/OCR vision 1d 160 / 160 20 8/8 n=1 34.3 run 2026-09-22
Test Category Score tok/s
D1 Doc/OCR — Transcription 20/20 39
D2 Doc/OCR — Document QA 20/20 35.8
D3 Doc/OCR — Field Extraction 20/20 31.4
D4 Doc/OCR — Document QA 20/20 27.2
D5 Doc/OCR — Transcription 20/20 27.1
D6 Doc/OCR — Field Extraction 20/20 27.7
D7 Doc/OCR — Degraded Input 20/20 32.5
D8 Doc/OCR — Degraded Input 20/20 33.2

Category mean Doc/OCR — Transcription 20Doc/OCR — Document QA 20Doc/OCR — Field Extraction 20Doc/OCR — Degraded Input 20

Doc/OCR — real-degraded tier 1d2 56.8 / 80 14.2 4/4 n=1 36.6 run 2026-09-21
Test Category Score tok/s
D9 Doc/OCR — Real Degraded Transcription 12/20 36.3
D10 Doc/OCR — Handwriting Extraction 16/20 32.3
D11 Doc/OCR — Degraded Table Transcription 12/20 36
D12 Doc/OCR — Degraded Table QA 17/20 36.2

Category mean Doc/OCR — Real Degraded Transcription 12Doc/OCR — Handwriting Extraction 16Doc/OCR — Degraded Table Transcription 12Doc/OCR — Degraded Table QA 17

Content-production depth 1e 109.2 / 120 18.2 6/6 n=1 38.5 run 2026-09-25
Test Category Score tok/s
E1 SOW & Requirements Decomposition 20/20 42.2
E2 SOW & Requirements Decomposition 20/20 40.4
E3 Executive & Proposal Content 18/20 35.5
E4 Executive & Proposal Content 16/20 38.6
E5 Executive & Proposal Content 16/20 36.2
E6 Executive & Proposal Content 19/20 38.1

Category mean SOW & Requirements Decomposition 20Executive & Proposal Content 17.3

Long-context retrieval & synthesis 1g 120 / 120 20 6/6 n=1 35.6 run 2026-09-24
Test Category Score tok/s
L1 Long-Context — Needle Retrieval 20/20 35.7
L2 Long-Context — Grounded QA with Distractors 20/20 28.6
L3 Long-Context — Ops Log Reasoning 20/20 20.4
L4 Long-Context — Faithful Summarization 20/20 26.4
L5 Long-Context — Instruction Retention 20/20 18.4
L6 Long-Context — Full-Haystack QA 20/20 24.9

Category mean Long-Context — Needle Retrieval 20Long-Context — Grounded QA with Distractors 20Long-Context — Ops Log Reasoning 20Long-Context — Faithful Summarization 20Long-Context — Instruction Retention 20Long-Context — Full-Haystack QA 20

Long-context multi-needle (MRCR) 1g2 24 / 60 8 3/3 n=1 22.5 run 2026-09-19
Test Category Score tok/s
L7 Long-Context — Multi-Needle Retrieval (MRCR) 20/20 20
L8 Long-Context — Multi-Needle Retrieval (MRCR) 4/20 24.2
Live one-shot builds (runtime-verified) 1h 48 / 80 12 4/4 n=1 36.1 run 2026-09-21
Test Category Score tok/s
H1 Web-Estate One-Shot Builds 20/20 40.9
H2 Web-Estate One-Shot Builds 12/20 37.9
H3 Web-Estate One-Shot Builds 16/20 32
H4 Web-Estate One-Shot Builds 0/20 33.8
Brand-constrained static build (negative-constraint adherence) 1j 88 / 100 17.6 5/5 n=1 32.1 run 2026-09-23
Test Category Score tok/s
J1 Brand-Constrained Static Build 20/20 32.3
J3 Brand-Constrained Static Build 20/20 33.4
J4 Brand-Constrained Static Build 20/20 33.8
J5 Brand-Constrained Static Build 11/20 30.8
J6 Brand-Constrained Static Build 17/20 30

Each test is LLM-judged 0–20 against a fixed rubric; suite max = tests × 20. An ADJ chip marks a cell carrying a subjective adjustment from the maintainer; the judged score is the one shown, is unchanged, and remains what the axes, ranks and tie bands are computed from. Hover for the reason. Rows with a ▶ expand to the per-test breakdown. Where fewer tests ran than the suite total, unrun tests count as zero toward the axis score. See methodology.