← Leaderboard

Bonsai 27B Ternary 2bit MLX

Tested Rank #12 of 33 · 5/8 GAUNTLET progress
Generalist — 82.5 / 100 G Agentic — 97.5 / 100 A Understanding — pending U Needle — pending N Thinking — 93.1 / 100 T Live — pending L Engineering — 99 / 100 E Throughput — 25.2 / 100 T

Hollow markers & dashed spokes: axis not yet scored

G
82.5
A
97.5
U
N
T
93.1
L
E
99
T
25.2

Specification

Parameters
27B
Architecture
qwen3_5
Size on disk
7.9 GB
Quantization
2bit (ternary, PrismML extreme compression)
Format
MLX
Reasoning (CoT)
Yes — emits reasoning tokens
Internal ID
M21
Mean speed
41.1 tok/s across suites
Stall census
2 stalls in 29 observed tests (6.9%)
Reasoning appetite
2,960 tokens mean · 15,892 max
Model card
lmstudio.ai

Suite results

Suite Score Avg / 20 Tests tok/s
General capability (13-task real-workload suite) 1a 214.5 / 260 16.5 13/13 40.7
Agentic tool-calling & protocol adherence 1b 156 / 160 19.5 8/8 48.9
Coding depth 1c 118.8 / 120 19.8 6/6 33.8

Each test is LLM-judged 0–20 against a fixed rubric; suite max = tests × 20. Where fewer tests ran than the suite total, unrun tests count as zero toward the axis score. See methodology.