← Leaderboard

Mistral Small 3.2 24B 8bit MLX

Keeper RAN THE GAUNTLET Rank #11 of 51 · 8/8 GAUNTLET progress
Generalist — 80.5 / 100 G Agentic — 91 / 100 A Understanding — 79 / 100 U Needle — 80.2 / 100 N Thinking — 100 / 100 T Live — 62.8 / 100 L Engineering — 80.5 / 100 E Throughput — 5.2 / 100 T
G
80.5
A
91
U
79
N
80.2
T
100
L
62.8
E
80.5
T
5.2

Specification

Parameters
24B
Architecture
mistral3
Size on disk
25.93 GB
Quantization
8bit
Format
MLX
Runtime
LM Studio
Class
Standard
Reasoning (CoT)
No
Internal ID
M13
Mean speed
18.2 tok/s across suites
Stall census
0 stalls in 54 observed tests (0.0%)
Model card
huggingface.co

Suite results

Suite Score Avg / 20 Tests n tok/s Run
General capability (13-task real-workload suite) 1a 209.3 / 260 16.1 13/13 n=2 21 run 2026-09-02
Test Category Score tok/s
A1 Dev/Ops Scripting 11/20 20.3
A2 Dev/Ops Scripting 17/20 19.8
A3 Dev/Ops Scripting 17/20 20.3
A4 Dev/Ops Scripting 15/20 20
A5 Dev/Ops Scripting 16/20 20.7
A6 Dev/Ops Scripting 20/20 20.8
B1 Document Processing 19/20 19.6
B2 Document Processing 20/20 19.7
B3 Document Processing 10/20 19.3
C1 Content Production 18/20 20.7
C2 Content Production 15/20 20.1
C3 Content Production 17/20 21.2
C4 Content Production 14/20 20.9

Category mean Dev/Ops Scripting 16Document Processing 16.3Content Production 16

Agentic tool-calling & protocol adherence 1b 145.6 / 160 18.2 8/8 n=1 21.7 run 2026-07-28
Test Category Score tok/s
WA1 Tool Calling & Protocol Adherence 20/20 5.2
WA2 Tool Calling & Protocol Adherence 20/20 17.3
WA3 Tool Calling & Protocol Adherence 16/20 17.8
WA4 Tool Calling & Protocol Adherence 20/20 16.4
WA5 Tool Calling & Protocol Adherence 20/20 20.1
WB1 Agent Robustness & State Management 20/20 20.6
WB2 Agent Robustness & State Management 19/20 20.5
WB3 Agent Robustness & State Management 11/20 18.8

Category mean Tool Calling & Protocol Adherence 19.2Agent Robustness & State Management 16.7

Coding depth 1c 144.9 / 180 16.1 9/9 n=1 16.1 run 2026-09-15
Test Category Score tok/s
K1 Coding — Concurrency 16/20 15.3
K2 Coding — Algorithms 20/20 14.9
K3 Coding — Security Review 13/20 14.2
K4 Coding — Performance 15/20 17.2
K5 Coding — Refactoring Under Constraints 20/20 13.7
K6 Coding — Test Writing 20/20 13.5
K7 Coding — React Component (typed) 9/20 19.1
K8 Coding — TypeScript Type System 17/20 18.2
K9 Coding — React Debugging 15/20 18.9

Category mean Coding — Concurrency 16Coding — Algorithms 20Coding — Security Review 13Coding — Performance 15Coding — Refactoring Under Constraints 20Coding — Test Writing 20Coding — React Component (typed) 9Coding — TypeScript Type System 17Coding — React Debugging 15

Doc/OCR vision 1d 142.4 / 160 17.8 8/8 n=1 19.9 run 2026-08-06
Test Category Score tok/s
D1 Doc/OCR — Transcription 20/20 16.4
D2 Doc/OCR — Document QA 20/20 8.7
D3 Doc/OCR — Field Extraction 20/20 15.3
D4 Doc/OCR — Document QA 11/20 12.9
D5 Doc/OCR — Transcription 11/20 15.5
D6 Doc/OCR — Field Extraction 20/20 12.8
D7 Doc/OCR — Degraded Input 20/20 16.2
D8 Doc/OCR — Degraded Input 20/20 14.4

Category mean Doc/OCR — Transcription 15.5Doc/OCR — Document QA 15.5Doc/OCR — Field Extraction 20Doc/OCR — Degraded Input 20

Doc/OCR — real-degraded tier 1d2 47.2 / 80 11.8 4/4 n=1 18.7 run 2026-08-14
Test Category Score tok/s
D9 Doc/OCR — Real Degraded Transcription 9/20 14.5
D10 Doc/OCR — Handwriting Extraction 16/20 10.3
D11 Doc/OCR — Degraded Table Transcription 12/20 14.3
D12 Doc/OCR — Degraded Table QA 10/20 6.5

Category mean Doc/OCR — Real Degraded Transcription 9Doc/OCR — Handwriting Extraction 16Doc/OCR — Degraded Table Transcription 12Doc/OCR — Degraded Table QA 10

Content-production depth 1e 88.8 / 120 14.8 6/6 n=1 20.4 run 2026-08-02
Test Category Score tok/s
E1 SOW & Requirements Decomposition 8/20 18.5
E2 SOW & Requirements Decomposition 20/20 19.1
E3 Executive & Proposal Content 16/20 20.3
E4 Executive & Proposal Content 8/20 21
E5 Executive & Proposal Content 17/20 20.1
E6 Executive & Proposal Content 20/20 19.3

Category mean SOW & Requirements Decomposition 14Executive & Proposal Content 15.3

Long-context retrieval & synthesis 1g 115.2 / 120 19.2 6/6 n=1 15.4 run 2026-08-06
Test Category Score tok/s
L1 Long-Context — Needle Retrieval 20/20 2.8
L2 Long-Context — Grounded QA with Distractors 15/20 1.4
L3 Long-Context — Ops Log Reasoning 20/20 2.1
L4 Long-Context — Faithful Summarization 20/20 6.9
L5 Long-Context — Instruction Retention 20/20 1.6
L6 Long-Context — Full-Haystack QA 20/20 0.5

Category mean Long-Context — Needle Retrieval 20Long-Context — Grounded QA with Distractors 15Long-Context — Ops Log Reasoning 20Long-Context — Faithful Summarization 20Long-Context — Instruction Retention 20Long-Context — Full-Haystack QA 20

Long-context multi-needle (MRCR) 1g2 29.1 / 60 9.7 3/3 n=1 12.3 run 2026-08-14
Test Category Score tok/s
L7 Long-Context — Multi-Needle Retrieval (MRCR) 20/20 5.1
L8 Long-Context — Multi-Needle Retrieval (MRCR) 5/20 3.3
L9 Long-Context — Multi-Needle Retrieval (MRCR) 4/20 2.5
Live one-shot builds (runtime-verified) 1h 50 / 80 12.5 4/4 n=1 15.6 run 2026-08-28
Test Category Score tok/s
H1 Web-Estate One-Shot Builds 16/20 16.9
H2 Web-Estate One-Shot Builds 12/20 15.4
H3 Web-Estate One-Shot Builds 10/20 15
H4 Web-Estate One-Shot Builds 12/20 15.2
Brand-constrained static build (negative-constraint adherence) 1j 63 / 100 12.6 5/5 n=1 20.8 run 2026-09-16
Test Category Score tok/s
J1 Brand-Constrained Static Build 17/20 20.1
J3 Brand-Constrained Static Build 11/20 20.4
J4 Brand-Constrained Static Build 6/20 20.5
J5 Brand-Constrained Static Build 15/20 21.4
J6 Brand-Constrained Static Build 14/20 21.5

Each test is LLM-judged 0–20 against a fixed rubric; suite max = tests × 20. An ADJ chip marks a cell carrying a subjective adjustment from the maintainer; the judged score is the one shown, is unchanged, and remains what the axes, ranks and tie bands are computed from. Hover for the reason. Rows with a ▶ expand to the per-test breakdown. Where fewer tests ran than the suite total, unrun tests count as zero toward the axis score. See methodology.