← Leaderboard

Devstral Small 2 24B 6bit MLX

Retired Rank #41 of 51 · 4/8 GAUNTLET progress
Generalist — 78 / 100 G Agentic — 97 / 100 A Understanding — pending U Needle — pending N Thinking — 100 / 100 T Live — pending L Engineering — pending E Throughput — 6.3 / 100 T

Hollow markers & dashed spokes: axis not yet scored

G
78
A
97
U
—
N
—
T
100
L
—
E
—
T
6.3

Specification

Parameters
24B
Architecture
mistral3
Size on disk
19.43 GB
Quantization
6bit
Format
MLX
Runtime
LM Studio
Class
Standard
Reasoning (CoT)
No
Internal ID
M22
Mean speed
22 tok/s across suites
Stall census
0 stalls in 22 observed tests (0.0%)
Model card
huggingface.co

Suite results

Suite Score Avg / 20 Tests n tok/s Run
General capability (13-task real-workload suite) 1a 202.8 / 260 15.6 13/13 n=1.1 17.7 run 2026-08-06
Test Category Score tok/s
A1 Dev/Ops Scripting 14/20 26.1
A2 Dev/Ops Scripting 14/20 24.9
A3 Dev/Ops Scripting 9/20 23.5
A4 Dev/Ops Scripting 18/20 23.8
A5 Dev/Ops Scripting 16/20 18.4
A6 Dev/Ops Scripting 20/20 16.8
B1 Document Processing 20/20 14.2
B2 Document Processing 20/20 15
B3 Document Processing 12/20 13.8
C1 Content Production 15/20 13.3
C2 Content Production 15/20 12.9
C3 Content Production 18/20 13.7
C4 Content Production 12/20 14.1

Category mean Dev/Ops Scripting 15.2Document Processing 17.3Content Production 15

Agentic tool-calling & protocol adherence 1b 155.2 / 160 19.4 8/8 n=1 26.3 run 2026-08-06
Test Category Score tok/s
WA1 Tool Calling & Protocol Adherence 20/20 20.1
WA2 Tool Calling & Protocol Adherence 20/20 22
WA3 Tool Calling & Protocol Adherence 20/20 22
WA4 Tool Calling & Protocol Adherence 20/20 20.7
WA5 Tool Calling & Protocol Adherence 20/20 25.3
WB1 Agent Robustness & State Management 20/20 25
WB2 Agent Robustness & State Management 20/20 24.6
WB3 Agent Robustness & State Management 15/20 23.8

Category mean Tool Calling & Protocol Adherence 20Agent Robustness & State Management 18.3

Each test is LLM-judged 0–20 against a fixed rubric; suite max = tests × 20. An ADJ chip marks a cell carrying a subjective adjustment from the maintainer; the judged score is the one shown, is unchanged, and remains what the axes, ranks and tie bands are computed from. Hover for the reason. Rows with a ▶ expand to the per-test breakdown. Where fewer tests ran than the suite total, unrun tests count as zero toward the axis score. See methodology.