← Leaderboard

Phi-3.5-mini-instruct 4bit MLXSMALL

Tested Rank #35 of 51 · 6/8 GAUNTLET progress
Generalist — 36 / 100 G Agentic — 26 / 100 A Understanding — pending U Needle — 37.2 / 100 N Thinking — pending T Live — 10.6 / 100 L Engineering — 25 / 100 E Throughput — 26 / 100 T

Hollow markers & dashed spokes: axis not yet scored

G
36
A
26
U
—
N
37.2
T
n/d
L
10.6
E
25
T
26

Specification

Parameters
3.8B
Architecture
phi3
Size on disk
2 GB
Quantization
4bit
Format
MLX
Runtime
LM Studio
Class
Small (under 8B total parameters — phone/edge class)
Reasoning (CoT)
No
Internal ID
M17
Mean speed
91.2 tok/s across suites
Stall census
0 stalls in 13 observed tests (0.0%)
Model card
huggingface.co

Suite results

Suite Score Avg / 20 Tests n tok/s Run
General capability (13-task real-workload suite) 1a 93.6 / 260 7.2 13/13 n=1 114.8 run 2026-08-04
Test Category Score tok/s
A1 Dev/Ops Scripting 2/20 76.3
A2 Dev/Ops Scripting 7/20 117.1
A3 Dev/Ops Scripting 1/20 146.5
A4 Dev/Ops Scripting 6/20 158.2
A5 Dev/Ops Scripting 5/20 71.8
A6 Dev/Ops Scripting 11/20 133.3
B1 Document Processing 11/20 131.2
B2 Document Processing 10/20 137.7
B3 Document Processing 7/20 56.8
C1 Content Production 13/20 157.4
C2 Content Production 8/20 60.8
C3 Content Production 8/20 127.4
C4 Content Production 5/20 57.7

Category mean Dev/Ops Scripting 5.3Document Processing 9.3Content Production 8.5

Agentic tool-calling & protocol adherence 1b 41.6 / 160 5.2 8/8 n=1 95.7 run 2026-09-22
Test Category Score tok/s
WA1 Tool Calling & Protocol Adherence 9/20 137.4
WA2 Tool Calling & Protocol Adherence 3/20 70
WA3 Tool Calling & Protocol Adherence 2/20 70.4
WA4 Tool Calling & Protocol Adherence 0/20 71.5
WA5 Tool Calling & Protocol Adherence 0/20 126.7
WB1 Agent Robustness & State Management 8/20 73.8
WB2 Agent Robustness & State Management 13/20 70.9
WB3 Agent Robustness & State Management 7/20 129.4

Category mean Tool Calling & Protocol Adherence 2.8Agent Robustness & State Management 9.3

Coding depth 1c 45 / 180 5 9/9 n=1 129.1 run 2026-09-29
Test Category Score tok/s
K1 Coding — Concurrency 3/20 122.4
K2 Coding — Algorithms 1/20 159.7
K3 Coding — Security Review 6/20 155
K4 Coding — Performance 12/20 69.3
K5 Coding — Refactoring Under Constraints 9/20 163.5
K6 Coding — Test Writing 5/20 59.9
K7 Coding — React Component (typed) 2/20 170.7
K8 Coding — TypeScript Type System 3/20 60.6
K9 Coding — React Debugging 4/20 104

Category mean Coding — Concurrency 3Coding — Algorithms 1Coding — Security Review 6Coding — Performance 12Coding — Refactoring Under Constraints 9Coding — Test Writing 5Coding — React Component (typed) 2Coding — TypeScript Type System 3Coding — React Debugging 4

Content-production depth 1e 27 / 120 4.5 6/6 n=1 125.7 run 2026-09-24
Test Category Score tok/s
E1 SOW & Requirements Decomposition 5/20 143.6
E2 SOW & Requirements Decomposition 2/20 66.6
E3 Executive & Proposal Content 7/20 123.7
E4 Executive & Proposal Content 5/20 162.9
E5 Executive & Proposal Content 8/20 165.5
E6 Executive & Proposal Content 0/20 61.9

Category mean SOW & Requirements Decomposition 3.5Executive & Proposal Content 5

Long-context retrieval & synthesis 1g 67 / 120 13.4 5/6 ● n=1 29.5 run 2026-09-23
Test Category Score tok/s
L1 Long-Context — Needle Retrieval 3/20 30.4
L2 Long-Context — Grounded QA with Distractors 20/20 6.2
L3 Long-Context — Ops Log Reasoning 8/20 9.5
L5 Long-Context — Instruction Retention 18/20 4.5
L6 Long-Context — Full-Haystack QA 18/20 1.4

Category mean Long-Context — Needle Retrieval 3Long-Context — Grounded QA with Distractors 20Long-Context — Ops Log Reasoning 8Long-Context — Instruction Retention 18Long-Context — Full-Haystack QA 18

Long-context multi-needle (MRCR) 1g2 0 / 60 0 3/3 n=1 23.7 run 2026-09-27
Test Category Score tok/s
L7 Long-Context — Multi-Needle Retrieval (MRCR) 0/20 23.1
Live one-shot builds (runtime-verified) 1h 2 / 80 0.5 4/4 n=1 102.3 run 2026-10-01
Test Category Score tok/s
H1 Web-Estate One-Shot Builds 0/20 53.7
H3 Web-Estate One-Shot Builds 1/20 51
Brand-constrained static build (negative-constraint adherence) 1j 17 / 100 3.4 5/5 n=1 108.5 run 2026-09-22
Test Category Score tok/s
J1 Brand-Constrained Static Build 7/20 171.6
J3 Brand-Constrained Static Build 2/20 73.1
J4 Brand-Constrained Static Build 2/20 130.8
J5 Brand-Constrained Static Build 6/20 73
J6 Brand-Constrained Static Build 0/20 69.7

Each test is LLM-judged 0–20 against a fixed rubric; suite max = tests × 20. An ADJ chip marks a cell carrying a subjective adjustment from the maintainer; the judged score is the one shown, is unchanged, and remains what the axes, ranks and tie bands are computed from. Hover for the reason. Rows with a ▶ expand to the per-test breakdown. Where fewer tests ran than the suite total, unrun tests count as zero toward the axis score. See methodology.