← Leaderboard

DeepSeek-R1-Distill-Llama 70B 8bit MLX

Retired Rank #48 of 51 · 2/8 GAUNTLET progress
Generalist — 62.5 / 100 G Agentic — pending A Understanding — pending U Needle — pending N Thinking — pending T Live — pending L Engineering — pending E Throughput — 1.6 / 100 T

Hollow markers & dashed spokes: axis not yet scored

G
62.5
A
—
U
—
N
—
T
n/d
L
—
E
—
T
1.6

Specification

Parameters
70B
Architecture
llama
Size on disk
75 GB
Quantization
8bit
Format
MLX
Runtime
LM Studio
Class
Standard
Reasoning (CoT)
Yes — emits reasoning tokens
Internal ID
M5
Mean speed
5.7 tok/s across suites
Stall census
0 stalls in 13 observed tests (0.0%)
Reasoning appetite
1,049 tokens mean · 6,741 max
Model card
huggingface.co

Suite results

Suite Score Avg / 20 Tests n tok/s Run
General capability (13-task real-workload suite) 1a 162.5 / 260 12.5 13/13 n=2 5.7 run 2026-09-02
Test Category Score tok/s
A1 Dev/Ops Scripting 7/20 6.1
A2 Dev/Ops Scripting 14/20 5.7
A3 Dev/Ops Scripting 19/20 5.6
A4 Dev/Ops Scripting 16/20 5.8
A5 Dev/Ops Scripting 5/20 5.8
A6 Dev/Ops Scripting 9/20 5.9
B1 Document Processing 19/20 5.8
B2 Document Processing 20/20 5.6
B3 Document Processing 5/20 5.3
C1 Content Production 10/20 5.7
C2 Content Production 13/20 5.6
C3 Content Production 12/20 5.6
C4 Content Production 14/20 5.8

Category mean Dev/Ops Scripting 11.7Document Processing 14.7Content Production 12.3

Each test is LLM-judged 0–20 against a fixed rubric; suite max = tests × 20. An ADJ chip marks a cell carrying a subjective adjustment from the maintainer; the judged score is the one shown, is unchanged, and remains what the axes, ranks and tie bands are computed from. Hover for the reason. Rows with a ▶ expand to the per-test breakdown. Where fewer tests ran than the suite total, unrun tests count as zero toward the axis score. See methodology.