← Leaderboard

Nemotron 3 Nano 30B-A3B 4bit MLX

Tested Rank #37 of 51 · 5/8 GAUNTLET progress
Generalist — 78.5 / 100 G Agentic — 92.5 / 100 A Understanding — pending U Needle — 75 / 100 N Thinking — 96.7 / 100 T Live — pending L Engineering — pending E Throughput — 39 / 100 T

Hollow markers & dashed spokes: axis not yet scored

G
78.5
A
92.5
U
—
N
75
T
96.7
L
—
E
—
T
39

Specification

Parameters
30B total / ~3B active per token
Architecture
nemotron3_nano (hybrid Mamba2/attention MoE)
Size on disk
17.79 GB
Quantization
4bit
Format
MLX
Runtime
LM Studio
Class
Standard
Reasoning (CoT)
Yes — emits reasoning tokens
Internal ID
M14
Mean speed
137.2 tok/s across suites
Stall census
1 stall in 30 observed tests (3.3%)
Reasoning appetite
1,413 tokens mean · 16,383 max
Model card
huggingface.co

Suite results

Suite Score Avg / 20 Tests n tok/s Run
General capability (13-task real-workload suite) 1a 204.1 / 260 15.7 13/13 n=1 142.6 run 2026-07-27
Test Category Score tok/s
A1 Dev/Ops Scripting 11/20 103.4
A2 Dev/Ops Scripting 18/20 137.2
A3 Dev/Ops Scripting 11/20 124.4
A4 Dev/Ops Scripting 16/20 113.7
A5 Dev/Ops Scripting 18/20 142
A6 Dev/Ops Scripting 0/20 112.7
B1 Document Processing 20/20 109.6
B2 Document Processing 20/20 128.4
B3 Document Processing 17/20 141.1
C1 Content Production 18/20 150.9
C2 Content Production 18/20 151.3
C3 Content Production 19/20 153.6
C4 Content Production 18/20 152.6

Category mean Dev/Ops Scripting 12.3Document Processing 19Content Production 18.3

Agentic tool-calling & protocol adherence 1b 148 / 160 18.5 8/8 n=1 162.8 run 2026-07-28
Test Category Score tok/s
WA1 Tool Calling & Protocol Adherence 20/20 17.3
WA2 Tool Calling & Protocol Adherence 20/20 83.4
WA3 Tool Calling & Protocol Adherence 20/20 118.4
WA4 Tool Calling & Protocol Adherence 20/20 77.2
WA5 Tool Calling & Protocol Adherence 20/20 111.6
WB1 Agent Robustness & State Management 17/20 156.1
WB2 Agent Robustness & State Management 20/20 155.8
WB3 Agent Robustness & State Management 11/20 142.1

Category mean Tool Calling & Protocol Adherence 20Agent Robustness & State Management 16

Long-context retrieval & synthesis 1g 120 / 120 20 6/6 n=1 140 run 2026-08-06
Test Category Score tok/s
L1 Long-Context — Needle Retrieval 20/20 61.3
L2 Long-Context — Grounded QA with Distractors 20/20 62.7
L3 Long-Context — Ops Log Reasoning 20/20 64.4
L4 Long-Context — Faithful Summarization 20/20 81.6
L5 Long-Context — Instruction Retention 20/20 52.4
L6 Long-Context — Full-Haystack QA 20/20 41.7

Category mean Long-Context — Needle Retrieval 20Long-Context — Grounded QA with Distractors 20Long-Context — Ops Log Reasoning 20Long-Context — Faithful Summarization 20Long-Context — Instruction Retention 20Long-Context — Full-Haystack QA 20

Long-context multi-needle (MRCR) 1g2 15 / 60 5 3/3 n=1 103.3 run 2026-08-14
Test Category Score tok/s
L7 Long-Context — Multi-Needle Retrieval (MRCR) 11/20 58.2
L8 Long-Context — Multi-Needle Retrieval (MRCR) 0/20 43.8
L9 Long-Context — Multi-Needle Retrieval (MRCR) 4/20 35.3

Each test is LLM-judged 0–20 against a fixed rubric; suite max = tests × 20. An ADJ chip marks a cell carrying a subjective adjustment from the maintainer; the judged score is the one shown, is unchanged, and remains what the axes, ranks and tie bands are computed from. Hover for the reason. Rows with a ▶ expand to the per-test breakdown. Where fewer tests ran than the suite total, unrun tests count as zero toward the axis score. See methodology.