← Leaderboard

Qwen3-VL 30B-A3B Instruct 4bit MLX

Tested Rank #27 of 51 · 7/8 GAUNTLET progress
Generalist — pending G Agentic — 85 / 100 A Understanding — 83.3 / 100 U Needle — 82.8 / 100 N Thinking — 100 / 100 T Live — 66.7 / 100 L Engineering — 84 / 100 E Throughput — 31 / 100 T

Hollow markers & dashed spokes: axis not yet scored

G
—
A
85
U
83.3
N
82.8
T
100
L
66.7
E
84
T
31

Specification

Parameters
30B total / 3B active per token
Architecture
qwen3_vl_moe
Size on disk
18 GB
Quantization
4bit
Format
MLX
Runtime
LM Studio
Class
Standard
Reasoning (CoT)
No
Internal ID
M23
Mean speed
109 tok/s across suites
Stall census
0 stalls in 17 observed tests (0.0%)
Model card
lmstudio.ai

Suite results

Suite Score Avg / 20 Tests n tok/s Run
Agentic tool-calling & protocol adherence 1b 136 / 160 17 8/8 n=1 133.4 run 2026-09-22
Test Category Score tok/s
WA1 Tool Calling & Protocol Adherence 13/20 57.6
WA2 Tool Calling & Protocol Adherence 20/20 67.3
WA3 Tool Calling & Protocol Adherence 15/20 71.9
WA4 Tool Calling & Protocol Adherence 20/20 55.2
WA5 Tool Calling & Protocol Adherence 20/20 107
WB1 Agent Robustness & State Management 20/20 124.1
WB2 Agent Robustness & State Management 20/20 119.8
WB3 Agent Robustness & State Management 8/20 85.3

Category mean Tool Calling & Protocol Adherence 17.6Agent Robustness & State Management 16

Coding depth 1c 151.2 / 180 16.8 9/9 n=1 129 run 2026-10-01
Test Category Score tok/s
K1 Coding — Concurrency 18/20 120.9
K2 Coding — Algorithms 20/20 116
K3 Coding — Security Review 14/20 119.6
K4 Coding — Performance 19/20 119.7
K5 Coding — Refactoring Under Constraints 20/20 125.1
K6 Coding — Test Writing 15/20 119.1
K7 Coding — React Component (typed) 10/20 126
K8 Coding — TypeScript Type System 17/20 118
K9 Coding — React Debugging 18/20 124.3

Category mean Coding — Concurrency 18Coding — Algorithms 20Coding — Security Review 14Coding — Performance 19Coding — Refactoring Under Constraints 20Coding — Test Writing 15Coding — React Component (typed) 10Coding — TypeScript Type System 17Coding — React Debugging 18

Doc/OCR vision 1d 151.2 / 160 18.9 8/8 n=1.1 120.2 run 2026-08-06
Test Category Score tok/s
D1 Doc/OCR — Transcription 20/20 68.2
D2 Doc/OCR — Document QA 20/20 19.6
D3 Doc/OCR — Field Extraction 20/20 60.2
D4 Doc/OCR — Document QA 11/20 32.6
D5 Doc/OCR — Transcription 20/20 59
D6 Doc/OCR — Field Extraction 20/20 42
D7 Doc/OCR — Degraded Input 20/20 75.4
D8 Doc/OCR — Degraded Input 20/20 70.5

Category mean Doc/OCR — Transcription 20Doc/OCR — Document QA 15.5Doc/OCR — Field Extraction 20Doc/OCR — Degraded Input 20

Doc/OCR — real-degraded tier 1d2 48.8 / 80 12.2 4/4 n=1 119.7 run 2026-08-14
Test Category Score tok/s
D9 Doc/OCR — Real Degraded Transcription 6/20 40.4
D10 Doc/OCR — Handwriting Extraction 16/20 26
D11 Doc/OCR — Degraded Table Transcription 10/20 84
D12 Doc/OCR — Degraded Table QA 17/20 29.9

Category mean Doc/OCR — Real Degraded Transcription 6Doc/OCR — Handwriting Extraction 16Doc/OCR — Degraded Table Transcription 10Doc/OCR — Degraded Table QA 17

Content-production depth 1e 97.8 / 120 16.3 6/6 n=1 129.4 run 2026-09-25
Test Category Score tok/s
E1 SOW & Requirements Decomposition 20/20 118.9
E2 SOW & Requirements Decomposition 18/20 111.7
E3 Executive & Proposal Content 16/20 120.4
E4 Executive & Proposal Content 13/20 126.6
E5 Executive & Proposal Content 16/20 120.6
E6 Executive & Proposal Content 15/20 116.2

Category mean SOW & Requirements Decomposition 19Executive & Proposal Content 15

Long-context retrieval & synthesis 1g 120 / 120 20 6/6 n=1 71.6 run 2026-09-24
Test Category Score tok/s
L1 Long-Context — Needle Retrieval 20/20 11.7
L2 Long-Context — Grounded QA with Distractors 20/20 4.6
L3 Long-Context — Ops Log Reasoning 20/20 7
L4 Long-Context — Faithful Summarization 20/20 29.1
L5 Long-Context — Instruction Retention 20/20 6.6
L6 Long-Context — Full-Haystack QA 20/20 1.6

Category mean Long-Context — Needle Retrieval 20Long-Context — Grounded QA with Distractors 20Long-Context — Ops Log Reasoning 20Long-Context — Faithful Summarization 20Long-Context — Instruction Retention 20Long-Context — Full-Haystack QA 20

Long-context multi-needle (MRCR) 1g2 29.1 / 60 9.7 3/3 n=1 42.4 run 2026-09-19
Test Category Score tok/s
L7 Long-Context — Multi-Needle Retrieval (MRCR) 20/20 19.2
L8 Long-Context — Multi-Needle Retrieval (MRCR) 5/20 9.7
L9 Long-Context — Multi-Needle Retrieval (MRCR) 4/20 6.1
Live one-shot builds (runtime-verified) 1h 42 / 80 10.5 4/4 n=1 120.7 run 2026-09-21
Test Category Score tok/s
H1 Web-Estate One-Shot Builds 14/20 120.1
H2 Web-Estate One-Shot Builds 10/20 124.4
H3 Web-Estate One-Shot Builds 9/20 119.2
H4 Web-Estate One-Shot Builds 9/20 112.2
Brand-constrained static build (negative-constraint adherence) 1j 78 / 100 15.6 5/5 n=1 114.4 run 2026-09-23
Test Category Score tok/s
J1 Brand-Constrained Static Build 18/20 104.2
J3 Brand-Constrained Static Build 14/20 108
J4 Brand-Constrained Static Build 18/20 104
J5 Brand-Constrained Static Build 14/20 110.8
J6 Brand-Constrained Static Build 14/20 118.1

Each test is LLM-judged 0–20 against a fixed rubric; suite max = tests × 20. An ADJ chip marks a cell carrying a subjective adjustment from the maintainer; the judged score is the one shown, is unchanged, and remains what the axes, ranks and tie bands are computed from. Hover for the reason. Rows with a ▶ expand to the per-test breakdown. Where fewer tests ran than the suite total, unrun tests count as zero toward the axis score. See methodology.