← Leaderboard

GPT-OSS 20B (OpenAI mxfp4)

Keeper Rank #33 of 51 · 6/8 GAUNTLET progress
Generalist — 62.3 / 100 G Agentic — 65 / 100 A Understanding — pending U Needle — 65.7 / 100 N Thinking — pending T Live — 62.4 / 100 L Engineering — 90 / 100 E Throughput — 33.4 / 100 T

Hollow markers & dashed spokes: axis not yet scored

G
62.3
A
65
U
—
N
65.7
T
n/d
L
62.4
E
90
T
33.4

Specification

Parameters
20B
Architecture
gpt_oss
Size on disk
12.1 GB
Quantization
mxfp4
Format
MLX
Runtime
LM Studio
Class
Standard
Reasoning (CoT)
No
Internal ID
G2
Mean speed
117.3 tok/s across suites
Stall census
0 stalls in 10 observed tests (0.0%)
Model card
lmstudio.ai

Suite results

Suite Score Avg / 20 Tests n tok/s Run
General capability (13-task real-workload suite) 1a 162 / 260 18 9/13 ● n=1.3 128.2 run 2026-09-02
Test Category Score tok/s
A1 Dev/Ops Scripting 10/20 75.3
A2 Dev/Ops Scripting 19/20 82.3
A3 Dev/Ops Scripting 18/20 97.4
A4 Dev/Ops Scripting 18/20 89.3
A5 Dev/Ops Scripting 17/20 87
A6 Dev/Ops Scripting 20/20 104.2
B1 Document Processing 20/20 106.3
B2 Document Processing 20/20 72
B3 Document Processing 20/20 44.1

Category mean Dev/Ops Scripting 17Document Processing 20

Agentic tool-calling & protocol adherence 1b 104 / 160 13 8/8 n=1 143.8 run 2026-09-16
Test Category Score tok/s
WA1 Tool Calling & Protocol Adherence 7/20 52.5
WA2 Tool Calling & Protocol Adherence 4/20 42.9
WA3 Tool Calling & Protocol Adherence 17/20 59.1
WA4 Tool Calling & Protocol Adherence 11/20 42.4
WA5 Tool Calling & Protocol Adherence 20/20 98.9
WB1 Agent Robustness & State Management 20/20 128.4
WB2 Agent Robustness & State Management 19/20 100.9
WB3 Agent Robustness & State Management 6/20 53.8

Category mean Tool Calling & Protocol Adherence 11.8Agent Robustness & State Management 15

Coding depth 1c 162 / 180 18 9/9 n=1 124 run 2026-09-15
Test Category Score tok/s
K1 Coding — Concurrency 20/20 92.9
K2 Coding — Algorithms 20/20 105.4
K3 Coding — Security Review 17/20 102
K4 Coding — Performance 15/20 100.1
K5 Coding — Refactoring Under Constraints 20/20 122
K6 Coding — Test Writing 16/20 116.8
K7 Coding — React Component (typed) 14/20 111.5
K8 Coding — TypeScript Type System 20/20 107.2
K9 Coding — React Debugging 20/20 107.6

Category mean Coding — Concurrency 20Coding — Algorithms 20Coding — Security Review 17Coding — Performance 15Coding — Refactoring Under Constraints 20Coding — Test Writing 16Coding — React Component (typed) 14Coding — TypeScript Type System 20Coding — React Debugging 20

Content-production depth 1e 91.8 / 120 15.3 6/6 n=1 131.6 run 2026-09-10
Test Category Score tok/s
E1 SOW & Requirements Decomposition 19/20 118.3
E2 SOW & Requirements Decomposition 20/20 83.8
E3 Executive & Proposal Content 17/20 122
E4 Executive & Proposal Content 7/20 121.9
E5 Executive & Proposal Content 15/20 114.5
E6 Executive & Proposal Content 14/20 58.5

Category mean SOW & Requirements Decomposition 19.5Executive & Proposal Content 13.3

Long-context retrieval & synthesis 1g 97.8 / 120 16.3 6/6 n=1 104.7 run 2026-09-10
Test Category Score tok/s
L1 Long-Context — Needle Retrieval 20/20 9.4
L2 Long-Context — Grounded QA with Distractors 20/20 4.7
L3 Long-Context — Ops Log Reasoning 12/20 9.3
L4 Long-Context — Faithful Summarization 6/20 28.2
L5 Long-Context — Instruction Retention 20/20 9.9
L6 Long-Context — Full-Haystack QA 20/20 1.7

Category mean Long-Context — Needle Retrieval 20Long-Context — Grounded QA with Distractors 20Long-Context — Ops Log Reasoning 12Long-Context — Faithful Summarization 6Long-Context — Instruction Retention 20Long-Context — Full-Haystack QA 20

Long-context multi-needle (MRCR) 1g2 20.4 / 60 6.8 3/3 n=1.7 71.5 run 2026-09-07
Test Category Score tok/s
L7 Long-Context — Multi-Needle Retrieval (MRCR) 7.5/20 26.5
L8 Long-Context — Multi-Needle Retrieval (MRCR) 9/20 19.6
L9 Long-Context — Multi-Needle Retrieval (MRCR) 4/20 3.3
Live one-shot builds (runtime-verified) 1h 56.8 / 80 14.2 4/4 n=1 125.8 run 2026-08-27
Test Category Score tok/s
H1 Web-Estate One-Shot Builds 15/20 120.8
H2 Web-Estate One-Shot Builds 15/20 123
H3 Web-Estate One-Shot Builds 13/20 116.8
H4 Web-Estate One-Shot Builds 14/20 108.4
Production replay (real agent workload) 1i 116.4 / 240 9.7 12/12 n=1 94.3 run 2026-09-26
Test Category Score tok/s
I1 Nightly Fix Replay 7/20 19.3
I2 Nightly Fix Replay 1/20 14.3
I3 Nightly Fix Replay 1/20 27.1
I4 Nightly Fix Replay 5/20 13
I5 Nightly Fix Replay 8/20 51.3
I6 Nightly Fix Replay 4/20 52.4
I7 Audit Judgment 15/20 46.9
I8 Audit Judgment 18/20 96
I9 Cron Dependency Reasoning 18/20 78.4
I10 Cron Dependency Reasoning 13/20 85.3
I11 Curation Judgment 12/20 98
I12 Curation Judgment 14/20 103.1

Category mean Nightly Fix Replay 4.3Audit Judgment 16.5Cron Dependency Reasoning 15.5Curation Judgment 13

Brand-constrained static build (negative-constraint adherence) 1j 89 / 100 17.8 5/5 n=1 132.1 run 2026-09-16
Test Category Score tok/s
J1 Brand-Constrained Static Build 18/20 81.5
J3 Brand-Constrained Static Build 20/20 104
J4 Brand-Constrained Static Build 20/20 86.9
J5 Brand-Constrained Static Build 14/20 103.4
J6 Brand-Constrained Static Build 17/20 104.6

Each test is LLM-judged 0–20 against a fixed rubric; suite max = tests × 20. An ADJ chip marks a cell carrying a subjective adjustment from the maintainer; the judged score is the one shown, is unchanged, and remains what the axes, ranks and tie bands are computed from. Hover for the reason. Rows with a ▶ expand to the per-test breakdown. Where fewer tests ran than the suite total, unrun tests count as zero toward the axis score. See methodology.