← Leaderboard

GPT-OSS 120B

Keeper Rank #25 of 51 · 7/8 GAUNTLET progress
Generalist — 65.1 / 100 G Agentic — 93 / 100 A Understanding — pending U Needle — 86.7 / 100 N Thinking — 100 / 100 T Live — 85 / 100 L Engineering — 90.5 / 100 E Throughput — 18.2 / 100 T

Hollow markers & dashed spokes: axis not yet scored

G
65.1
A
93
U
—
N
86.7
T
100
L
85
E
90.5
T
18.2

Specification

Parameters
120B
Architecture
gpt_oss
Size on disk
63.41 GB
Quantization
mxfp4
Format
MLX
Runtime
LM Studio
Class
Standard
Reasoning (CoT)
No
Internal ID
G4
Mean speed
64 tok/s across suites
Stall census
0 stalls in 15 observed tests (0.0%)
Model card
huggingface.co

Suite results

Suite Score Avg / 20 Tests n tok/s Run
General capability (13-task real-workload suite) 1a 169.2 / 260 18.8 9/13 ● n=1 78.4 run 2026-09-02
Test Category Score tok/s
A1 Dev/Ops Scripting 20/20 49.9
A2 Dev/Ops Scripting 18/20 35.7
A3 Dev/Ops Scripting 18/20 52.8
A4 Dev/Ops Scripting 20/20 21
A5 Dev/Ops Scripting 14/20 53.9
A6 Dev/Ops Scripting 20/20 56.2
B1 Document Processing 20/20 45.1
B2 Document Processing 20/20 23.8
B3 Document Processing 20/20 21.3

Category mean Dev/Ops Scripting 18.3Document Processing 20

Agentic tool-calling & protocol adherence 1b 148.8 / 160 18.6 8/8 n=1 55.9 run 2026-09-16
Test Category Score tok/s
WA1 Tool Calling & Protocol Adherence 19/20 12.7
WA2 Tool Calling & Protocol Adherence 20/20 13.1
WA3 Tool Calling & Protocol Adherence 20/20 17.8
WA4 Tool Calling & Protocol Adherence 19/20 16.4
WA5 Tool Calling & Protocol Adherence 20/20 38.3
WB1 Agent Robustness & State Management 20/20 44.6
WB2 Agent Robustness & State Management 20/20 36.1
WB3 Agent Robustness & State Management 11/20 11.1

Category mean Tool Calling & Protocol Adherence 19.6Agent Robustness & State Management 17

Coding depth 1c 162.9 / 180 18.1 9/9 n=1 54.7 run 2026-09-15
Test Category Score tok/s
K1 Coding — Concurrency 20/20 42.8
K2 Coding — Algorithms 18/20 48.2
K3 Coding — Security Review 16/20 34.7
K4 Coding — Performance 19/20 47
K5 Coding — Refactoring Under Constraints 20/20 52.7
K6 Coding — Test Writing 20/20 40.4
K7 Coding — React Component (typed) 11/20 33.3
K8 Coding — TypeScript Type System 19/20 26.6
K9 Coding — React Debugging 20/20 25.3

Category mean Coding — Concurrency 20Coding — Algorithms 18Coding — Security Review 16Coding — Performance 19Coding — Refactoring Under Constraints 20Coding — Test Writing 20Coding — React Component (typed) 11Coding — TypeScript Type System 19Coding — React Debugging 20

Content-production depth 1e 112.8 / 120 18.8 6/6 n=1 82.9 run 2026-08-02
Test Category Score tok/s
E1 SOW & Requirements Decomposition 19/20 53.5
E2 SOW & Requirements Decomposition 20/20 35.9
E3 Executive & Proposal Content 18/20 71.4
E4 Executive & Proposal Content 17/20 64.4
E5 Executive & Proposal Content 19/20 55.5
E6 Executive & Proposal Content 20/20 29.4

Category mean SOW & Requirements Decomposition 19.5Executive & Proposal Content 18.5

Long-context retrieval & synthesis 1g 120 / 120 20 6/6 n=1 64.1 run 2026-09-10
Test Category Score tok/s
L1 Long-Context — Needle Retrieval 20/20 2.5
L2 Long-Context — Grounded QA with Distractors 20/20 1.7
L3 Long-Context — Ops Log Reasoning 20/20 5.1
L4 Long-Context — Faithful Summarization 20/20 12.7
L5 Long-Context — Instruction Retention 20/20 3.2
L6 Long-Context — Full-Haystack QA 20/20 1.2

Category mean Long-Context — Needle Retrieval 20Long-Context — Grounded QA with Distractors 20Long-Context — Ops Log Reasoning 20Long-Context — Faithful Summarization 20Long-Context — Instruction Retention 20Long-Context — Full-Haystack QA 20

Long-context multi-needle (MRCR) 1g2 36 / 60 12 3/3 n=1 47.2 run 2026-08-31
Test Category Score tok/s
L7 Long-Context — Multi-Needle Retrieval (MRCR) 20/20 12.8
L8 Long-Context — Multi-Needle Retrieval (MRCR) 13/20 8.5
L9 Long-Context — Multi-Needle Retrieval (MRCR) 3/20 2.3
Live one-shot builds (runtime-verified) 1h 68 / 80 17 4/4 n=1 73 run 2026-09-02
Test Category Score tok/s
H1 Web-Estate One-Shot Builds 19/20 61.8
H2 Web-Estate One-Shot Builds 16/20 43
H3 Web-Estate One-Shot Builds 17/20 63.9
H4 Web-Estate One-Shot Builds 16/20 27.7
Brand-constrained static build (negative-constraint adherence) 1j 85 / 100 17 5/5 n=1 55.8 run 2026-09-16
Test Category Score tok/s
J1 Brand-Constrained Static Build 18/20 25.9
J3 Brand-Constrained Static Build 20/20 33.5
J4 Brand-Constrained Static Build 20/20 25.1
J5 Brand-Constrained Static Build 14/20 19.3
J6 Brand-Constrained Static Build 13/20 23

Each test is LLM-judged 0–20 against a fixed rubric; suite max = tests × 20. An ADJ chip marks a cell carrying a subjective adjustment from the maintainer; the judged score is the one shown, is unchanged, and remains what the axes, ranks and tie bands are computed from. Hover for the reason. Rows with a ▶ expand to the per-test breakdown. Where fewer tests ran than the suite total, unrun tests count as zero toward the axis score. See methodology.