Hollow markers & dashed spokes: axis not yet scored
G
53
A
—
U
—
N
—
T
n/d
L
—
E
—
T
44
Specification
- Parameters
- 20B
- Architecture
- gpt_oss
- Size on disk
- 11.15 GB
- Quantization
- Format
- MLX
- Runtime
- LM Studio
- Class
- Standard
- Reasoning (CoT)
- Yes — emits reasoning tokens
- Internal ID
- G3
- Mean speed
- 154.5 tok/s across suites
- Stall census
- 0 stalls in 9 observed tests (0.0%)
- Model card
- huggingface.co
Suite results
Suite Score Avg / 20 Tests n tok/s Run
General capability (13-task real-workload suite) 1a 137.7 / 260 15.3 9/13 ● n=1 154.5 run 2026-09-02
| Test | Category | Score | tok/s | |
|---|---|---|---|---|
| A1 | Dev/Ops Scripting | 14/20 | 46.9 | |
| A2 | Dev/Ops Scripting | 17/20 | 76.7 | |
| A3 | Dev/Ops Scripting | 5/20 | 45.2 | |
| A4 | Dev/Ops Scripting | 19/20 | 63.1 | |
| A5 | Dev/Ops Scripting | 16/20 | 71 | |
| A6 | Dev/Ops Scripting | 18/20 | 77.9 | |
| B1 | Document Processing | 20/20 | 77.6 | |
| B2 | Document Processing | 19/20 | 40.6 | |
| B3 | Document Processing | 10/20 | 22.9 |
Category mean Dev/Ops Scripting 14.8Document Processing 16.3
Each test is LLM-judged 0–20 against a fixed rubric; suite max = tests × 20. An ADJ chip marks a cell carrying a subjective adjustment from the maintainer; the judged score is the one shown, is unchanged, and remains what the axes, ranks and tie bands are computed from. Hover for the reason. Rows with a ▶ expand to the per-test breakdown. Where fewer tests ran than the suite total, unrun tests count as zero toward the axis score. See methodology.