Hollow markers & dashed spokes: axis not yet scored
G
87
A
99
U
—
N
—
T
100
L
—
E
—
T
19.8
Specification
- Parameters
- 118B total / 8B active (MoE, 10-of-256 experts + 1 shared)
- Architecture
- laguna
- Size on disk
- 71.2 GB
- Quantization
- Q4_K_M
- Format
- GGUF
- Runtime
- LM Studio
- Class
- Standard
- Reasoning (CoT)
- No
- Internal ID
- M20
- Mean speed
- 69.6 tok/s across suites
- Stall census
- 0 stalls in 29 observed tests (0.0%)
- Model card
- huggingface.co
Suite results
Suite Score Avg / 20 Tests n tok/s Run
General capability (13-task real-workload suite) 1a 226.2 / 260 17.4 13/13 n=1.2 70 run 2026-08-04
| Test | Category | Score | tok/s | |
|---|---|---|---|---|
| A1 | Dev/Ops Scripting | 17/20 | 64.9 | |
| A2 | Dev/Ops Scripting | 18/20 | 54.5 | |
| A3 | Dev/Ops Scripting | 7/20 | 64.1 | |
| A4 | Dev/Ops Scripting | 19/20 | 57.6 | |
| A5 | Dev/Ops Scripting | 18/20 | 60.3 | |
| A6 | Dev/Ops Scripting | 20/20 | 59.7 | |
| B1 | Document Processing | 20/20 | 56.1 | |
| B2 | Document Processing | 20/20 | 60.7 | |
| B3 | Document Processing | 16/20 | 62.4 | |
| C1 | Content Production | 20/20 | 61.4 | |
| C2 | Content Production | 18/20 | 60.7 | |
| C3 | Content Production | 15/20 | 63.7 | |
| C4 | Content Production | 20/20 | 47.4 |
Category mean Dev/Ops Scripting 16.5Document Processing 18.7Content Production 18.3
Agentic tool-calling & protocol adherence 1b 158.4 / 160 19.8 8/8 n=1 69.2 run 2026-08-04
| Test | Category | Score | tok/s | |
|---|---|---|---|---|
| WA1 | Tool Calling & Protocol Adherence | 20/20 | 20.5 | |
| WA2 | Tool Calling & Protocol Adherence | 20/20 | 16.7 | |
| WA3 | Tool Calling & Protocol Adherence | 20/20 | 17.4 | |
| WA4 | Tool Calling & Protocol Adherence | 20/20 | 21.6 | |
| WA5 | Tool Calling & Protocol Adherence | 20/20 | 50.3 | |
| WB1 | Agent Robustness & State Management | 19/20 | 58.5 | |
| WB2 | Agent Robustness & State Management | 20/20 | 55.6 | |
| WB3 | Agent Robustness & State Management | 19/20 | 58.5 |
Category mean Tool Calling & Protocol Adherence 20Agent Robustness & State Management 19.3
Each test is LLM-judged 0–20 against a fixed rubric; suite max = tests × 20. An ADJ chip marks a cell carrying a subjective adjustment from the maintainer; the judged score is the one shown, is unchanged, and remains what the axes, ranks and tie bands are computed from. Hover for the reason. Rows with a ▶ expand to the per-test breakdown. Where fewer tests ran than the suite total, unrun tests count as zero toward the axis score. See methodology.