Hollow markers & dashed spokes: axis not yet scored
Specification
- Parameters
- 35.95B total / ~3B active per token (MoE, 256 experts, 8 used)
- Architecture
- qwen35moe
- Size on disk
- 20 GB
- Quantization
- Q4_K_M
- Format
- GGUF
- Class
- Standard
- Reasoning (CoT)
- Yes — emits reasoning tokens
- Internal ID
- M34
- Mean speed
- 82 tok/s across suites
- Model card
- huggingface.co
Suite results
General capability (13-task real-workload suite) 1a 235.3 / 260 18.1 13/13 84.3 run 2026-09-06
| Test | Category | Score | tok/s | |
|---|---|---|---|---|
| A1 | Dev/Ops Scripting | 20/20 | 92.5 | |
| A2 | Dev/Ops Scripting | 18/20 | 99.5 | |
| A3 | Dev/Ops Scripting | 20/20 | 101 | |
| A4 | Dev/Ops Scripting | 19/20 | 82.2 | |
| A5 | Dev/Ops Scripting | 20/20 | 76 | |
| A6 | Dev/Ops Scripting | 20/20 | 81 | |
| B1 | Document Processing | 20/20 | 100.2 | |
| B2 | Document Processing | 20/20 | 82.7 | |
| B3 | Document Processing | 0/20 | — | |
| C1 | Content Production | 20/20 | 75.1 | |
| C2 | Content Production | 18/20 | 73.5 | |
| C3 | Content Production | 20/20 | 72.4 | |
| C4 | Content Production | 20/20 | 75.4 |
Category mean Dev/Ops Scripting 19.5Document Processing 13.3Content Production 19.5
Agentic tool-calling & protocol adherence 1b 158.4 / 160 19.8 8/8 94.5 run 2026-09-06
| Test | Category | Score | tok/s | |
|---|---|---|---|---|
| WA1 | Tool Calling & Protocol Adherence | 20/20 | 96.2 | |
| WA2 | Tool Calling & Protocol Adherence | 20/20 | 88 | |
| WA3 | Tool Calling & Protocol Adherence | 20/20 | 93.7 | |
| WA4 | Tool Calling & Protocol Adherence | 20/20 | 73.8 | |
| WA5 | Tool Calling & Protocol Adherence | 20/20 | 103.6 | |
| WB1 | Agent Robustness & State Management | 20/20 | 95.5 | |
| WB2 | Agent Robustness & State Management | 20/20 | 109.3 | |
| WB3 | Agent Robustness & State Management | 18/20 | 96.2 |
Category mean Tool Calling & Protocol Adherence 20Agent Robustness & State Management 19.3
Coding depth 1c 120 / 120 20 6/6 97.2 run 2026-09-06
| Test | Category | Score | tok/s | |
|---|---|---|---|---|
| K1 | Coding — Concurrency | 20/20 | 101.7 | |
| K2 | Coding — Algorithms | 20/20 | 123.3 | |
| K3 | Coding — Security Review | 20/20 | 101.2 | |
| K4 | Coding — Performance | 20/20 | 87.6 | |
| K5 | Coding — Refactoring Under Constraints | 20/20 | 79.6 | |
| K6 | Coding — Test Writing | 20/20 | 89.6 |
Category mean Coding — Concurrency 20Coding — Algorithms 20Coding — Security Review 20Coding — Performance 20Coding — Refactoring Under Constraints 20Coding — Test Writing 20
Doc/OCR vision 1d 160 / 160 20 8/8 51.1 run 2026-09-06
| Test | Category | Score | tok/s | |
|---|---|---|---|---|
| D1 | Doc/OCR — Transcription | 20/20 | 36.7 | |
| D2 | Doc/OCR — Document QA | 20/20 | 32.3 | |
| D3 | Doc/OCR — Field Extraction | 20/20 | 54.8 | |
| D4 | Doc/OCR — Document QA | 20/20 | 44.9 | |
| D5 | Doc/OCR — Transcription | 20/20 | 61.8 | |
| D6 | Doc/OCR — Field Extraction | 20/20 | 35 | |
| D7 | Doc/OCR — Degraded Input | 20/20 | 72 | |
| D8 | Doc/OCR — Degraded Input | 20/20 | 71 |
Category mean Doc/OCR — Transcription 20Doc/OCR — Document QA 20Doc/OCR — Field Extraction 20Doc/OCR — Degraded Input 20
Doc/OCR — real-degraded tier 1d2 51.2 / 80 12.8 4/4 75 run 2026-09-06
| Test | Category | Score | tok/s | |
|---|---|---|---|---|
| D9 | Doc/OCR — Real Degraded Transcription | 7/20 | 79.1 | |
| D10 | Doc/OCR — Handwriting Extraction | 17/20 | 52.8 | |
| D11 | Doc/OCR — Degraded Table Transcription | 14/20 | 92.1 | |
| D12 | Doc/OCR — Degraded Table QA | 13/20 | 76.2 |
Category mean Doc/OCR — Real Degraded Transcription 7Doc/OCR — Handwriting Extraction 17Doc/OCR — Degraded Table Transcription 14Doc/OCR — Degraded Table QA 13
Live one-shot builds (runtime-verified) 1h 59.1 / 80 19.7 3/4 90.1 run 2026-09-06
| Test | Category | Score | tok/s | |
|---|---|---|---|---|
| H1 | Web-Estate One-Shot Builds | 19/20 | 103.6 | |
| H2 | Web-Estate One-Shot Builds | 20/20 | 84.9 | |
| H4 | Web-Estate One-Shot Builds | 20/20 | 81.7 |
Each test is LLM-judged 0–20 against a fixed rubric; suite max = tests × 20. Rows with a ▶ expand to the per-test breakdown. Where fewer tests ran than the suite total, unrun tests count as zero toward the axis score. See methodology.