← Leaderboard
Gemma-4 26B-A4B 8bit MLX
Daily driver RAN THE GAUNTLET Rank #2 of 33 · 8/8 GAUNTLET progress
G
95.5
A
98
U
79.7
N
77.8
T
88
L
55.5
E
91
T
40.6
Specification
- Parameters
- 26B total / 4B active per token
- Architecture
- gemma4
- Size on disk
- 28 GB
- Quantization
- 8bit
- Format
- MLX
- Reasoning (CoT)
- No
- Internal ID
- M2
- Mean speed
- 66.2 tok/s across suites
- Stall census
- 10 stalls in 83 observed tests (12.0%)
- Reasoning appetite
- 3,860 tokens mean · 16,381 max
- Model card
- lmstudio.ai
Suite results
| Suite | Score | Avg / 20 | Tests | tok/s |
|---|---|---|---|---|
| General capability (13-task real-workload suite) 1a | 248.3 / 260 | 19.1 | 13/13 | 78.1 |
| Agentic tool-calling & protocol adherence 1b | 156.8 / 160 | 19.6 | 8/8 | 70.5 |
| Coding depth 1c | 109.2 / 120 | 18.2 | 6/6 | 66.1 |
| Doc/OCR vision 1d | 156 / 160 | 19.5 | 8/8 | 78 |
| Doc/OCR — real-degraded tier 1d2 | 35.2 / 80 | 8.8 | 4/4 | 68 |
| Content-production depth 1e | 111 / 120 | 18.5 | 6/6 | 80.9 |
| Long-context retrieval & synthesis 1g | 120 / 120 | 20 | 6/6 | 53.4 |
| Long-context multi-needle (MRCR) 1g2 | 20.1 / 60 | 6.7 | 3/3 | 46.2 |
| Live one-shot builds (runtime-verified) 1h | 62.4 / 80 | 15.6 | 4/4 | — |
| Production replay (real agent workload) 1i | 115.2 / 240 | 9.6 | 12/12 | 54.9 |
Each test is LLM-judged 0–20 against a fixed rubric; suite max = tests × 20. Where fewer tests ran than the suite total, unrun tests count as zero toward the axis score. See methodology.