Local-LLM proving ground
Proving local agents before they act.
GAUNTLET Bench runs local models through real workloads — agentic, live, verified — and publishes the evidence.
All scores measured on Apple M5 Max, 128 GB unified memory, LM Studio
Current podium
Gemma-4 26B-A4B 8bit MLX
Qwen3.6-27B Fable-Fusion-711 Uncensored Heretic NEO-MAX MTPLX 8bit MLX
Leaderboard — 51 models · click a column to sort
| #▲ | Model▼ | Card | Params▼ | Quant / Format▼ | G▼ | A▼ | U▼ | N▼ | T▼ | L▼ | E▼ | T▼ | tok/s▼ | Progress▼ | Evidence▼ |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| =1 | Qwen3.8-27B Q6_K GGUF | Card | 27B dense | Q6_K GGUF | 99 | 96 | 93.3 | 91.2 | 96.5 | 83 | 86.7 | 5.7 | 20 | 8/8 | n=1 ● |
| =1 | Gemma-4 26B-A4B 8bit MLX | Card | 26B total / 4B active per token | 8bit MLX | 95.5 | 99.5 | 81 | 77.8 | 88 | 68 | 98.5 | 19.8 | 69.6 | 8/8 | n=1 |
| =1 | Qwen3.6-27B Fable-Fusion-711 Uncensored Heretic NEO-MAX MTPLX 8bit MLX | Card | 27B | 8bit MLX | 95.5 | 99 | 90.8 | 73.5 | 91.5 | 53 | 88.4 | 4.5 | 15.7 | 8/8 | n=1 ● |
| =1 | Qwen3.8-27B Q4_K_M GGUF | Card | 27B dense | Q4_K_M GGUF | 95 | 100 | 95.3 | 91.2 | 95.7 | 85.6 | 94.5 | 7.6 | 26.7 | 8/8 | n=1 ● |
| =1 | Qwen3.6-35B-A3B 4bit MLX | Card | 35B total / 3B active per token | 4bit MLX | 94 | 78 | 93.7 | 91.2 | 97.1 | 76.7 | 96.5 | 30.7 | 108 | 8/8 | n=1 |
| =6 | Qwen3.6-35B-A3B Unsloth Dynamic UD-Q8_K_XL MLX (Brooooooklyn) | Card | 35B total / 3B active per token | UD-Q8_K_XL mixed-precision MLX | 93.5 | 96 | 79.7 | 91.2 | 100 | 83.3 | 98 | 23.7 | 83.2 | 8/8 | n=1 |
| =6 | Qwen3.6-27B Dense 8bit MLX | Card | 27B dense | 8bit MLX | 90.5 | 100 | 96.3 | 100 | 97.2 | 79 | 77.8 | 4.2 | 14.6 | 8/8 | n=1 ● |
| =6 | Qwen3.5-122B-A10B 4bit MLX | Card | 122B total / 10B active per token | 4bit MLX | 90 | 94.5 | 100 | 91.2 | 100 | 84.4 | 98.5 | 14.2 | 50 | 8/8 | n=1 |
| =9 | Gemma-4 E4B 4B 8bit MLX SMALL | Card | 4B | 8bit MLX | 84.5 | 94.5 | 82.3 | 83.8 | 100 | 79 | 85.5 | 22.3 | 78.3 | 8/8 | n=1 |
| =9 | Bonsai 27B Ternary 2bit MLX | Card | 27B | 2bit (ternary, PrismML extreme compression) MLX | 81 | 92 | 90.3 | 80 | 93.1 | 75.6 | 96 | 10.4 | 36.7 | 8/8 | n=1 |
| =9 | Mistral Small 3.2 24B 8bit MLX | Card | 24B | 8bit MLX | 80.5 | 91 | 79 | 80.2 | 100 | 62.8 | 80.5 | 5.2 | 18.2 | 8/8 | n=1 |
| =12 | Qwen3.8-27B Q6_K GGUF lmstudio-community (Ollama, MBP - identical-file arm of M25b) OLLAMA | Card | 27B dense | Q6_K GGUF | 98.5 | 100 | 94.7 | 91.2 | n/d | 95 | 97 | 4.4 | 15.6 | 7/8 | n=1 |
| =12 | Qwen3.8-27B GSQ-RCO IQ3_S GGUF ISTA-DASLab (Ollama, MBP — identical-file arm of M35) OLLAMA | Card | 27B dense | IQ3_S GGUF | 98 | 100 | 92 | 91.2 | n/d | 84.2 | 99 | 6.2 | 21.7 | 7/8 | n=1 |
| =12 | Qwen3.8-27B UD-Q6_K_XL GGUF (Ollama, identical-file control) OLLAMA | Card | 27B dense | Q6_K_XL GGUF | 95.5 | 100 | 95 | 91.2 | n/d | 95.3 | 98 | 4.3 | 15.2 | 7/8 | n=1 |
| =12 | Qwen3.8-27B nvfp4 MLX (Ollama, MBP) OLLAMA | Card | 27B dense | nvfp4 MLX | 95 | 100 | 93 | 89.8 | n/d | 48.6 | 97 | 13.7 | 48.3 | 7/8 | n=1 |
| =12 | Gemma-4 26B-A4B MLX (Ollama, MBP) OLLAMA | Card | 26B total / 4B active per token | MLX | 93.5 | 87.5 | 73 | 77.5 | n/d | 49.7 | 73.5 | 34.4 | 120.8 | 7/8 | n=1 |
| =17 | Ornith-1.5-35B-A3B Q4_K_M GGUF | Card | 35.95B total / ~3B active per token (MoE, 256 experts, 8 used) | Q4_K_M GGUF | 90.5 | 99 | 88 | 100 | n/d | 72.7 | 88.5 | 27.2 | 95.6 | 7/8 | n=1 ● |
| =17 | Qwen3.6-27B Fable-Fusion-711 Uncensored Heretic NEO-MAX Q8_0 GGUF | Card | 27B | Q8_0 GGUF | 89.5 | 100 | 84.2 | 91.2 | 97.7 | — | 97 | 3.7 | 13.1 | 7/8 | n=1 ● |
| =17 | Qwen3-Coder-Next 80B 4bit MLX | Card | 80B | 4bit MLX | 88.5 | 94.5 | — | 64.4 | 100 | 87.8 | 89.5 | 23.3 | 81.8 | 7/8 | n=1 ● |
| =17 | Ornith-1.0-35B 4bit MLX | Card | 35B total (MoE; active-per-token not confirmed in model card) | 4bit MLX | 86.3 | 98 | 94.7 | 90.2 | n/d | 90.4 | 99 | 35 | 122.9 | 7/8 | n=1 ● |
| =21 | Qwen3.8-27B MLX 4bit (lmstudio-community) | Card | 27B dense | 4bit MLX | 82.9 | 100 | 91.3 | 91.2 | n/d | 82.8 | 99 | 6.8 | 23.9 | 7/8 | n=1 ● |
| =21 | Gemma-4 31B IT 4bit MLX | Card | 31.27B | 4bit MLX | 82.1 | 100 | 85.4 | 91.2 | n/d | 82.2 | 98 | 5.1 | 18 | 7/8 | n=1 ● |
| =21 | Gemma-4 E4B 7.5B Q4_K_M GGUF SMALL | Card | 7.5B | Q4_K_M GGUF | 78 | 89.5 | 84.7 | 64.4 | n/d | 72.3 | 75 | 23.6 | 82.9 | 7/8 | n=1 ● |
| =24 | Gemma-3 12B IT QAT 4bit MLX | Card | 12.19B | 4bit (QAT) MLX | 67.5 | 92 | 54 | 67.8 | n/d | 57.8 | 69.5 | 16.1 | 56.4 | 7/8 | n=1 |
| =24 | GPT-OSS 120B | Card | 120B | mxfp4 MLX | 65.1 | 93 | — | 86.7 | 100 | 85 | 90.5 | 18.2 | 64 | 7/8 | n=1 ● |
| 26 | Gemma 3 4B QAT 4bit MLX SMALL | Card | 4B | QAT 4bit MLX | 57.5 | 86 | 49.3 | 46.8 | n/d | 37.9 | 51 | 46.3 | 162.6 | 7/8 | n=1 |
| 27 | Qwen3-VL 30B-A3B Instruct 4bit MLX | Card | 30B total / 3B active per token | 4bit MLX | — | 85 | 83.3 | 82.8 | 100 | 66.7 | 84 | 31 | 109 | 7/8 | n=1 |
| =28 | Qwen3.8-27B Uncensored HauhauCS Aggressive Q5_K_P GGUF (Ollama, MBP — identical-file arm of M29) OLLAMA | Card | 27B dense | Q5_K_P GGUF | 97.5 | 97 | — | 91.2 | n/d | 94.6 | 99 | 4.4 | 15.6 | 6/8 | n=1 |
| =28 | Qwen3.8-27B Uncensored HauhauCS Aggressive Q6_K_P GGUF (Ollama, MBP — identical-file arm of M27) OLLAMA | Card | 27B dense | Q6_K_P GGUF | 97.5 | 98 | — | 91.2 | n/d | 96.8 | 99 | 3.9 | 13.8 | 6/8 | n=1 |
| =28 | Qwen3.8-27B MLX 4bit lmstudio-community (Ollama, MBP — identical-weights arm of M25h) OLLAMA | Card | 27B dense | 4bit MLX | 93.5 | 100 | — | 91.2 | n/d | 85 | 97 | 7.2 | 25.1 | 6/8 | n=1 |
| =28 | Qwen3.8-27B MLX 8bit lmstudio-community (Ollama, MBP — quant-ladder rung) OLLAMA | Card | 27B dense | 8bit MLX | 93 | 100 | — | 91.2 | n/d | 96.6 | 99 | 4.5 | 15.9 | 6/8 | n=1 |
| 32 | Qwen3-4B-2507 (non-thinking) 4bit MLX SMALL | Card | 4B | 4bit MLX | 68 | 93 | — | 80 | n/d | 46.6 | 56 | 38 | 133.6 | 6/8 | n=1 |
| 33 | GPT-OSS 20B (OpenAI mxfp4) | Card | 20B | mxfp4 MLX | 62.3 | 65 | — | 65.7 | n/d | 62.4 | 90 | 33.4 | 117.3 | 6/8 | n=1 ● |
| 34 | DeepSeek-R1-0528-Qwen3-8B 4bit MLX | Card | 8.19B | 4bit MLX | 42 | 93 | — | 56.2 | n/d | 33.9 | 44 | 19.8 | 69.6 | 6/8 | n=1 ● |
| 35 | Phi-3.5-mini-instruct 4bit MLX SMALL | Card | 3.8B | 4bit MLX | 36 | 26 | — | 37.2 | n/d | 10.6 | 25 | 26 | 91.2 | 6/8 | n=1 ● |
| 36 | Qwen3.8-27B MLX 6bit (lmstudio-community) | Card | 27B dense | 6bit MLX | 90 | 97 | 74 | 73.5 | n/d | — | — | 4.4 | 15.4 | 5/8 | n=1 ● |
| 37 | Nemotron 3 Nano 30B-A3B 4bit MLX | Card | 30B total / ~3B active per token | 4bit MLX | 78.5 | 92.5 | — | 75 | 96.7 | — | — | 39 | 137.2 | 5/8 | n=1 |
| 38 | Gemma 3 1B QAT 4bit MLX SMALL | Card | 1B | QAT 4bit MLX | 28.5 | 33 | — | — | n/d | 7.8 | 18.5 | 100 | 351.3 | 5/8 | n=1 |
| 39 | Qwen3.8-27B Q8_0 GGUF | Card | 27B dense | Q8_0 GGUF | 94 | 95.5 | — | — | 100 | — | — | 7.8 | 27.6 | 4/8 | n=1 |
| 40 | Laguna S 2.1 (Q4_K_M GGUF) | Card | 118B total / 8B active (MoE, 10-of-256 experts + 1 shared) | Q4_K_M GGUF | 87 | 99 | — | — | 100 | — | — | 19.8 | 69.6 | 4/8 | n=1 |
| =41 | Devstral Small 2 24B 6bit MLX | Card | 24B | 6bit MLX | 78 | 97 | — | — | 100 | — | — | 6.3 | 22 | 4/8 | n=1 |
| =41 | Muse-Glimmer 30B (GGUF, kquant 17GB) | Card | 30B (unconfirmed -- filename-derived, see note) | kquant (unconfirmed exact scheme -- filename says "kquant", not a standard llama.cpp quant label) GGUF | 74 | 76 | — | — | 100 | — | — | 5.9 | 20.7 | 4/8 | n=1 |
| 43 | Qwen3.8-27B MLX 5bit (lmstudio-community) | Card | 27B dense | 5bit MLX | 74.6 | 100 | — | — | n/d | — | — | 4.5 | 15.8 | 3/8 | n=1 ● |
| 44 | Qwen3.6-35B-A3B MLX 8bit Uniform (lmstudio-community) | Card | 35B total / 3B active per token | 8bit uniform MLX | 88 | — | — | — | n/d | — | — | 25.3 | 88.8 | 2/8 | n=2 |
| =45 | Hermes-4 70B 4bit MLX | Card | 70B | 4bit MLX | 74 | — | — | — | n/d | — | — | 3 | 10.4 | 2/8 | n=2 |
| =45 | Devstral Small 2 24B 4bit MLX | Card | 24B | 4bit MLX | 73.5 | — | — | — | n/d | — | — | 10.2 | 35.8 | 2/8 | n=2.1 |
| 47 | Qwen3.8-27B MTPLX Optimized-Speed 4bit MLX (Youssofal) | Card | 27B dense | 4bit MLX | 68.1 | — | — | — | n/d | — | — | 6.3 | 22.2 | 2/8 | n=1 ● |
| =48 | DeepSeek-R1-Distill-Llama 70B 8bit MLX | Card | 70B | 8bit MLX | 62.5 | — | — | — | n/d | — | — | 1.6 | 5.7 | 2/8 | n=2 |
| =48 | DeepSeek-R1-Distill-Qwen-32B 8bit MLX | Card | 32B | 8bit MLX | 60.5 | — | — | — | n/d | — | — | 3.7 | 13.1 | 2/8 | n=2 |
| 50 | GPT-OSS Safeguard 20B MLX | Card | 20B | MLX | 53 | — | — | — | n/d | — | — | 44 | 154.5 | 2/8 | n=1 ● |
| 51 | Qwen3.8-27B MLX 8bit (lmstudio-community) | Card | 27B dense | 8bit MLX | — | 97.5 | — | — | n/d | — | — | 4.9 | 17.1 | 2/8 | n=1 |
SMALL
= under 8B total parameters (phone/edge class) — scored on the same suites, ranked in the same list; compare within class.
Axes are 0–100. — = suite not yet run for that model.
n/d = Thinking axis needs ≥15 observed tests to score (see methodology).
Default order: GAUNTLET progress, then Generalist score.
How much to trust a gap
Most of these scores are a single observation. 306 of
362 published model×suite cells (85%) are
n=1. The Evidence column gives the fewest repeats
behind each row.
Ranks marked =N are ties, not orderings. We measure run-to-run noise from byte-identical arms — the same weights file served by two runtimes, where any gap is measurement error by construction. Across 23 paired observations: mean Δ 0.1, mean |Δ| 0.6, sd 1.2, and the largest gap between two identical arms is 4 points. Anything inside 1 point shares a rank. That band is about one sd — it is a floor, not a confidence interval.
A dot (●) means a score was computed on a subset. When a generation dies mid-stream we record the test as did-not-finish rather than scoring it zero, because the failure is the serving layer's, not the model's. But those failures are not random — they fall on long, hard generations, and across this corpus the excluded tests score about 0.8 points lower than the ones that survived. So a partial score reads high, and the dot tells you which rows to discount. Hover it for the exact coverage.
One host, one operator, no external replication. Everything above is measured on this corpus and is published so you can check it — see the methodology.