Local-LLM proving ground

Proving local agents before they act.

GAUNTLET Bench runs local models through real workloads — agentic, live, verified — and publishes the evidence.

51models 11suites 2,484per-test results

All scores measured on Apple M5 Max, 128 GB unified memory, LM Studio

Current podium

1 Generalist — 99 / 100 G Agentic — 96 / 100 A Understanding — 93.3 / 100 U Needle — 91.2 / 100 N Thinking — 96.5 / 100 T Live — 83 / 100 L Engineering — 86.7 / 100 E Throughput — 5.7 / 100 T

Qwen3.8-27B Q6_K GGUF

27B dense · Q6_K GGUF

8/8 20 tok/s
2 Generalist — 95.5 / 100 G Agentic — 99.5 / 100 A Understanding — 81 / 100 U Needle — 77.8 / 100 N Thinking — 88 / 100 T Live — 68 / 100 L Engineering — 98.5 / 100 E Throughput — 19.8 / 100 T

Gemma-4 26B-A4B 8bit MLX

26B total / 4B active per token · 8bit MLX

8/8 69.6 tok/s

Leaderboard — 51 models · click a column to sort

#▲ Model▼ Card Params▼ Quant / Format▼ G▼ A▼ U▼ N▼ T▼ L▼ E▼ T▼ tok/s▼ Progress▼ Evidence▼
=1 Qwen3.8-27B Q6_K GGUF Card 27B dense Q6_K GGUF 999693.391.296.58386.75.7 20 8/8 n=1  ●
=1 Gemma-4 26B-A4B 8bit MLX Card 26B total / 4B active per token 8bit MLX 95.599.58177.8886898.519.8 69.6 8/8 n=1
=1 Qwen3.6-27B Fable-Fusion-711 Uncensored Heretic NEO-MAX MTPLX 8bit MLX Card 27B 8bit MLX 95.59990.873.591.55388.44.5 15.7 8/8 n=1  ●
=1 Qwen3.8-27B Q4_K_M GGUF Card 27B dense Q4_K_M GGUF 9510095.391.295.785.694.57.6 26.7 8/8 n=1  ●
=1 Qwen3.6-35B-A3B 4bit MLX Card 35B total / 3B active per token 4bit MLX 947893.791.297.176.796.530.7 108 8/8 n=1
=6 Qwen3.6-35B-A3B Unsloth Dynamic UD-Q8_K_XL MLX (Brooooooklyn) Card 35B total / 3B active per token UD-Q8_K_XL mixed-precision MLX 93.59679.791.210083.39823.7 83.2 8/8 n=1
=6 Qwen3.6-27B Dense 8bit MLX Card 27B dense 8bit MLX 90.510096.310097.27977.84.2 14.6 8/8 n=1  ●
=6 Qwen3.5-122B-A10B 4bit MLX Card 122B total / 10B active per token 4bit MLX 9094.510091.210084.498.514.2 50 8/8 n=1
=9 Gemma-4 E4B 4B 8bit MLX SMALL Card 4B 8bit MLX 84.594.582.383.81007985.522.3 78.3 8/8 n=1
=9 Bonsai 27B Ternary 2bit MLX Card 27B 2bit (ternary, PrismML extreme compression) MLX 819290.38093.175.69610.4 36.7 8/8 n=1
=9 Mistral Small 3.2 24B 8bit MLX Card 24B 8bit MLX 80.5917980.210062.880.55.2 18.2 8/8 n=1
=12 Qwen3.8-27B Q6_K GGUF lmstudio-community (Ollama, MBP - identical-file arm of M25b) OLLAMA Card 27B dense Q6_K GGUF 98.510094.791.2n/d95974.4 15.6 7/8 n=1
=12 Qwen3.8-27B GSQ-RCO IQ3_S GGUF ISTA-DASLab (Ollama, MBP — identical-file arm of M35) OLLAMA Card 27B dense IQ3_S GGUF 981009291.2n/d84.2996.2 21.7 7/8 n=1
=12 Qwen3.8-27B UD-Q6_K_XL GGUF (Ollama, identical-file control) OLLAMA Card 27B dense Q6_K_XL GGUF 95.51009591.2n/d95.3984.3 15.2 7/8 n=1
=12 Qwen3.8-27B nvfp4 MLX (Ollama, MBP) OLLAMA Card 27B dense nvfp4 MLX 951009389.8n/d48.69713.7 48.3 7/8 n=1
=12 Gemma-4 26B-A4B MLX (Ollama, MBP) OLLAMA Card 26B total / 4B active per token MLX 93.587.57377.5n/d49.773.534.4 120.8 7/8 n=1
=17 Ornith-1.5-35B-A3B Q4_K_M GGUF Card 35.95B total / ~3B active per token (MoE, 256 experts, 8 used) Q4_K_M GGUF 90.59988100n/d72.788.527.2 95.6 7/8 n=1  ●
=17 Qwen3.6-27B Fable-Fusion-711 Uncensored Heretic NEO-MAX Q8_0 GGUF Card 27B Q8_0 GGUF 89.510084.291.297.7—973.7 13.1 7/8 n=1  ●
=17 Qwen3-Coder-Next 80B 4bit MLX Card 80B 4bit MLX 88.594.5—64.410087.889.523.3 81.8 7/8 n=1  ●
=17 Ornith-1.0-35B 4bit MLX Card 35B total (MoE; active-per-token not confirmed in model card) 4bit MLX 86.39894.790.2n/d90.49935 122.9 7/8 n=1  ●
=21 Qwen3.8-27B MLX 4bit (lmstudio-community) Card 27B dense 4bit MLX 82.910091.391.2n/d82.8996.8 23.9 7/8 n=1  ●
=21 Gemma-4 31B IT 4bit MLX Card 31.27B 4bit MLX 82.110085.491.2n/d82.2985.1 18 7/8 n=1  ●
=21 Gemma-4 E4B 7.5B Q4_K_M GGUF SMALL Card 7.5B Q4_K_M GGUF 7889.584.764.4n/d72.37523.6 82.9 7/8 n=1  ●
=24 Gemma-3 12B IT QAT 4bit MLX Card 12.19B 4bit (QAT) MLX 67.5925467.8n/d57.869.516.1 56.4 7/8 n=1
=24 GPT-OSS 120B Card 120B mxfp4 MLX 65.193—86.71008590.518.2 64 7/8 n=1  ●
26 Gemma 3 4B QAT 4bit MLX SMALL Card 4B QAT 4bit MLX 57.58649.346.8n/d37.95146.3 162.6 7/8 n=1
27 Qwen3-VL 30B-A3B Instruct 4bit MLX Card 30B total / 3B active per token 4bit MLX —8583.382.810066.78431 109 7/8 n=1
=28 Qwen3.8-27B Uncensored HauhauCS Aggressive Q5_K_P GGUF (Ollama, MBP — identical-file arm of M29) OLLAMA Card 27B dense Q5_K_P GGUF 97.597—91.2n/d94.6994.4 15.6 6/8 n=1
=28 Qwen3.8-27B Uncensored HauhauCS Aggressive Q6_K_P GGUF (Ollama, MBP — identical-file arm of M27) OLLAMA Card 27B dense Q6_K_P GGUF 97.598—91.2n/d96.8993.9 13.8 6/8 n=1
=28 Qwen3.8-27B MLX 4bit lmstudio-community (Ollama, MBP — identical-weights arm of M25h) OLLAMA Card 27B dense 4bit MLX 93.5100—91.2n/d85977.2 25.1 6/8 n=1
=28 Qwen3.8-27B MLX 8bit lmstudio-community (Ollama, MBP — quant-ladder rung) OLLAMA Card 27B dense 8bit MLX 93100—91.2n/d96.6994.5 15.9 6/8 n=1
32 Qwen3-4B-2507 (non-thinking) 4bit MLX SMALL Card 4B 4bit MLX 6893—80n/d46.65638 133.6 6/8 n=1
33 GPT-OSS 20B (OpenAI mxfp4) Card 20B mxfp4 MLX 62.365—65.7n/d62.49033.4 117.3 6/8 n=1  ●
34 DeepSeek-R1-0528-Qwen3-8B 4bit MLX Card 8.19B 4bit MLX 4293—56.2n/d33.94419.8 69.6 6/8 n=1  ●
35 Phi-3.5-mini-instruct 4bit MLX SMALL Card 3.8B 4bit MLX 3626—37.2n/d10.62526 91.2 6/8 n=1  ●
36 Qwen3.8-27B MLX 6bit (lmstudio-community) Card 27B dense 6bit MLX 90977473.5n/d——4.4 15.4 5/8 n=1  ●
37 Nemotron 3 Nano 30B-A3B 4bit MLX Card 30B total / ~3B active per token 4bit MLX 78.592.5—7596.7——39 137.2 5/8 n=1
38 Gemma 3 1B QAT 4bit MLX SMALL Card 1B QAT 4bit MLX 28.533——n/d7.818.5100 351.3 5/8 n=1
39 Qwen3.8-27B Q8_0 GGUF Card 27B dense Q8_0 GGUF 9495.5——100——7.8 27.6 4/8 n=1
40 Laguna S 2.1 (Q4_K_M GGUF) Card 118B total / 8B active (MoE, 10-of-256 experts + 1 shared) Q4_K_M GGUF 8799——100——19.8 69.6 4/8 n=1
=41 Devstral Small 2 24B 6bit MLX Card 24B 6bit MLX 7897——100——6.3 22 4/8 n=1
=41 Muse-Glimmer 30B (GGUF, kquant 17GB) Card 30B (unconfirmed -- filename-derived, see note) kquant (unconfirmed exact scheme -- filename says "kquant", not a standard llama.cpp quant label) GGUF 7476——100——5.9 20.7 4/8 n=1
43 Qwen3.8-27B MLX 5bit (lmstudio-community) Card 27B dense 5bit MLX 74.6100——n/d——4.5 15.8 3/8 n=1  ●
44 Qwen3.6-35B-A3B MLX 8bit Uniform (lmstudio-community) Card 35B total / 3B active per token 8bit uniform MLX 88———n/d——25.3 88.8 2/8 n=2
=45 Hermes-4 70B 4bit MLX Card 70B 4bit MLX 74———n/d——3 10.4 2/8 n=2
=45 Devstral Small 2 24B 4bit MLX Card 24B 4bit MLX 73.5———n/d——10.2 35.8 2/8 n=2.1
47 Qwen3.8-27B MTPLX Optimized-Speed 4bit MLX (Youssofal) Card 27B dense 4bit MLX 68.1———n/d——6.3 22.2 2/8 n=1  ●
=48 DeepSeek-R1-Distill-Llama 70B 8bit MLX Card 70B 8bit MLX 62.5———n/d——1.6 5.7 2/8 n=2
=48 DeepSeek-R1-Distill-Qwen-32B 8bit MLX Card 32B 8bit MLX 60.5———n/d——3.7 13.1 2/8 n=2
50 GPT-OSS Safeguard 20B MLX Card 20B MLX 53———n/d——44 154.5 2/8 n=1  ●
51 Qwen3.8-27B MLX 8bit (lmstudio-community) Card 27B dense 8bit MLX —97.5——n/d——4.9 17.1 2/8 n=1

SMALL = under 8B total parameters (phone/edge class) — scored on the same suites, ranked in the same list; compare within class.
Axes are 0–100. — = suite not yet run for that model. n/d = Thinking axis needs ≥15 observed tests to score (see methodology). Default order: GAUNTLET progress, then Generalist score.

How much to trust a gap

Most of these scores are a single observation. 306 of 362 published model×suite cells (85%) are n=1. The Evidence column gives the fewest repeats behind each row.

Ranks marked =N are ties, not orderings. We measure run-to-run noise from byte-identical arms — the same weights file served by two runtimes, where any gap is measurement error by construction. Across 23 paired observations: mean Δ 0.1, mean |Δ| 0.6, sd 1.2, and the largest gap between two identical arms is 4 points. Anything inside 1 point shares a rank. That band is about one sd — it is a floor, not a confidence interval.

A dot (●) means a score was computed on a subset. When a generation dies mid-stream we record the test as did-not-finish rather than scoring it zero, because the failure is the serving layer's, not the model's. But those failures are not random — they fall on long, hard generations, and across this corpus the excluded tests score about 0.8 points lower than the ones that survived. So a partial score reads high, and the dot tells you which rows to discount. Hover it for the exact coverage.

One host, one operator, no external replication. Everything above is measured on this corpus and is published so you can check it — see the methodology.