← Leaderboard

Laguna S 2.1 (Q4_K_M GGUF)

Retired Rank #40 of 51 · 4/8 GAUNTLET progress
Generalist — 87 / 100 G Agentic — 99 / 100 A Understanding — pending U Needle — pending N Thinking — 100 / 100 T Live — pending L Engineering — pending E Throughput — 19.8 / 100 T

Hollow markers & dashed spokes: axis not yet scored

G
87
A
99
U
—
N
—
T
100
L
—
E
—
T
19.8

Specification

Parameters
118B total / 8B active (MoE, 10-of-256 experts + 1 shared)
Architecture
laguna
Size on disk
71.2 GB
Quantization
Q4_K_M
Format
GGUF
Runtime
LM Studio
Class
Standard
Reasoning (CoT)
No
Internal ID
M20
Mean speed
69.6 tok/s across suites
Stall census
0 stalls in 29 observed tests (0.0%)
Model card
huggingface.co

Suite results

Suite Score Avg / 20 Tests n tok/s Run
General capability (13-task real-workload suite) 1a 226.2 / 260 17.4 13/13 n=1.2 70 run 2026-08-04
Test Category Score tok/s
A1 Dev/Ops Scripting 17/20 64.9
A2 Dev/Ops Scripting 18/20 54.5
A3 Dev/Ops Scripting 7/20 64.1
A4 Dev/Ops Scripting 19/20 57.6
A5 Dev/Ops Scripting 18/20 60.3
A6 Dev/Ops Scripting 20/20 59.7
B1 Document Processing 20/20 56.1
B2 Document Processing 20/20 60.7
B3 Document Processing 16/20 62.4
C1 Content Production 20/20 61.4
C2 Content Production 18/20 60.7
C3 Content Production 15/20 63.7
C4 Content Production 20/20 47.4

Category mean Dev/Ops Scripting 16.5Document Processing 18.7Content Production 18.3

Agentic tool-calling & protocol adherence 1b 158.4 / 160 19.8 8/8 n=1 69.2 run 2026-08-04
Test Category Score tok/s
WA1 Tool Calling & Protocol Adherence 20/20 20.5
WA2 Tool Calling & Protocol Adherence 20/20 16.7
WA3 Tool Calling & Protocol Adherence 20/20 17.4
WA4 Tool Calling & Protocol Adherence 20/20 21.6
WA5 Tool Calling & Protocol Adherence 20/20 50.3
WB1 Agent Robustness & State Management 19/20 58.5
WB2 Agent Robustness & State Management 20/20 55.6
WB3 Agent Robustness & State Management 19/20 58.5

Category mean Tool Calling & Protocol Adherence 20Agent Robustness & State Management 19.3

Each test is LLM-judged 0–20 against a fixed rubric; suite max = tests × 20. An ADJ chip marks a cell carrying a subjective adjustment from the maintainer; the judged score is the one shown, is unchanged, and remains what the axes, ranks and tie bands are computed from. Hover for the reason. Rows with a ▶ expand to the per-test breakdown. Where fewer tests ran than the suite total, unrun tests count as zero toward the axis score. See methodology.