← Leaderboard

Muse-Glimmer 30B (GGUF, kquant 17GB)

tested-offloaded Rank #42 of 51 · 4/8 GAUNTLET progress
Generalist — 74 / 100 G Agentic — 76 / 100 A Understanding — pending U Needle — pending N Thinking — 100 / 100 T Live — pending L Engineering — pending E Throughput — 5.9 / 100 T

Hollow markers & dashed spokes: axis not yet scored

G
74
A
76
U
—
N
—
T
100
L
—
E
—
T
5.9

Specification

Parameters
30B (unconfirmed -- filename-derived, see note)
Architecture
muse-glimmer
Size on disk
17 GB
Quantization
kquant (unconfirmed exact scheme -- filename says "kquant", not a standard llama.cpp quant label)
Format
GGUF
Runtime
LM Studio
Class
Standard
Reasoning (CoT)
No
Internal ID
M24
Mean speed
20.7 tok/s across suites
Stall census
0 stalls in 22 observed tests (0.0%)
Model card
huggingface.co

Suite results

Suite Score Avg / 20 Tests n tok/s Run
General capability (13-task real-workload suite) 1a 192.4 / 260 14.8 13/13 n=1.1 15.4 run 2026-08-13
Test Category Score tok/s
A1 Dev/Ops Scripting 13/20 20.6
A2 Dev/Ops Scripting 15/20 12.1
A3 Dev/Ops Scripting 15/20 14.3
A4 Dev/Ops Scripting 12/20 16.1
A5 Dev/Ops Scripting 17/20 16.1
A6 Dev/Ops Scripting 12/20 16.3
B1 Document Processing 17/20 13.4
B2 Document Processing 15/20 13.6
B3 Document Processing 20/20 14.8
C1 Content Production 15/20 16.2
C2 Content Production 16/20 16.7
C3 Content Production 14/20 14.7
C4 Content Production 11/20 15.5

Category mean Dev/Ops Scripting 14Document Processing 17.3Content Production 14

Agentic tool-calling & protocol adherence 1b 121.6 / 160 15.2 8/8 n=1 26 run 2026-08-13
Test Category Score tok/s
WA1 Tool Calling & Protocol Adherence 11/20 26
WA2 Tool Calling & Protocol Adherence 14/20 21.2
WA3 Tool Calling & Protocol Adherence 14/20 24.2
WA4 Tool Calling & Protocol Adherence 17/20 14.4
WA5 Tool Calling & Protocol Adherence 17/20 24.5
WB1 Agent Robustness & State Management 16/20 22.4
WB2 Agent Robustness & State Management 18/20 18.1
WB3 Agent Robustness & State Management 15/20 16.9

Category mean Tool Calling & Protocol Adherence 14.6Agent Robustness & State Management 16.3

Each test is LLM-judged 0–20 against a fixed rubric; suite max = tests × 20. An ADJ chip marks a cell carrying a subjective adjustment from the maintainer; the judged score is the one shown, is unchanged, and remains what the axes, ranks and tie bands are computed from. Hover for the reason. Rows with a ▶ expand to the per-test breakdown. Where fewer tests ran than the suite total, unrun tests count as zero toward the axis score. See methodology.