← Leaderboard

Gemma-4 E4B 7.5B Q4_K_M GGUFSMALL

Keeper Rank #23 of 51 · 7/8 GAUNTLET progress
Generalist — 78 / 100 G Agentic — 89.5 / 100 A Understanding — 84.7 / 100 U Needle — 64.4 / 100 N Thinking — pending T Live — 72.3 / 100 L Engineering — 75 / 100 E Throughput — 23.6 / 100 T

Hollow markers & dashed spokes: axis not yet scored

G
78
A
89.5
U
84.7
N
64.4
T
n/d
L
72.3
E
75
T
23.6

Specification

Parameters
7.5B
Architecture
gemma4
Size on disk
6.33 GB
Quantization
Q4_K_M
Format
GGUF
Runtime
LM Studio
Class
Small (under 8B total parameters — phone/edge class)
Reasoning (CoT)
Yes — emits reasoning tokens
Internal ID
M1a
Mean speed
82.9 tok/s across suites
Stall census
0 stalls in 13 observed tests (0.0%)
Reasoning appetite
861 tokens mean · 2,100 max
Model card
lmstudio.ai

Suite results

Suite Score Avg / 20 Tests n tok/s Run
General capability (13-task real-workload suite) 1a 202.8 / 260 15.6 13/13 n=2.2 88.6 run 2026-09-02
Test Category Score tok/s
A1 Dev/Ops Scripting 11/20 102.7
A2 Dev/Ops Scripting 16/20 99.7
A3 Dev/Ops Scripting 10/20 97.4
A4 Dev/Ops Scripting 19/20 100
A5 Dev/Ops Scripting 12/20 101.3
A6 Dev/Ops Scripting 20/20 101.7
B1 Document Processing 20/20 83
B2 Document Processing 20/20 71.6
B3 Document Processing 8/20 76.2
C1 Content Production 18/20 89
C2 Content Production 17/20 76.1
C3 Content Production 20/20 79.8
C4 Content Production 12/20 83.3

Category mean Dev/Ops Scripting 14.7Document Processing 16Content Production 16.8

Agentic tool-calling & protocol adherence 1b 143.2 / 160 17.9 8/8 n=1 101.2 run 2026-09-16
Test Category Score tok/s
WA1 Tool Calling & Protocol Adherence 14/20 63.8
WA2 Tool Calling & Protocol Adherence 20/20 80.8
WA3 Tool Calling & Protocol Adherence 20/20 63.9
WA4 Tool Calling & Protocol Adherence 20/20 74
WA5 Tool Calling & Protocol Adherence 20/20 88.5
WB1 Agent Robustness & State Management 20/20 97.6
WB2 Agent Robustness & State Management 20/20 96.3
WB3 Agent Robustness & State Management 9/20 96

Category mean Tool Calling & Protocol Adherence 18.8Agent Robustness & State Management 16.3

Coding depth 1c 135 / 180 15 9/9 n=1 93.1 run 2026-09-18
Test Category Score tok/s
K1 Coding — Concurrency 18/20 92.8
K2 Coding — Algorithms 20/20 93.7
K3 Coding — Security Review 13/20 95.1
K4 Coding — Performance 15/20 92
K5 Coding — Refactoring Under Constraints 20/20 90.6
K6 Coding — Test Writing 11/20 92
K7 Coding — React Component (typed) 11/20 91.2
K8 Coding — TypeScript Type System 11/20 90.4
K9 Coding — React Debugging 16/20 85

Category mean Coding — Concurrency 18Coding — Algorithms 20Coding — Security Review 13Coding — Performance 15Coding — Refactoring Under Constraints 20Coding — Test Writing 11Coding — React Component (typed) 11Coding — TypeScript Type System 11Coding — React Debugging 16

Doc/OCR vision 1d 152 / 160 19 8/8 n=1 87 run 2026-09-20
Test Category Score tok/s
D1 Doc/OCR — Transcription 20/20 53.2
D2 Doc/OCR — Document QA 20/20 70.7
D3 Doc/OCR — Field Extraction 20/20 59.5
D4 Doc/OCR — Document QA 20/20 66.3
D5 Doc/OCR — Transcription 12/20 71.3
D6 Doc/OCR — Field Extraction 20/20 60.3
D7 Doc/OCR — Degraded Input 20/20 83.9
D8 Doc/OCR — Degraded Input 20/20 74.9

Category mean Doc/OCR — Transcription 16Doc/OCR — Document QA 20Doc/OCR — Field Extraction 20Doc/OCR — Degraded Input 20

Doc/OCR — real-degraded tier 1d2 51.2 / 80 12.8 4/4 n=1 74.5 run 2026-09-20
Test Category Score tok/s
D9 Doc/OCR — Real Degraded Transcription 8/20 54
D10 Doc/OCR — Handwriting Extraction 17/20 69.6
D11 Doc/OCR — Degraded Table Transcription 9/20 63.3
D12 Doc/OCR — Degraded Table QA 17/20 66

Category mean Doc/OCR — Real Degraded Transcription 8Doc/OCR — Handwriting Extraction 17Doc/OCR — Degraded Table Transcription 9Doc/OCR — Degraded Table QA 17

Content-production depth 1e 87 / 120 14.5 6/6 n=1 85.8 run 2026-09-16
Test Category Score tok/s
E1 SOW & Requirements Decomposition 20/20 80.8
E2 SOW & Requirements Decomposition 20/20 82.9
E3 Executive & Proposal Content 20/20 81.8
E4 Executive & Proposal Content 10/20 85.2
E5 Executive & Proposal Content 17/20 85.2
E6 Executive & Proposal Content 0/20 86.8

Category mean SOW & Requirements Decomposition 20Executive & Proposal Content 11.8

Long-context retrieval & synthesis 1g 100 / 120 20 5/6 ● n=1 69.7 run 2026-09-17
Test Category Score tok/s
L1 Long-Context — Needle Retrieval 20/20 52.5
L2 Long-Context — Grounded QA with Distractors 20/20 44.8
L3 Long-Context — Ops Log Reasoning 20/20 34.9
L4 Long-Context — Faithful Summarization 20/20 45.1
L5 Long-Context — Instruction Retention 20/20 39.7

Category mean Long-Context — Needle Retrieval 20Long-Context — Grounded QA with Distractors 20Long-Context — Ops Log Reasoning 20Long-Context — Faithful Summarization 20Long-Context — Instruction Retention 20

Long-context multi-needle (MRCR) 1g2 15.9 / 60 5.3 3/3 n=1 73.2 run 2026-08-31
Test Category Score tok/s
L7 Long-Context — Multi-Needle Retrieval (MRCR) 8/20 50.9
L8 Long-Context — Multi-Needle Retrieval (MRCR) 4/20 14.6
L9 Long-Context — Multi-Needle Retrieval (MRCR) 4/20 8.4
Live one-shot builds (runtime-verified) 1h 51.2 / 80 12.8 4/4 n=1 72.1 run 2026-09-04
Test Category Score tok/s
H1 Web-Estate One-Shot Builds 15/20 60.5
H2 Web-Estate One-Shot Builds 16/20 55.3
H3 Web-Estate One-Shot Builds 9/20 84.5
H4 Web-Estate One-Shot Builds 11/20 85.3
Brand-constrained static build (negative-constraint adherence) 1j 79 / 100 15.8 5/5 n=1 84 run 2026-09-16
Test Category Score tok/s
J1 Brand-Constrained Static Build 17/20 78
J3 Brand-Constrained Static Build 16/20 81.3
J4 Brand-Constrained Static Build 20/20 82.1
J5 Brand-Constrained Static Build 15/20 84.2
J6 Brand-Constrained Static Build 11/20 87

Each test is LLM-judged 0–20 against a fixed rubric; suite max = tests × 20. An ADJ chip marks a cell carrying a subjective adjustment from the maintainer; the judged score is the one shown, is unchanged, and remains what the axes, ranks and tie bands are computed from. Hover for the reason. Rows with a ▶ expand to the per-test breakdown. Where fewer tests ran than the suite total, unrun tests count as zero toward the axis score. See methodology.