← Leaderboard

Ornith-1.5-35B-A3B Q4_K_M GGUF

Tested Rank #8 of 40 · 6/8 GAUNTLET progress
Generalist — 90.5 / 100 G Agentic — 99 / 100 A Understanding — 88 / 100 U Needle — pending N Thinking — pending T Live — 73.9 / 100 L Engineering — 100 / 100 E Throughput — 34.5 / 100 T

Hollow markers & dashed spokes: axis not yet scored

G
90.5
A
99
U
88
N
T
n/d
L
73.9
E
100
T
34.5

Specification

Parameters
35.95B total / ~3B active per token (MoE, 256 experts, 8 used)
Architecture
qwen35moe
Size on disk
20 GB
Quantization
Q4_K_M
Format
GGUF
Class
Standard
Reasoning (CoT)
Yes — emits reasoning tokens
Internal ID
M34
Mean speed
82 tok/s across suites
Model card
huggingface.co

Suite results

Suite Score Avg / 20 Tests tok/s Run
General capability (13-task real-workload suite) 1a 235.3 / 260 18.1 13/13 84.3 run 2026-09-06
Test Category Score tok/s
A1 Dev/Ops Scripting 20/20 92.5
A2 Dev/Ops Scripting 18/20 99.5
A3 Dev/Ops Scripting 20/20 101
A4 Dev/Ops Scripting 19/20 82.2
A5 Dev/Ops Scripting 20/20 76
A6 Dev/Ops Scripting 20/20 81
B1 Document Processing 20/20 100.2
B2 Document Processing 20/20 82.7
B3 Document Processing 0/20
C1 Content Production 20/20 75.1
C2 Content Production 18/20 73.5
C3 Content Production 20/20 72.4
C4 Content Production 20/20 75.4

Category mean Dev/Ops Scripting 19.5Document Processing 13.3Content Production 19.5

Agentic tool-calling & protocol adherence 1b 158.4 / 160 19.8 8/8 94.5 run 2026-09-06
Test Category Score tok/s
WA1 Tool Calling & Protocol Adherence 20/20 96.2
WA2 Tool Calling & Protocol Adherence 20/20 88
WA3 Tool Calling & Protocol Adherence 20/20 93.7
WA4 Tool Calling & Protocol Adherence 20/20 73.8
WA5 Tool Calling & Protocol Adherence 20/20 103.6
WB1 Agent Robustness & State Management 20/20 95.5
WB2 Agent Robustness & State Management 20/20 109.3
WB3 Agent Robustness & State Management 18/20 96.2

Category mean Tool Calling & Protocol Adherence 20Agent Robustness & State Management 19.3

Coding depth 1c 120 / 120 20 6/6 97.2 run 2026-09-06
Test Category Score tok/s
K1 Coding — Concurrency 20/20 101.7
K2 Coding — Algorithms 20/20 123.3
K3 Coding — Security Review 20/20 101.2
K4 Coding — Performance 20/20 87.6
K5 Coding — Refactoring Under Constraints 20/20 79.6
K6 Coding — Test Writing 20/20 89.6

Category mean Coding — Concurrency 20Coding — Algorithms 20Coding — Security Review 20Coding — Performance 20Coding — Refactoring Under Constraints 20Coding — Test Writing 20

Doc/OCR vision 1d 160 / 160 20 8/8 51.1 run 2026-09-06
Test Category Score tok/s
D1 Doc/OCR — Transcription 20/20 36.7
D2 Doc/OCR — Document QA 20/20 32.3
D3 Doc/OCR — Field Extraction 20/20 54.8
D4 Doc/OCR — Document QA 20/20 44.9
D5 Doc/OCR — Transcription 20/20 61.8
D6 Doc/OCR — Field Extraction 20/20 35
D7 Doc/OCR — Degraded Input 20/20 72
D8 Doc/OCR — Degraded Input 20/20 71

Category mean Doc/OCR — Transcription 20Doc/OCR — Document QA 20Doc/OCR — Field Extraction 20Doc/OCR — Degraded Input 20

Doc/OCR — real-degraded tier 1d2 51.2 / 80 12.8 4/4 75 run 2026-09-06
Test Category Score tok/s
D9 Doc/OCR — Real Degraded Transcription 7/20 79.1
D10 Doc/OCR — Handwriting Extraction 17/20 52.8
D11 Doc/OCR — Degraded Table Transcription 14/20 92.1
D12 Doc/OCR — Degraded Table QA 13/20 76.2

Category mean Doc/OCR — Real Degraded Transcription 7Doc/OCR — Handwriting Extraction 17Doc/OCR — Degraded Table Transcription 14Doc/OCR — Degraded Table QA 13

Live one-shot builds (runtime-verified) 1h 59.1 / 80 19.7 3/4 90.1 run 2026-09-06
Test Category Score tok/s
H1 Web-Estate One-Shot Builds 19/20 103.6
H2 Web-Estate One-Shot Builds 20/20 84.9
H4 Web-Estate One-Shot Builds 20/20 81.7

Each test is LLM-judged 0–20 against a fixed rubric; suite max = tests × 20. Rows with a expand to the per-test breakdown. Where fewer tests ran than the suite total, unrun tests count as zero toward the axis score. See methodology.