Suite results
Suite Score Avg / 20 Tests n tok/s Run
▶ General capability (13-task real-workload suite) 1a 149.5 / 260 11.5 13/13 n=1 171.8 run 2026-08-04
| Test | Category | Score | | tok/s |
| A1 | Dev/Ops Scripting | 8/20 | | 167.1 |
| A2 | Dev/Ops Scripting | 12/20 | | 166.2 |
| A3 | Dev/Ops Scripting | 5/20 | | 161.7 |
| A4 | Dev/Ops Scripting | 5/20 | | 145.4 |
| A5 | Dev/Ops Scripting | 8/20 | | 166.6 |
| A6 | Dev/Ops Scripting | 12/20 | | 165.8 |
| B1 | Document Processing | 18/20 | | 158.3 |
| B2 | Document Processing | 20/20 | | 163.6 |
| B3 | Document Processing | 8/20 | | 167.4 |
| C1 | Content Production | 18/20 | | 167 |
| C2 | Content Production | 10/20 | | 164.8 |
| C3 | Content Production | 15/20 | | 167.1 |
| C4 | Content Production | 11/20 | | 162.5 |
Category mean Dev/Ops Scripting 8.3Document Processing 15.3Content Production 13.5
▶ Agentic tool-calling & protocol adherence 1b 137.6 / 160 17.2 8/8 n=1 171.3 run 2026-09-22
| Test | Category | Score | | tok/s |
| WA1 | Tool Calling & Protocol Adherence | 20/20 | | 102.8 |
| WA2 | Tool Calling & Protocol Adherence | 20/20 | | 111.1 |
| WA3 | Tool Calling & Protocol Adherence | 13/20 | | 115.7 |
| WA4 | Tool Calling & Protocol Adherence | 20/20 | | 96.2 |
| WA5 | Tool Calling & Protocol Adherence | 19/20 | | 139.6 |
| WB1 | Agent Robustness & State Management | 18/20 | | 166.1 |
| WB2 | Agent Robustness & State Management | 20/20 | | 167.1 |
| WB3 | Agent Robustness & State Management | 8/20 | | 130.5 |
Category mean Tool Calling & Protocol Adherence 18.4Agent Robustness & State Management 15.3
▶ Coding depth 1c 91.8 / 180 10.2 9/9 n=1 169.5 run 2026-09-29
| Test | Category | Score | | tok/s |
| K1 | Coding — Concurrency | 16/20 | | 145.5 |
| K2 | Coding — Algorithms | 8/20 | | 166.3 |
| K3 | Coding — Security Review | 8/20 | | 168 |
| K4 | Coding — Performance | 18/20 | | 167.2 |
| K5 | Coding — Refactoring Under Constraints | 10/20 | | 167.9 |
| K6 | Coding — Test Writing | 12/20 | | 163.4 |
| K7 | Coding — React Component (typed) | 8/20 | | 165.4 |
| K8 | Coding — TypeScript Type System | 6/20 | | 161.9 |
| K9 | Coding — React Debugging | 6/20 | | 167 |
Category mean Coding — Concurrency 16Coding — Algorithms 8Coding — Security Review 8Coding — Performance 18Coding — Refactoring Under Constraints 10Coding — Test Writing 12Coding — React Component (typed) 8Coding — TypeScript Type System 6Coding — React Debugging 6
▶ Doc/OCR vision 1d 87.2 / 160 10.9 8/8 n=1 169.6 run 2026-09-22
| Test | Category | Score | | tok/s |
| D1 | Doc/OCR — Transcription | 6/20 | | 119.4 |
| D2 | Doc/OCR — Document QA | 11/20 | | 87.8 |
| D3 | Doc/OCR — Field Extraction | 3/20 | | 133.7 |
| D4 | Doc/OCR — Document QA | 13/20 | | 118 |
| D5 | Doc/OCR — Transcription | 6/20 | | 124.4 |
| D6 | Doc/OCR — Field Extraction | 17/20 | | 126 |
| D7 | Doc/OCR — Degraded Input | 11/20 | | 137.8 |
| D8 | Doc/OCR — Degraded Input | 20/20 | | 133.1 |
Category mean Doc/OCR — Transcription 6Doc/OCR — Document QA 12Doc/OCR — Field Extraction 10Doc/OCR — Degraded Input 15.5
▶ Doc/OCR — real-degraded tier 1d2 31.2 / 80 7.8 4/4 n=1 176.5 run 2026-09-20
| Test | Category | Score | | tok/s |
| D9 | Doc/OCR — Real Degraded Transcription | 3/20 | | 85.5 |
| D10 | Doc/OCR — Handwriting Extraction | 16/20 | | 104.3 |
| D11 | Doc/OCR — Degraded Table Transcription | 3/20 | | 118 |
| D12 | Doc/OCR — Degraded Table QA | 9/20 | | 74.6 |
Category mean Doc/OCR — Real Degraded Transcription 3Doc/OCR — Handwriting Extraction 16Doc/OCR — Degraded Table Transcription 3Doc/OCR — Degraded Table QA 9
▶ Content-production depth 1e 73.2 / 120 12.2 6/6 n=1 150.9 run 2026-09-24
| Test | Category | Score | | tok/s |
| E1 | SOW & Requirements Decomposition | 15/20 | | 129.3 |
| E2 | SOW & Requirements Decomposition | 18/20 | | 148.4 |
| E3 | Executive & Proposal Content | 12/20 | | 137.9 |
| E4 | Executive & Proposal Content | 6/20 | | 142.9 |
| E5 | Executive & Proposal Content | 18/20 | | 144.7 |
| E6 | Executive & Proposal Content | 4/20 | | 159.1 |
Category mean SOW & Requirements Decomposition 16.5Executive & Proposal Content 10
▶ Long-context retrieval & synthesis 1g 76.2 / 120 12.7 6/6 n=1 152.3 run 2026-09-22
| Test | Category | Score | | tok/s |
| L1 | Long-Context — Needle Retrieval | 20/20 | | 19.7 |
| L2 | Long-Context — Grounded QA with Distractors | 7/20 | | 15.4 |
| L3 | Long-Context — Ops Log Reasoning | 6/20 | | 45.7 |
| L4 | Long-Context — Faithful Summarization | 6/20 | | 67.2 |
| L5 | Long-Context — Instruction Retention | 17/20 | | 16.9 |
| L6 | Long-Context — Full-Haystack QA | 20/20 | | 5.4 |
Category mean Long-Context — Needle Retrieval 20Long-Context — Grounded QA with Distractors 7Long-Context — Ops Log Reasoning 6Long-Context — Faithful Summarization 6Long-Context — Instruction Retention 17Long-Context — Full-Haystack QA 20
▶ Long-context multi-needle (MRCR) 1g2 8.1 / 60 2.7 3/3 n=1 137.3 run 2026-09-18
| Test | Category | Score | | tok/s |
| L7 | Long-Context — Multi-Needle Retrieval (MRCR) | 0/20 | | 58.2 |
| L8 | Long-Context — Multi-Needle Retrieval (MRCR) | 4/20 | | 24.4 |
| L9 | Long-Context — Multi-Needle Retrieval (MRCR) | 4/20 | | 107.3 |
▶ Live one-shot builds (runtime-verified) 1h 19.2 / 80 4.8 4/4 n=1 169.3 run 2026-09-18
| Test | Category | Score | | tok/s |
| H1 | Web-Estate One-Shot Builds | 7/20 | | 166.6 |
| H2 | Web-Estate One-Shot Builds | 6/20 | | 167.2 |
| H3 | Web-Estate One-Shot Builds | 6/20 | | 167.1 |
| H4 | Web-Estate One-Shot Builds | 0/20 | | 165.4 |
▶ Brand-constrained static build (negative-constraint adherence) 1j 49 / 100 9.8 5/5 n=1 157.4 run 2026-09-22
| Test | Category | Score | | tok/s |
| J1 | Brand-Constrained Static Build | 10/20 | | 140.3 |
| J3 | Brand-Constrained Static Build | 12/20 | | 146.7 |
| J4 | Brand-Constrained Static Build | 12/20 | | 147 |
| J5 | Brand-Constrained Static Build | 7/20 | | 153.7 |
| J6 | Brand-Constrained Static Build | 8/20 | | 166.1 |
Each test is LLM-judged 0–20 against a fixed rubric; suite max = tests × 20.
An ADJ chip marks a cell carrying a subjective
adjustment from the maintainer; the judged score is the one shown, is unchanged,
and remains what the axes, ranks and tie bands are computed from. Hover for the reason.
Rows with a ▶ expand to the per-test breakdown.
Where fewer tests ran than the suite total, unrun tests count as zero toward the axis score.
See methodology.