Suite results
Suite Score Avg / 20 Tests n tok/s Run
▶ General capability (13-task real-workload suite) 1a 93.6 / 260 7.2 13/13 n=1 114.8 run 2026-08-04
| Test | Category | Score | | tok/s |
| A1 | Dev/Ops Scripting | 2/20 | | 76.3 |
| A2 | Dev/Ops Scripting | 7/20 | | 117.1 |
| A3 | Dev/Ops Scripting | 1/20 | | 146.5 |
| A4 | Dev/Ops Scripting | 6/20 | | 158.2 |
| A5 | Dev/Ops Scripting | 5/20 | | 71.8 |
| A6 | Dev/Ops Scripting | 11/20 | | 133.3 |
| B1 | Document Processing | 11/20 | | 131.2 |
| B2 | Document Processing | 10/20 | | 137.7 |
| B3 | Document Processing | 7/20 | | 56.8 |
| C1 | Content Production | 13/20 | | 157.4 |
| C2 | Content Production | 8/20 | | 60.8 |
| C3 | Content Production | 8/20 | | 127.4 |
| C4 | Content Production | 5/20 | | 57.7 |
Category mean Dev/Ops Scripting 5.3Document Processing 9.3Content Production 8.5
▶ Agentic tool-calling & protocol adherence 1b 41.6 / 160 5.2 8/8 n=1 95.7 run 2026-09-22
| Test | Category | Score | | tok/s |
| WA1 | Tool Calling & Protocol Adherence | 9/20 | | 137.4 |
| WA2 | Tool Calling & Protocol Adherence | 3/20 | | 70 |
| WA3 | Tool Calling & Protocol Adherence | 2/20 | | 70.4 |
| WA4 | Tool Calling & Protocol Adherence | 0/20 | | 71.5 |
| WA5 | Tool Calling & Protocol Adherence | 0/20 | | 126.7 |
| WB1 | Agent Robustness & State Management | 8/20 | | 73.8 |
| WB2 | Agent Robustness & State Management | 13/20 | | 70.9 |
| WB3 | Agent Robustness & State Management | 7/20 | | 129.4 |
Category mean Tool Calling & Protocol Adherence 2.8Agent Robustness & State Management 9.3
▶ Coding depth 1c 45 / 180 5 9/9 n=1 129.1 run 2026-09-29
| Test | Category | Score | | tok/s |
| K1 | Coding — Concurrency | 3/20 | | 122.4 |
| K2 | Coding — Algorithms | 1/20 | | 159.7 |
| K3 | Coding — Security Review | 6/20 | | 155 |
| K4 | Coding — Performance | 12/20 | | 69.3 |
| K5 | Coding — Refactoring Under Constraints | 9/20 | | 163.5 |
| K6 | Coding — Test Writing | 5/20 | | 59.9 |
| K7 | Coding — React Component (typed) | 2/20 | | 170.7 |
| K8 | Coding — TypeScript Type System | 3/20 | | 60.6 |
| K9 | Coding — React Debugging | 4/20 | | 104 |
Category mean Coding — Concurrency 3Coding — Algorithms 1Coding — Security Review 6Coding — Performance 12Coding — Refactoring Under Constraints 9Coding — Test Writing 5Coding — React Component (typed) 2Coding — TypeScript Type System 3Coding — React Debugging 4
▶ Content-production depth 1e 27 / 120 4.5 6/6 n=1 125.7 run 2026-09-24
| Test | Category | Score | | tok/s |
| E1 | SOW & Requirements Decomposition | 5/20 | | 143.6 |
| E2 | SOW & Requirements Decomposition | 2/20 | | 66.6 |
| E3 | Executive & Proposal Content | 7/20 | | 123.7 |
| E4 | Executive & Proposal Content | 5/20 | | 162.9 |
| E5 | Executive & Proposal Content | 8/20 | | 165.5 |
| E6 | Executive & Proposal Content | 0/20 | | 61.9 |
Category mean SOW & Requirements Decomposition 3.5Executive & Proposal Content 5
▶ Long-context retrieval & synthesis 1g 67 / 120 13.4 5/6 ● n=1 29.5 run 2026-09-23
| Test | Category | Score | | tok/s |
| L1 | Long-Context — Needle Retrieval | 3/20 | | 30.4 |
| L2 | Long-Context — Grounded QA with Distractors | 20/20 | | 6.2 |
| L3 | Long-Context — Ops Log Reasoning | 8/20 | | 9.5 |
| L5 | Long-Context — Instruction Retention | 18/20 | | 4.5 |
| L6 | Long-Context — Full-Haystack QA | 18/20 | | 1.4 |
Category mean Long-Context — Needle Retrieval 3Long-Context — Grounded QA with Distractors 20Long-Context — Ops Log Reasoning 8Long-Context — Instruction Retention 18Long-Context — Full-Haystack QA 18
▶ Long-context multi-needle (MRCR) 1g2 0 / 60 0 3/3 n=1 23.7 run 2026-09-27
| Test | Category | Score | | tok/s |
| L7 | Long-Context — Multi-Needle Retrieval (MRCR) | 0/20 | | 23.1 |
▶ Live one-shot builds (runtime-verified) 1h 2 / 80 0.5 4/4 n=1 102.3 run 2026-10-01
| Test | Category | Score | | tok/s |
| H1 | Web-Estate One-Shot Builds | 0/20 | | 53.7 |
| H3 | Web-Estate One-Shot Builds | 1/20 | | 51 |
▶ Brand-constrained static build (negative-constraint adherence) 1j 17 / 100 3.4 5/5 n=1 108.5 run 2026-09-22
| Test | Category | Score | | tok/s |
| J1 | Brand-Constrained Static Build | 7/20 | | 171.6 |
| J3 | Brand-Constrained Static Build | 2/20 | | 73.1 |
| J4 | Brand-Constrained Static Build | 2/20 | | 130.8 |
| J5 | Brand-Constrained Static Build | 6/20 | | 73 |
| J6 | Brand-Constrained Static Build | 0/20 | | 69.7 |
Each test is LLM-judged 0–20 against a fixed rubric; suite max = tests × 20.
An ADJ chip marks a cell carrying a subjective
adjustment from the maintainer; the judged score is the one shown, is unchanged,
and remains what the axes, ranks and tie bands are computed from. Hover for the reason.
Rows with a ▶ expand to the per-test breakdown.
Where fewer tests ran than the suite total, unrun tests count as zero toward the axis score.
See methodology.