Suite results
Suite Score Avg / 20 Tests n tok/s Run
▶ General capability (13-task real-workload suite) 1a 230.1 / 260 17.7 13/13 n=5 86.5 run 2026-09-02
| Test | Category | Score | | tok/s |
| A1 | Dev/Ops Scripting | 8.6/20 | | 76.3 |
| A2 | Dev/Ops Scripting | 16.2/20 | | 76.2 |
| A3 | Dev/Ops Scripting | 18.4/20 | | 74.6 |
| A4 | Dev/Ops Scripting | 19/20 | | 70.7 |
| A5 | Dev/Ops Scripting | 20/20 | | 75.6 |
| A6 | Dev/Ops Scripting | 20/20 | | 75.4 |
| B1 | Document Processing | 20/20 | | 70.8 |
| B2 | Document Processing | 19.2/20 | | 71.8 |
| B3 | Document Processing | 18.8/20 | | 76.5 |
| C1 | Content Production | 18.4/20 | | 77.4 |
| C2 | Content Production | 17.6/20 | | 73.8 |
| C3 | Content Production | 17/20 | | 77.5 |
| C4 | Content Production | 16.6/20 | | 74.1 |
Category mean Dev/Ops Scripting 17Document Processing 19.3Content Production 17.4
▶ Agentic tool-calling & protocol adherence 1b 151.2 / 160 18.9 8/8 n=1 89.6 run 2026-07-28
| Test | Category | Score | | tok/s |
| WA1 | Tool Calling & Protocol Adherence | 20/20 | | 4.6 |
| WA2 | Tool Calling & Protocol Adherence | 20/20 | | 31.4 |
| WA3 | Tool Calling & Protocol Adherence | 20/20 | | 35.8 |
| WA4 | Tool Calling & Protocol Adherence | 20/20 | | 26 |
| WA5 | Tool Calling & Protocol Adherence | 20/20 | | 61.6 |
| WB1 | Agent Robustness & State Management | 20/20 | | 81.7 |
| WB2 | Agent Robustness & State Management | 20/20 | | 79 |
| WB3 | Agent Robustness & State Management | 11/20 | | 49.1 |
Category mean Tool Calling & Protocol Adherence 20Agent Robustness & State Management 17
▶ Coding depth 1c 161.1 / 180 17.9 9/9 n=1 78 run 2026-09-15
| Test | Category | Score | | tok/s |
| K1 | Coding — Concurrency | 19/20 | | 62.5 |
| K2 | Coding — Algorithms | 20/20 | | 67 |
| K3 | Coding — Security Review | 13/20 | | 70.8 |
| K4 | Coding — Performance | 20/20 | | 71.4 |
| K5 | Coding — Refactoring Under Constraints | 20/20 | | 74.9 |
| K6 | Coding — Test Writing | 18/20 | | 69.2 |
| K7 | Coding — React Component (typed) | 16/20 | | 79.3 |
| K8 | Coding — TypeScript Type System | 19/20 | | 68.6 |
| K9 | Coding — React Debugging | 16/20 | | 80.7 |
Category mean Coding — Concurrency 19Coding — Algorithms 20Coding — Security Review 13Coding — Performance 20Coding — Refactoring Under Constraints 20Coding — Test Writing 18Coding — React Component (typed) 16Coding — TypeScript Type System 19Coding — React Debugging 16
▶ Content-production depth 1e 100.2 / 120 16.7 6/6 n=1 87.3 run 2026-08-02
| Test | Category | Score | | tok/s |
| E1 | SOW & Requirements Decomposition | 20/20 | | 81.9 |
| E2 | SOW & Requirements Decomposition | 16/20 | | 67.2 |
| E3 | Executive & Proposal Content | 17/20 | | 84.4 |
| E4 | Executive & Proposal Content | 15/20 | | 78.5 |
| E5 | Executive & Proposal Content | 17/20 | | 73.1 |
| E6 | Executive & Proposal Content | 15/20 | | 76.5 |
Category mean SOW & Requirements Decomposition 18Executive & Proposal Content 16
▶ Long-context retrieval & synthesis 1g 100 / 120 20 5/6 ● n=1 81.6 run 2026-09-17
| Test | Category | Score | | tok/s |
| L1 | Long-Context — Needle Retrieval | 20/20 | | 7.8 |
| L2 | Long-Context — Grounded QA with Distractors | 20/20 | | 4.7 |
| L3 | Long-Context — Ops Log Reasoning | 20/20 | | 16.2 |
| L4 | Long-Context — Faithful Summarization | 20/20 | | 29.8 |
| L5 | Long-Context — Instruction Retention | 20/20 | | 7.3 |
Category mean Long-Context — Needle Retrieval 20Long-Context — Grounded QA with Distractors 20Long-Context — Ops Log Reasoning 20Long-Context — Faithful Summarization 20Long-Context — Instruction Retention 20
▶ Long-context multi-needle (MRCR) 1g2 15.9 / 60 5.3 3/3 n=1 69.3 run 2026-09-02
| Test | Category | Score | | tok/s |
| L7 | Long-Context — Multi-Needle Retrieval (MRCR) | 7/20 | | 28.8 |
| L8 | Long-Context — Multi-Needle Retrieval (MRCR) | 5/20 | | 16.8 |
| L9 | Long-Context — Multi-Needle Retrieval (MRCR) | 4/20 | | 9.5 |
▶ Live one-shot builds (runtime-verified) 1h 68 / 80 17 4/4 n=1 74.3 run 2026-09-04
| Test | Category | Score | | tok/s |
| H1 | Web-Estate One-Shot Builds | 19/20 | | 62 |
| H2 | Web-Estate One-Shot Builds | 19/20 | | 83.8 |
| H3 | Web-Estate One-Shot Builds | 17/20 | | 64.7 |
| H4 | Web-Estate One-Shot Builds | 13/20 | | 77.2 |
▶ Brand-constrained static build (negative-constraint adherence) 1j 90 / 100 18 5/5 n=1 88.2 run 2026-09-16
| Test | Category | Score | | tok/s |
| J1 | Brand-Constrained Static Build | 20/20 | | 85.1 |
| J3 | Brand-Constrained Static Build | 20/20 | | 83.7 |
| J4 | Brand-Constrained Static Build | 17/20 | | 73.9 |
| J5 | Brand-Constrained Static Build | 17/20 | | 86.4 |
| J6 | Brand-Constrained Static Build | 16/20 | | 84.2 |
Each test is LLM-judged 0–20 against a fixed rubric; suite max = tests × 20.
An ADJ chip marks a cell carrying a subjective
adjustment from the maintainer; the judged score is the one shown, is unchanged,
and remains what the axes, ranks and tie bands are computed from. Hover for the reason.
Rows with a ▶ expand to the per-test breakdown.
Where fewer tests ran than the suite total, unrun tests count as zero toward the axis score.
See methodology.