Suite results
Suite Score Avg / 20 Tests n tok/s Run
▶ General capability (13-task real-workload suite) 1a 162 / 260 18 9/13 ● n=1.3 128.2 run 2026-09-02
| Test | Category | Score | | tok/s |
| A1 | Dev/Ops Scripting | 10/20 | | 75.3 |
| A2 | Dev/Ops Scripting | 19/20 | | 82.3 |
| A3 | Dev/Ops Scripting | 18/20 | | 97.4 |
| A4 | Dev/Ops Scripting | 18/20 | | 89.3 |
| A5 | Dev/Ops Scripting | 17/20 | | 87 |
| A6 | Dev/Ops Scripting | 20/20 | | 104.2 |
| B1 | Document Processing | 20/20 | | 106.3 |
| B2 | Document Processing | 20/20 | | 72 |
| B3 | Document Processing | 20/20 | | 44.1 |
Category mean Dev/Ops Scripting 17Document Processing 20
▶ Agentic tool-calling & protocol adherence 1b 104 / 160 13 8/8 n=1 143.8 run 2026-09-16
| Test | Category | Score | | tok/s |
| WA1 | Tool Calling & Protocol Adherence | 7/20 | | 52.5 |
| WA2 | Tool Calling & Protocol Adherence | 4/20 | | 42.9 |
| WA3 | Tool Calling & Protocol Adherence | 17/20 | | 59.1 |
| WA4 | Tool Calling & Protocol Adherence | 11/20 | | 42.4 |
| WA5 | Tool Calling & Protocol Adherence | 20/20 | | 98.9 |
| WB1 | Agent Robustness & State Management | 20/20 | | 128.4 |
| WB2 | Agent Robustness & State Management | 19/20 | | 100.9 |
| WB3 | Agent Robustness & State Management | 6/20 | | 53.8 |
Category mean Tool Calling & Protocol Adherence 11.8Agent Robustness & State Management 15
▶ Coding depth 1c 162 / 180 18 9/9 n=1 124 run 2026-09-15
| Test | Category | Score | | tok/s |
| K1 | Coding — Concurrency | 20/20 | | 92.9 |
| K2 | Coding — Algorithms | 20/20 | | 105.4 |
| K3 | Coding — Security Review | 17/20 | | 102 |
| K4 | Coding — Performance | 15/20 | | 100.1 |
| K5 | Coding — Refactoring Under Constraints | 20/20 | | 122 |
| K6 | Coding — Test Writing | 16/20 | | 116.8 |
| K7 | Coding — React Component (typed) | 14/20 | | 111.5 |
| K8 | Coding — TypeScript Type System | 20/20 | | 107.2 |
| K9 | Coding — React Debugging | 20/20 | | 107.6 |
Category mean Coding — Concurrency 20Coding — Algorithms 20Coding — Security Review 17Coding — Performance 15Coding — Refactoring Under Constraints 20Coding — Test Writing 16Coding — React Component (typed) 14Coding — TypeScript Type System 20Coding — React Debugging 20
▶ Content-production depth 1e 91.8 / 120 15.3 6/6 n=1 131.6 run 2026-09-10
| Test | Category | Score | | tok/s |
| E1 | SOW & Requirements Decomposition | 19/20 | | 118.3 |
| E2 | SOW & Requirements Decomposition | 20/20 | | 83.8 |
| E3 | Executive & Proposal Content | 17/20 | | 122 |
| E4 | Executive & Proposal Content | 7/20 | | 121.9 |
| E5 | Executive & Proposal Content | 15/20 | | 114.5 |
| E6 | Executive & Proposal Content | 14/20 | | 58.5 |
Category mean SOW & Requirements Decomposition 19.5Executive & Proposal Content 13.3
▶ Long-context retrieval & synthesis 1g 97.8 / 120 16.3 6/6 n=1 104.7 run 2026-09-10
| Test | Category | Score | | tok/s |
| L1 | Long-Context — Needle Retrieval | 20/20 | | 9.4 |
| L2 | Long-Context — Grounded QA with Distractors | 20/20 | | 4.7 |
| L3 | Long-Context — Ops Log Reasoning | 12/20 | | 9.3 |
| L4 | Long-Context — Faithful Summarization | 6/20 | | 28.2 |
| L5 | Long-Context — Instruction Retention | 20/20 | | 9.9 |
| L6 | Long-Context — Full-Haystack QA | 20/20 | | 1.7 |
Category mean Long-Context — Needle Retrieval 20Long-Context — Grounded QA with Distractors 20Long-Context — Ops Log Reasoning 12Long-Context — Faithful Summarization 6Long-Context — Instruction Retention 20Long-Context — Full-Haystack QA 20
▶ Long-context multi-needle (MRCR) 1g2 20.4 / 60 6.8 3/3 n=1.7 71.5 run 2026-09-07
| Test | Category | Score | | tok/s |
| L7 | Long-Context — Multi-Needle Retrieval (MRCR) | 7.5/20 | | 26.5 |
| L8 | Long-Context — Multi-Needle Retrieval (MRCR) | 9/20 | | 19.6 |
| L9 | Long-Context — Multi-Needle Retrieval (MRCR) | 4/20 | | 3.3 |
▶ Live one-shot builds (runtime-verified) 1h 56.8 / 80 14.2 4/4 n=1 125.8 run 2026-08-27
| Test | Category | Score | | tok/s |
| H1 | Web-Estate One-Shot Builds | 15/20 | | 120.8 |
| H2 | Web-Estate One-Shot Builds | 15/20 | | 123 |
| H3 | Web-Estate One-Shot Builds | 13/20 | | 116.8 |
| H4 | Web-Estate One-Shot Builds | 14/20 | | 108.4 |
▶ Production replay (real agent workload) 1i 116.4 / 240 9.7 12/12 n=1 94.3 run 2026-09-26
| Test | Category | Score | | tok/s |
| I1 | Nightly Fix Replay | 7/20 | | 19.3 |
| I2 | Nightly Fix Replay | 1/20 | | 14.3 |
| I3 | Nightly Fix Replay | 1/20 | | 27.1 |
| I4 | Nightly Fix Replay | 5/20 | | 13 |
| I5 | Nightly Fix Replay | 8/20 | | 51.3 |
| I6 | Nightly Fix Replay | 4/20 | | 52.4 |
| I7 | Audit Judgment | 15/20 | | 46.9 |
| I8 | Audit Judgment | 18/20 | | 96 |
| I9 | Cron Dependency Reasoning | 18/20 | | 78.4 |
| I10 | Cron Dependency Reasoning | 13/20 | | 85.3 |
| I11 | Curation Judgment | 12/20 | | 98 |
| I12 | Curation Judgment | 14/20 | | 103.1 |
Category mean Nightly Fix Replay 4.3Audit Judgment 16.5Cron Dependency Reasoning 15.5Curation Judgment 13
▶ Brand-constrained static build (negative-constraint adherence) 1j 89 / 100 17.8 5/5 n=1 132.1 run 2026-09-16
| Test | Category | Score | | tok/s |
| J1 | Brand-Constrained Static Build | 18/20 | | 81.5 |
| J3 | Brand-Constrained Static Build | 20/20 | | 104 |
| J4 | Brand-Constrained Static Build | 20/20 | | 86.9 |
| J5 | Brand-Constrained Static Build | 14/20 | | 103.4 |
| J6 | Brand-Constrained Static Build | 17/20 | | 104.6 |
Each test is LLM-judged 0–20 against a fixed rubric; suite max = tests × 20.
An ADJ chip marks a cell carrying a subjective
adjustment from the maintainer; the judged score is the one shown, is unchanged,
and remains what the axes, ranks and tie bands are computed from. Hover for the reason.
Rows with a ▶ expand to the per-test breakdown.
Where fewer tests ran than the suite total, unrun tests count as zero toward the axis score.
See methodology.