Suite results
Suite Score Avg / 20 Tests n tok/s Run
▶ General capability (13-task real-workload suite) 1a 248.3 / 260 19.1 13/13 n=4.9 47.1 run 2026-09-02
| Test | Category | Score | | tok/s |
| A1 | Dev/Ops Scripting | 20/20 | | 86.6 |
| A2 | Dev/Ops Scripting | 18/20 | | 85.8 |
| A3 | Dev/Ops Scripting | 18/20 | | 74.7 |
| A4 | Dev/Ops Scripting | 18/20 | | 72.4 |
| A5 | Dev/Ops Scripting | 20/20 | | 72.7 |
| A6 | Dev/Ops Scripting | 20/20 | | 73.3 |
| B1 | Document Processing | 20/20 | | 75 |
| B2 | Document Processing | 20/20 | | 77.3 |
| B3 | Document Processing | 20/20 | | 77.6 |
| C1 | Content Production | 18/20 | | 77.6 |
| C2 | Content Production | 18/20 | | 76.7 |
| C3 | Content Production | 20/20 | | 75.8 |
| C4 | Content Production | 19/20 | | 74.6 |
Category mean Dev/Ops Scripting 19Document Processing 20Content Production 18.8
▶ Agentic tool-calling & protocol adherence 1b 159.2 / 160 19.9 8/8 n=4 83.9 run 2026-09-12
| Test | Category | Score | | tok/s |
| WA1 | Tool Calling & Protocol Adherence | 20/20 | | 55.7 |
| WA2 | Tool Calling & Protocol Adherence | 20/20 | | 74.2 |
| WA3 | Tool Calling & Protocol Adherence | 20/20 | | 69.1 |
| WA4 | Tool Calling & Protocol Adherence | 20/20 | | 65.8 |
| WA5 | Tool Calling & Protocol Adherence | 20/20 | | 71.6 |
| WB1 | Agent Robustness & State Management | 20/20 | | 87.2 |
| WB2 | Agent Robustness & State Management | 20/20 | | 85.7 |
| WB3 | Agent Robustness & State Management | 19/20 | | 85.9 |
Category mean Tool Calling & Protocol Adherence 20Agent Robustness & State Management 19.7
▶ Coding depth 1c 177.3 / 180 19.7 9/9 n=1 72.4 run 2026-09-15
| Test | Category | Score | | tok/s |
| K1 | Coding — Concurrency | 20/20 | | 77.2 |
| K2 | Coding — Algorithms | 20/20 | | 74.9 |
| K3 | Coding — Security Review | 20/20 | | 69.2 |
| K4 | Coding — Performance | 18/20 | | 63.5 |
| K5 | Coding — Refactoring Under Constraints | 20/20 | | 57.8 |
| K6 | Coding — Test Writing | 20/20 | | 54.3 |
| K7 | Coding — React Component (typed) | 20/20 | | 80.3 |
| K8 | Coding — TypeScript Type System | 19/20 | | 82.7 |
| K9 | Coding — React Debugging | 20/20 | | 78.9 |
Category mean Coding — Concurrency 20Coding — Algorithms 20Coding — Security Review 20Coding — Performance 18Coding — Refactoring Under Constraints 20Coding — Test Writing 20Coding — React Component (typed) 20Coding — TypeScript Type System 19Coding — React Debugging 20
▶ Doc/OCR vision 1d 158.4 / 160 19.8 8/8 n=1.1 82 run 2026-08-06
| Test | Category | Score | | tok/s |
| D1 | Doc/OCR — Transcription | 20/20 | | 83.2 |
| D2 | Doc/OCR — Document QA | 20/20 | | 82.3 |
| D3 | Doc/OCR — Field Extraction | 20/20 | | 78.9 |
| D4 | Doc/OCR — Document QA | 20/20 | | 80.8 |
| D5 | Doc/OCR — Transcription | 18/20 | | 80.2 |
| D6 | Doc/OCR — Field Extraction | 20/20 | | 71.4 |
| D7 | Doc/OCR — Degraded Input | 20/20 | | 74.3 |
| D8 | Doc/OCR — Degraded Input | 20/20 | | 72.5 |
Category mean Doc/OCR — Transcription 19Doc/OCR — Document QA 20Doc/OCR — Field Extraction 20Doc/OCR — Degraded Input 20
▶ Doc/OCR — real-degraded tier 1d2 36 / 80 9 4/4 n=1 68.9 run 2026-08-14
| Test | Category | Score | | tok/s |
| D9 | Doc/OCR — Real Degraded Transcription | 8/20 | | 82.6 |
| D10 | Doc/OCR — Handwriting Extraction | 0/20 | | 62.1 |
| D11 | Doc/OCR — Degraded Table Transcription | 13/20 | | 61.7 |
| D12 | Doc/OCR — Degraded Table QA | 15/20 | | 65.8 |
Category mean Doc/OCR — Real Degraded Transcription 8Doc/OCR — Handwriting Extraction 0Doc/OCR — Degraded Table Transcription 13Doc/OCR — Degraded Table QA 15
▶ Content-production depth 1e 114 / 120 19 6/6 n=1.2 81.5 run 2026-08-02
| Test | Category | Score | | tok/s |
| E1 | SOW & Requirements Decomposition | 20/20 | | 84.4 |
| E2 | SOW & Requirements Decomposition | 20/20 | | 82.7 |
| E3 | Executive & Proposal Content | 20/20 | | 78.4 |
| E4 | Executive & Proposal Content | 15/20 | | 83 |
| E5 | Executive & Proposal Content | 19/20 | | 79.2 |
| E6 | Executive & Proposal Content | 20/20 | | 74.1 |
Category mean SOW & Requirements Decomposition 20Executive & Proposal Content 18.5
▶ Long-context retrieval & synthesis 1g 120 / 120 20 6/6 n=1.2 71 run 2026-08-06
| Test | Category | Score | | tok/s |
| L1 | Long-Context — Needle Retrieval | 20/20 | | 60.9 |
| L2 | Long-Context — Grounded QA with Distractors | 20/20 | | 65.4 |
| L3 | Long-Context — Ops Log Reasoning | 20/20 | | 48.8 |
| L4 | Long-Context — Faithful Summarization | 20/20 | | 54.6 |
| L5 | Long-Context — Instruction Retention | 20/20 | | 57.1 |
| L6 | Long-Context — Full-Haystack QA | 20/20 | | 33.9 |
Category mean Long-Context — Needle Retrieval 20Long-Context — Grounded QA with Distractors 20Long-Context — Ops Log Reasoning 20Long-Context — Faithful Summarization 20Long-Context — Instruction Retention 20Long-Context — Full-Haystack QA 20
▶ Long-context multi-needle (MRCR) 1g2 20.1 / 60 6.7 3/3 n=1 50.5 run 2026-08-14
| Test | Category | Score | | tok/s |
| L7 | Long-Context — Multi-Needle Retrieval (MRCR) | 20/20 | | 61 |
| L8 | Long-Context — Multi-Needle Retrieval (MRCR) | 0/20 | | 41.5 |
| L9 | Long-Context — Multi-Needle Retrieval (MRCR) | 0/20 | | 36 |
▶ Live one-shot builds (runtime-verified) 1h 68 / 80 17 4/4 n=1 79.8 run 2026-08-15
| Test | Category | Score | | tok/s |
| H1 | Web-Estate One-Shot Builds | 15/20 | | 73.5 |
| H2 | Web-Estate One-Shot Builds | 19/20 | | 78.2 |
| H3 | Web-Estate One-Shot Builds | 17/20 | | 79.4 |
| H4 | Web-Estate One-Shot Builds | 17/20 | | 74.5 |
▶ Production replay (real agent workload) 1i 141.6 / 240 11.8 12/12 n=1 56.6 run 2026-08-16
| Test | Category | Score | | tok/s |
| I1 | Nightly Fix Replay | 0/20 | | 59.2 |
| I2 | Nightly Fix Replay | 0/20 | | 47.5 |
| I3 | Nightly Fix Replay | 8/20 | | 36.7 |
| I4 | Nightly Fix Replay | 15/20 | | 51.7 |
| I5 | Nightly Fix Replay | 13/20 | | 51.1 |
| I6 | Nightly Fix Replay | 0/20 | | 50.9 |
| I7 | Audit Judgment | 15/20 | | 59.3 |
| I8 | Audit Judgment | 20/20 | | 65.2 |
| I9 | Cron Dependency Reasoning | 19/20 | | 56.7 |
| I10 | Cron Dependency Reasoning | 15/20 | | 56.8 |
| I11 | Curation Judgment | 20/20 | | 62.5 |
| I12 | Curation Judgment | 17/20 | | 61 |
Category mean Nightly Fix Replay 6Audit Judgment 17.5Cron Dependency Reasoning 17Curation Judgment 18.5
▶ Brand-constrained static build (negative-constraint adherence) 1j 76 / 100 15.2 5/5 n=1 72.2 run 2026-09-13
| Test | Category | Score | | tok/s |
| J1 | Brand-Constrained Static Build | 12/20 | | 64.3 |
| J3 | Brand-Constrained Static Build | 20/20 | | 68.6 |
| J4 | Brand-Constrained Static Build | 20/20 | | 69.8 |
| J5 | Brand-Constrained Static Build | 10/20 | | 77 |
| J6 | Brand-Constrained Static Build | 14/20 | | 78.4 |
Each test is LLM-judged 0–20 against a fixed rubric; suite max = tests × 20.
An ADJ chip marks a cell carrying a subjective
adjustment from the maintainer; the judged score is the one shown, is unchanged,
and remains what the axes, ranks and tie bands are computed from. Hover for the reason.
Rows with a ▶ expand to the per-test breakdown.
Where fewer tests ran than the suite total, unrun tests count as zero toward the axis score.
See methodology.