Suite results
Suite Score Avg / 20 Tests n tok/s Run
▶ Agentic tool-calling & protocol adherence 1b 136 / 160 17 8/8 n=1 133.4 run 2026-09-22
| Test | Category | Score | | tok/s |
| WA1 | Tool Calling & Protocol Adherence | 13/20 | | 57.6 |
| WA2 | Tool Calling & Protocol Adherence | 20/20 | | 67.3 |
| WA3 | Tool Calling & Protocol Adherence | 15/20 | | 71.9 |
| WA4 | Tool Calling & Protocol Adherence | 20/20 | | 55.2 |
| WA5 | Tool Calling & Protocol Adherence | 20/20 | | 107 |
| WB1 | Agent Robustness & State Management | 20/20 | | 124.1 |
| WB2 | Agent Robustness & State Management | 20/20 | | 119.8 |
| WB3 | Agent Robustness & State Management | 8/20 | | 85.3 |
Category mean Tool Calling & Protocol Adherence 17.6Agent Robustness & State Management 16
▶ Coding depth 1c 151.2 / 180 16.8 9/9 n=1 129 run 2026-10-01
| Test | Category | Score | | tok/s |
| K1 | Coding — Concurrency | 18/20 | | 120.9 |
| K2 | Coding — Algorithms | 20/20 | | 116 |
| K3 | Coding — Security Review | 14/20 | | 119.6 |
| K4 | Coding — Performance | 19/20 | | 119.7 |
| K5 | Coding — Refactoring Under Constraints | 20/20 | | 125.1 |
| K6 | Coding — Test Writing | 15/20 | | 119.1 |
| K7 | Coding — React Component (typed) | 10/20 | | 126 |
| K8 | Coding — TypeScript Type System | 17/20 | | 118 |
| K9 | Coding — React Debugging | 18/20 | | 124.3 |
Category mean Coding — Concurrency 18Coding — Algorithms 20Coding — Security Review 14Coding — Performance 19Coding — Refactoring Under Constraints 20Coding — Test Writing 15Coding — React Component (typed) 10Coding — TypeScript Type System 17Coding — React Debugging 18
▶ Doc/OCR vision 1d 151.2 / 160 18.9 8/8 n=1.1 120.2 run 2026-08-06
| Test | Category | Score | | tok/s |
| D1 | Doc/OCR — Transcription | 20/20 | | 68.2 |
| D2 | Doc/OCR — Document QA | 20/20 | | 19.6 |
| D3 | Doc/OCR — Field Extraction | 20/20 | | 60.2 |
| D4 | Doc/OCR — Document QA | 11/20 | | 32.6 |
| D5 | Doc/OCR — Transcription | 20/20 | | 59 |
| D6 | Doc/OCR — Field Extraction | 20/20 | | 42 |
| D7 | Doc/OCR — Degraded Input | 20/20 | | 75.4 |
| D8 | Doc/OCR — Degraded Input | 20/20 | | 70.5 |
Category mean Doc/OCR — Transcription 20Doc/OCR — Document QA 15.5Doc/OCR — Field Extraction 20Doc/OCR — Degraded Input 20
▶ Doc/OCR — real-degraded tier 1d2 48.8 / 80 12.2 4/4 n=1 119.7 run 2026-08-14
| Test | Category | Score | | tok/s |
| D9 | Doc/OCR — Real Degraded Transcription | 6/20 | | 40.4 |
| D10 | Doc/OCR — Handwriting Extraction | 16/20 | | 26 |
| D11 | Doc/OCR — Degraded Table Transcription | 10/20 | | 84 |
| D12 | Doc/OCR — Degraded Table QA | 17/20 | | 29.9 |
Category mean Doc/OCR — Real Degraded Transcription 6Doc/OCR — Handwriting Extraction 16Doc/OCR — Degraded Table Transcription 10Doc/OCR — Degraded Table QA 17
▶ Content-production depth 1e 97.8 / 120 16.3 6/6 n=1 129.4 run 2026-09-25
| Test | Category | Score | | tok/s |
| E1 | SOW & Requirements Decomposition | 20/20 | | 118.9 |
| E2 | SOW & Requirements Decomposition | 18/20 | | 111.7 |
| E3 | Executive & Proposal Content | 16/20 | | 120.4 |
| E4 | Executive & Proposal Content | 13/20 | | 126.6 |
| E5 | Executive & Proposal Content | 16/20 | | 120.6 |
| E6 | Executive & Proposal Content | 15/20 | | 116.2 |
Category mean SOW & Requirements Decomposition 19Executive & Proposal Content 15
▶ Long-context retrieval & synthesis 1g 120 / 120 20 6/6 n=1 71.6 run 2026-09-24
| Test | Category | Score | | tok/s |
| L1 | Long-Context — Needle Retrieval | 20/20 | | 11.7 |
| L2 | Long-Context — Grounded QA with Distractors | 20/20 | | 4.6 |
| L3 | Long-Context — Ops Log Reasoning | 20/20 | | 7 |
| L4 | Long-Context — Faithful Summarization | 20/20 | | 29.1 |
| L5 | Long-Context — Instruction Retention | 20/20 | | 6.6 |
| L6 | Long-Context — Full-Haystack QA | 20/20 | | 1.6 |
Category mean Long-Context — Needle Retrieval 20Long-Context — Grounded QA with Distractors 20Long-Context — Ops Log Reasoning 20Long-Context — Faithful Summarization 20Long-Context — Instruction Retention 20Long-Context — Full-Haystack QA 20
▶ Long-context multi-needle (MRCR) 1g2 29.1 / 60 9.7 3/3 n=1 42.4 run 2026-09-19
| Test | Category | Score | | tok/s |
| L7 | Long-Context — Multi-Needle Retrieval (MRCR) | 20/20 | | 19.2 |
| L8 | Long-Context — Multi-Needle Retrieval (MRCR) | 5/20 | | 9.7 |
| L9 | Long-Context — Multi-Needle Retrieval (MRCR) | 4/20 | | 6.1 |
▶ Live one-shot builds (runtime-verified) 1h 42 / 80 10.5 4/4 n=1 120.7 run 2026-09-21
| Test | Category | Score | | tok/s |
| H1 | Web-Estate One-Shot Builds | 14/20 | | 120.1 |
| H2 | Web-Estate One-Shot Builds | 10/20 | | 124.4 |
| H3 | Web-Estate One-Shot Builds | 9/20 | | 119.2 |
| H4 | Web-Estate One-Shot Builds | 9/20 | | 112.2 |
▶ Brand-constrained static build (negative-constraint adherence) 1j 78 / 100 15.6 5/5 n=1 114.4 run 2026-09-23
| Test | Category | Score | | tok/s |
| J1 | Brand-Constrained Static Build | 18/20 | | 104.2 |
| J3 | Brand-Constrained Static Build | 14/20 | | 108 |
| J4 | Brand-Constrained Static Build | 18/20 | | 104 |
| J5 | Brand-Constrained Static Build | 14/20 | | 110.8 |
| J6 | Brand-Constrained Static Build | 14/20 | | 118.1 |
Each test is LLM-judged 0–20 against a fixed rubric; suite max = tests × 20.
An ADJ chip marks a cell carrying a subjective
adjustment from the maintainer; the judged score is the one shown, is unchanged,
and remains what the axes, ranks and tie bands are computed from. Hover for the reason.
Rows with a ▶ expand to the per-test breakdown.
Where fewer tests ran than the suite total, unrun tests count as zero toward the axis score.
See methodology.