Live artifacts
Arena
Every build ships here exactly as the model wrote it — one file, no edits. Static scores grade the code; the browser grades the truth. Each test below has its own page with every published model's build, side by side.
Read the round first: 79/80 on Paper, Cloth Won't Pull in the Browser — the writeup behind the 1H set: a near-perfect static score, then a human loaded the files. The four models it names are marked BROWSER-VERIFIED on their test pages, with Dave's actual verdicts quoted.
Artifacts run in sandboxed iframes (scripts only). Not every model ran every test — each page lists exactly who is published for that one.
Every one of the 293 published files was run once through that same sandboxed embed to check for JavaScript errors before anyone loaded them by hand — 10 throw before the page ever renders. Each is flagged JS ERROR ON LOAD on its test page, artifact shown exactly as written, same as everything else here.
1H · Live web-estate builds
One-shot web-estate builds, scored on whether they work and then checked by hand in a real browser — the axis where a static judge and a human loading the page do not always agree.
H1 · Arcade game
A one-shot canvas arcade game, played as shipped.
32 models published →
H2 · Cloth simulation
A verlet cloth sim — an r/LocalLLaMA favorite. Try to pull the cloth with your mouse.
32 models published →
H3 · Parallax car — community prompt
The community parallax-car prompt, run verbatim — grammar quirks and all — for public comparability.
29 models published →
H4 · Pricing component
A static-site pricing section. Each model shipped an Astro component — source, not a standalone page — so this test is shown as the exact file the model wrote, viewable but not runnable on its own.
4 models published →
H5 · Cloth, loose prompt
The community’s loose cloth prompt, verbatim — “think it through… simulate everything before you write a line of code”. A design-judgment axis rather than a spec axis; scored separately from the 1H/80 total.
29 models published →
1J · Brand-constrained static builds
A different question from 1H: not whether a model can build something that works, but whether it can build something that obeys — a written brief with hard constraints, where the failure mode is a page that looks fine and breaks a rule.
J1 · Doctrine homepage — negative constraints
A single-page build under a brand doctrine written mostly as prohibitions. The hard part is not what the model draws — it is everything it was told not to.
35 models published →
J3 · Self-contradicting brief
A brief that quietly contradicts itself. The test is not which instruction the model follows — it is whether it notices the conflict and says so, instead of silently picking a side.
36 models published →
J4 · Redline revision — fix only what is flagged
Three specific violations, nothing else touched. A surgical-edit test: fix exactly what is flagged and leave the rest of the file exactly as it was.
35 models published →
J5 · Blueprint grid — every element boxed
The inverse constraint: where J1 forbids boxes, J5 demands a visible rule around every element. Same model, opposite instruction — which is the point of running both.
32 models published →
J6 · RFC datasheet — teal/cream grid, mono body
A literal, detailed spec run verbatim: teal-and-cream palette, monospace body, RFC-style datasheet layout. The question is fidelity to spec, not creative judgment.
33 models published →