Live artifacts

Arena

Every build ships here exactly as the model wrote it — one file, no edits. Static scores grade the code; the browser grades the truth. Each test below has its own page with every published model's build, side by side.

Read the round first: 79/80 on Paper, Cloth Won't Pull in the Browser — the writeup behind the 1H set: a near-perfect static score, then a human loaded the files. The four models it names are marked BROWSER-VERIFIED on their test pages, with Dave's actual verdicts quoted.

Artifacts run in sandboxed iframes (scripts only). Not every model ran every test — each page lists exactly who is published for that one.

Every one of the 293 published files was run once through that same sandboxed embed to check for JavaScript errors before anyone loaded them by hand — 10 throw before the page ever renders. Each is flagged JS ERROR ON LOAD on its test page, artifact shown exactly as written, same as everything else here.

1H · Live web-estate builds

One-shot web-estate builds, scored on whether they work and then checked by hand in a real browser — the axis where a static judge and a human loading the page do not always agree.

1J · Brand-constrained static builds

A different question from 1H: not whether a model can build something that works, but whether it can build something that obeys — a written brief with hard constraints, where the failure mode is a page that looks fine and breaks a rule.