Arena

Every Live-axis build ships here exactly as the model wrote it — one file, no edits. Static scores grade the code; the browser grades the truth.

Read the round first: 79/80 on Paper, Cloth Won't Pull in the Browser — the writeup behind these artifacts: a near-perfect static score, then a human loaded the files. Every browser-verification verdict quoted below comes from that article.

Artifacts run in sandboxed iframes (scripts only). Suite totals shown are the static 1H score out of 80; per-test static scores aren't published, so none are shown.

H1 · Arcade game

A one-shot canvas arcade game, played as shipped.

Browser verification reshuffled this podium: the in-browser winner was the 35B MoE, not the static leader.

H2 · Cloth simulation

A verlet cloth sim — an r/LocalLLaMA favorite. Try to pull the cloth with your mouse.

Browser verification: the cloth interaction failed 0-for-4 across the entire field (worst case: the cloth disappears on click). The cloth renders and hangs; the mouse is where it breaks.

Qwen3.8-27B Q6_K GGUF

1H suite: 79/80 Open artifact ↗

Browser check: renders and hangs beautifully, but it won’t pull with the mouse and tearing behaves oddly — likely a mouse-coordinate/device-pixel-ratio mapping bug.

H3 · Parallax car — community prompt

The community parallax-car prompt, run verbatim — grammar quirks and all — for public comparability.

Browser verification: two of the four models drew the car backwards.

H4 · Pricing component

A static-site pricing section. Each model shipped an Astro component — source, not a standalone page — so this test is shown as the exact file the model wrote, viewable but not runnable on its own.

H5 · Cloth, loose prompt

The community’s loose cloth prompt, verbatim — “think it through… simulate everything before you write a line of code”. A design-judgment axis rather than a spec axis; scored separately from the 1H/80 total.

Qwen3.8-27B Q6_K GGUF

1H suite: 79/80 H5: 17/20 Open artifact ↗

This artifact is the low-reasoning-effort rerun. At its production medium effort the model stalled outright — 24,575 reasoning tokens, 22.7 minutes, zero-token answer — so there was nothing to ship from that run.