<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"><channel><title>GAUNTLET Bench — Rounds</title><description>Round-by-round writeups from the GAUNTLET local-model benchmark: what ran, what broke, and what the scores actually mean.</description><link>https://gauntletbench.com/</link><item><title>One Model Claimed 63 Escalations and Delivered 2: First Results from the Production-Replay Suite</title><link>https://gauntletbench.com/rounds/production-replay/</link><guid isPermaLink="true">https://gauntletbench.com/rounds/production-replay/</guid><description>We replayed a real production agent&apos;s nightly documentation work through four local models and scored them against what actually shipped — including a truthfulness check on their own work reports. One model fabricated its numbers. Only one answered all twelve tasks.</description><pubDate>Sun, 16 Aug 2026 00:00:00 GMT</pubDate></item><item><title>79/80 on Paper, Cloth Won&apos;t Pull in the Browser: Why We Now Run Every Front-End Artifact</title><link>https://gauntletbench.com/rounds/browser-verification/</link><guid isPermaLink="true">https://gauntletbench.com/rounds/browser-verification/</guid><description>A new browser-artifact suite produced a near-perfect 79/80 static score — then a human loaded the artifacts. The cloth sim doesn&apos;t respond to the mouse, the car is drawn backwards, and all four models in the field failed the same interaction the same way.</description><pubDate>Sat, 15 Aug 2026 00:00:00 GMT</pubDate></item><item><title>More Bits Isn&apos;t More Brain: A 5-Quant Ladder Where Q6_K Beat Q8_0</title><link>https://gauntletbench.com/rounds/quant-ladder/</link><guid isPermaLink="true">https://gauntletbench.com/rounds/quant-ladder/</guid><description>We ran the same 27B model at five quantization levels through the same 21-test gauntlet at temp 0. Quality was not monotonic with bits: Q6_K took the all-time roster record while Q8_0 — 12 GB heavier than Q4 — bought exactly zero extra points.</description><pubDate>Sat, 15 Aug 2026 00:00:00 GMT</pubDate></item><item><title>The Knob That Wasn&apos;t Attached: Sideloaded MLX Models Silently Lose Their Effort Controls</title><link>https://gauntletbench.com/rounds/sideloaded-mlx-drop/</link><guid isPermaLink="true">https://gauntletbench.com/rounds/sideloaded-mlx-drop/</guid><description>Two MLX builds of the same model kept stalling at settings that worked fine on their GGUF siblings. The reason: LM Studio accepts the reasoning-effort parameter for sideloaded MLX models at the API — then silently discards it.</description><pubDate>Sat, 15 Aug 2026 00:00:00 GMT</pubDate></item><item><title>The Shipped Default Cost 39 Points: When &quot;Think Harder&quot; Makes a Model Worse</title><link>https://gauntletbench.com/rounds/xhigh-minus-39/</link><guid isPermaLink="true">https://gauntletbench.com/rounds/xhigh-minus-39/</guid><description>Qwen3.8-27B ships with reasoning effort set to xhigh. Head-to-head at a 49K token budget, xhigh scored 218/260 against medium&apos;s 257/260 — 5.2x the reasoning tokens, 3.5x the wall time, for a net minus-39.</description><pubDate>Sat, 15 Aug 2026 00:00:00 GMT</pubDate></item></channel></rss>