Contribute
The leaderboard is single-operator by design. The test surface doesn't have to be.
Phase 1 — now: ideas in, results ours
Every score on this site is run by one operator, on one controlled host, through one pipeline. That isn't a limitation we're working around — it's what makes the numbers comparable. Community-run results would mean uncontrolled hardware, uncontrolled configs, and a leaderboard measuring setups instead of models.
What we want from you right now:
- Test ideas. Workloads you'd actually run on a local model — especially ones where current models embarrass themselves.
- Model requests. A model you think belongs in the gauntlet. Tell us why, link the weights.
Both go through GitHub issues. No forms, no queue theater — an issue tracker we actually read.
Phase 2 — planned: contributed test designs
Full test designs contributed by the community: you design it, we run it on the controlled host, you get credit on the resulting scores. Contributed tests go through the same contamination policy as everything else — prompts and rubrics stay private once they're live.
Community-submitted results
Maybe, eventually, as a clearly separated dataset. Never mixed into the verified leaderboard. A number we didn't measure is a number we won't publish as ours.