Same prompt, temperature 0
Every model gets the identical request. No second try unless the first run never produced a file. A fluke is part of the score. That is the point of one shot.
This is the bench behind the Fully Tested videos. Not a lab chart. Every ranked model wrote the same games, the same sales page, and sat the same two graded tasks. Open theirs.
One battery. Harder tests count more, so a perfect penguin cannot hide a broken game. The printable object is the heaviest. The two agent tests are graded by a checker, not by me watching.
Every model gets the identical request. No second try unless the first run never produced a file. A fluke is part of the score. That is the point of one shot.
Games and pages are scored against a written rubric, out of 100. The books test and the skepticism test are auto-graded. Taste does not move those two.
Printable object ×4. Halo, tactics, and Diablo ×3. Drone, books, and the sales page ×2. Penguin and the audit ×1. The rank is that weighted average, rounded.
A model with two beautiful games and seven blanks does not get a rank. Unfinished runs sit under the board until the battery is complete.
The games you open are the files the model wrote. They run in a sandbox. They cannot reach the network, other than the one Three.js file the shooter tests are allowed to load. If a game is blank, that is the model’s file, not a broken site.