Fully TestedSeason 3 @mattjohnston_ai

Same nine prompts. Then you play what they built.

This is the bench behind the Fully Tested videos. Not a lab chart. Every ranked model wrote the same games, the same sales page, and sat the same two graded tasks. Open theirs.

The nine tests

One battery. Harder tests count more, so a perfect penguin cannot hide a broken game. The printable object is the heaviest. The two agent tests are graded by a checker, not by me watching.

How a number gets on the board

Same prompt, temperature 0

Every model gets the identical request. No second try unless the first run never produced a file. A fluke is part of the score. That is the point of one shot.

70 for correct, 30 for craft

Games and pages are scored against a written rubric, out of 100. The books test and the skepticism test are auto-graded. Taste does not move those two.

Hard tests count more

Printable object ×4. Halo, tactics, and Diablo ×3. Drone, books, and the sales page ×2. Penguin and the audit ×1. The rank is that weighted average, rounded.

No score until all nine exist

A model with two beautiful games and seven blanks does not get a rank. Unfinished runs sit under the board until the battery is complete.

The games you open are the files the model wrote. They run in a sandbox. They cannot reach the network, other than the one Three.js file the shooter tests are allowed to load. If a game is blank, that is the model’s file, not a broken site.