3D objects AI agents built from one written task, measured in a fixed studio.
BAY 01 Gallery
Every test we run. Open one to see how each build did.
Each test is one object that every model was asked to build from the same description. The picture is the best build by measured checks. Looks are not measured: where builds pass the same checks, a person picked the cover by eye. Press Play to run it live.
BAY 02 Inspection wall
Every build of this test.
Press Play to run a build. The lights are its checks. What it said opens in one panel, for the build you press.
The test
BEST BUILD
BAY 03 Compare bench
Put any two results side by side.
Pick a build and an effort level for each side. A model is never on both sides. Same studio, same camera, same clock. One fact per row; rows that differ are marked. There is no total.
BAY 04 Stats
The numbers, apart from the pictures.
Every number below is measured on the builds themselves. Nothing is typed in by hand, and a number that was not recorded says so instead of showing a guess.
Leaderboard
Model against model, one chart per stat
The leaderboard’s numbers all at once: every model side by side on each stat, each stage and each group of tests. The marked bar is the best on that stat. Only the Overall chart mixes kinds of number, and it says so; a value no run recorded gets no bar.
This test, build by build
How to read the numbers
- Claims. What the model said about its build, in its own words. A claim counts as true or false only where it can be tested; the rest read “not checked”.
- Checks. The same checks for every build of a test.
- Parts and triangles. Counted from the object as delivered.
- Time, tokens and cost. As reported by the tool the model ran in. Where a tool reports no cost, it is worked out from what the run used at the listed price, and marked. A value that was not recorded says so; it is never guessed.
- One blend. The Overall tab is a weighted average of six scores and says how it is worked out. Every other tab shows one measured thing; Cost against checks sets two side by side without adding them up.
BAY 05 What the scores mean
What is measured, and what these results cannot tell you.
Every model is given the same thing to build and is measured the same way. Here is what a check means, and where to be careful reading the numbers.
The same for every model
- The same task. Every model gets exactly the same description of the object to build.
- One try. Each model builds it once and cannot test or fix it first. What you see is that first attempt, never the best of several.
- Unchanged. What you see is what the model delivered. Nothing is fixed or tidied.
- Private tasks. The task descriptions are not published, so a model cannot be tuned to them. Each test shows a short description instead.
What a check means
- Look. The parts it needs are there, can be seen, and sit where they belong.
- Move. Left alone, the right parts move by themselves, at the right pace.
- Work. It does its job: a wrong input changes nothing, the right one does what the real object would.
- Adapt. It still works when we change a value.
- Physics. Where the object should follow real physics, we judge where things end up against an independent calculation.
- Its own claims. What the model said about its build, in its own words. A claim is marked true or false only where it can be tested.
Same checks, very different look
What these results cannot tell you
In short: few tests, one try per model on each test, looks are not measured, and every cost is a list price.
- Few tests, of one kind. Every test here is one small object built in one try. A model’s place could change with more tests or other kinds of work.
- One try is one try. A model can give different results on the same task. A difference between two models on one test may be chance.
- Results, not method. A build that reaches the right result passes, however it got there.
- Not everything is measured. We measure which parts exist, how they move and where they sit. We do not measure whether it looks good.
- Cost is the list price for that run, reported by the tool or worked out from what the run used. On a subscription you do not pay that amount per run.
- One-shot building, not everyday coding. Models here cannot run their code and try again, which real work allows.
- Checked failures. When a build failed, the failure was checked against the task before it was counted.
BAY 06 Line-up
Pick models. Keep adding.
Put as many models side by side as you like: one column each, the same numbers as the leaderboard, and every test’s build as the studio pictured it. Press a model to add it or take it out.
Across every test
One row per measured thing. “Best” marks the best value among the models you picked; it is left out when they are level.
Test by test
Each picture is that model’s one build of the test, seen from above. Looks are not measured. Open a test to see its builds run.
SCENES Scene benchmarks
Whole scenes, not single objects.
Coming later: whole scenes, where several things have to work together in one place. Nothing has been run yet.
Nothing run yet
0 scene tests · 0 builds. Every result on this site so far is an object benchmark.
APPS App benchmarks
Small working apps, not objects.
Coming later: small apps, checked on whether they take input, refuse bad input and keep what you saved. Nothing has been run yet.
Nothing run yet
0 app tests · 0 builds.
TOOLS Agent tools
The same model, through different agent tools.
Coming later: the same model through different coding tools, to show how much the tool changes the result. Nothing has been run yet.
Nothing run yet
0 tools compared · every build so far is Claude Code.
FIX-IT Fix-it benchmarks
Given a broken build, can it find the fault and repair it?
A model is given a build that does not work and one try to repair it. The table shows how many checks pass before and after, whether the repair broke anything that worked, and how much of the original it kept.
Nothing run yet
0 repair tests · 0 repairs.
ROUNDS Long-task benchmarks
Built in rounds: does it still work after the fourth change?
One object is built in steps: each round asks for a change to what the same model built before. The table shows how many checks pass after each round, and whether a change broke something that worked earlier.
Nothing run yet
0 rounds run · the first object, a tower clock in four rounds, is written and waiting.
JUDGE Judgement benchmarks
The task disagrees with itself. Does the model say so?
Real instructions are not always right. Here the task contains two statements that cannot both be true. What counts is whether the model says so: “yes” means it pointed out the clash, “partly” means it only listed something as not done, “no” means it said nothing.
Nothing run yet
0 models asked · the first task, a gearbox, is written and waiting.
07 Looks
Looks, picked by eye
Every test, top to bottom. In each test the same task was built by every model on the bench; each picture names its model. The studio does not measure looks. One person, the owner of this bench, looked at each test’s builds from above and picked the one that looks best; that build is marked “picked by eye”. It is one person’s taste, not a measurement. Press “Play this row” to run its builds and judge for yourself; one test runs at a time.