BANNERALL
Chapter one · the style sweep

Choosing a look

written 2026-09-02

We asked a 3D-generation service to make the same watchtower, wolf and foot soldier in 10 different art styles, wrote down how we would judge them before we looked, and then watched the experiment declare itself broken.

Bannerfall's board is built from thousands of 3D models. We are trying to find out whether a machine can make them, and the first thing you have to settle is what they should look like — because if you pick the look after you see the results, you have not run an experiment, you have picked a favourite.

So this one was set up the other way round. The rules were written, committed and pushed first. Then 36 models were ordered. Then we looked. What follows is what came back, including the parts where the experiment turned round and told us it was wrong.

Anything set like 0.0665 was measured by a machine and generated into this page from the committed results. Everything else was written by a person.


The rules went in before the pictures came out

Everything about the judgement — which styles were allowed to win, how much a shape had to change before the change counted, what would disqualify a model outright — was fixed in a document and pushed to the server at 2026-08-31T22:26:52Z. The first model was ordered at 2026-08-31T22:27:07.804Z. That is 15.8 seconds later.

Here is the caveat, and it matters. You cannot see that ordering in this project's main history, because branches are squashed into a single commit when they land — the rules and the results arrived in the same commit, after the fact. What actually pins it down is the commit a2be209d on the branch, whose id is written into the code that ran the sweep and is checked by a test. "Trust me, we wrote it down first" is precisely the thing writing it down first is supposed to replace, so: don't trust it, check that commit.

A watchtower, a wolf, and a foot soldier walk into a render farm

There are 3 subjects, described once in plain words and never changed: a ruined stone watchtower, a large four-legged wolf, and a standing foot soldier with a spear and a round shield. Each was ordered 12 times — once with no style at all, once more with no style at a different random seed so we could measure how much the machine disagrees with itself, and then once for each of the 10 style names the service offers. Every one was given the same budget of 4000 triangles.

That second no-style order is the most useful thing in the whole design. Ask for the same wolf twice and you do not get the same wolf. The gap between those two is the noise floor, and any style that does not move a model further than that has not done anything you could not have got by rolling the dice again.

It took 4h 47m end to end, one model at a time, at a median of 7.9 minutes each — 4.8 hours of graphics card, for 36 models.

The question it was built to answer, answered: no

The real question was whether a style name changes a model's shape or only its paint. This matters for a strategy board, where you read a unit by its outline at a glance. If style changed shape, each faction could have its own style and its armies would be distinguishable across the map. If style only changed the paint, that idea is dead.

The rule set beforehand: an id had to push the outline further than a reseed does, by more than 0.10, in at least 2 of the 3 subjects, and at least 3 ids had to manage it.

None did. Of the 10 ids, 0 cleared it. Averaged over the 3 subjects, the largest displacement any id managed belongs to chibi, at 0.0665 — short of the floor before you even get to the question of breadth.

Individual casts are a separate population and they are worth stating separately, because they are not all inside the floor: 2 of the 30 styled models did move an outline past it, the largest being chibi at 0.1913. But a style that moves one subject and leaves the other two where they were is not a style that changes form, and both of those casts are the same subject. That is what the rule was asking about when it demanded 2 subjects and 3 ids. On that question, nothing.

Style moves surface, not silhouette. Per-faction art styles were ruled out by the data rather than by anyone's opinion, which is the only way you want that kind of thing ruled out.

largest silhouette displacement, any id

0.0665
The line, fixed in advance: 0.10 — "the style changed the form"Did not reach it.

Then the verdict announced that its own rule was wrong

One of the pre-registered disqualifications said a character or creature must be at least as tall as it is wide — a sensible way to catch a delivery that came back as a flat smear. It was written against the subject's category.

The creature is a wolf on four legs. A wolf is longer than it is tall. Every single one of the 12 wolves failed that rule, including the un-styled control, which measured 0.949 against a floor of 1.00. The control is the baseline the whole comparison is measured from. When your baseline fails your own gate, the gate is not measuring what you thought.

The verdict this produced was reassuring: outcome (d), style: null. Changing nothing was best. That is exactly the shape of answer you should distrust, and it was a property of the rule, not of the models. It was written down as a defect rather than quietly patched, and the rule was replaced in a separate piece of work: the aspect check is now keyed to the pose that was actually requested, with a band for a quadruped of 0.25 to 1.25 instead of a single floor. Under the corrected rule the outcome becomes (a) — one id programme-wide, toon.

The correction was made with all the data already on screen, and it makes the more convenient answer easier to reach. That is the direction to be suspicious in, so both rules are still runnable and the report re-runs the whole judgement under each. The part that decides everything — style moves surface, not silhouette — comes out identical either way.

the un-styled wolf, against the upright rule

0.949
The line, fixed in advance: 1.00 — "at least as tall as it is wide"Did not reach it.

A cartoon face, carved into a watchtower

Look at what the chibi style did to the watchtower. It did not restyle the stonework. It put two eyes and a mouth on the tower.

The instrument never saw it. The measurement compares outlines from a fixed camera, and a face painted and pressed into a wall does not change an outline: this tower scored 0.8158 against a noise floor of 0.8448, a difference of 0.0290 — well inside the range you get from re-rolling the same order. Nothing in the numbers flagged it. What did flag it was the proportions: it is the largest deviation in the whole sweep, off by 23.4 percent from the control tower.

This is the honest limit of the thing. An outline measurement answers "did the shape move". It cannot answer "is this still the object I asked for". That second question was answered by looking, which is why the pictures are on this page instead of a table of scores.

the watchtower with a face, against the seed-noise floor

0.0290
The line, fixed in advance: 0.10 — "the outline moved at all"Did not reach it.
A ruined stone watchtower generated with the style setting "no style", shown twice: the service's own render on the left and this project's render of the delivered geometry on the right.no style
A ruined stone watchtower generated with the style setting "chibi", shown twice: the service's own render on the left and this project's render of the delivered geometry on the right.chibi
The watchtower as ordered, and the same watchtower with a face. In each pair the left frame is the service's own render and the right is ours, of the actual geometry it delivered.

Half the wolves came back with their heads inside their chests

The wolf was the unstable subject from the start. Two no-style orders of the same wolf agree with each other only to 0.6135, against 0.8448 for the watchtower and 0.8689 for the soldier. A wolf disagrees with itself about twice as much as anything else does, which puts a hard ceiling on what any comparison between wolves can show.

Then there is what they look like. 6 of the 12 came back with the head lowered and fused into the chest and ruff, legs merged into a slab, tail reduced to a lump. The other 6 are an animal: a muzzle, ears, four legs, a tail that reads as a tail.

There is a pattern in which is which — every clean one carries a cartoonish style name, every collapsed one carries a painterly or photoreal one — and there is a plausible story about why, involving fur being modelled as strands and eating the triangle budget. Both of those are worth a real experiment.

They are not worth believing yet, and here is why, stated before the interesting part rather than after it: the un-styled control is in the collapsed group, and the second no-style order — the same wolf, one seed away — is in the clean group. Random chance alone can move a wolf across this line, and there is exactly one of each. This is a hypothesis. It has its own ticket and it needs a repeat.

A four-legged wolf generated with the style setting "no style", shown twice: the service's own render on the left and this project's render of the delivered geometry on the right.no styleoutline 1.0000head collapsed
A four-legged wolf generated with the style setting "no style, reseeded", shown twice: the service's own render on the left and this project's render of the delivered geometry on the right.no style, reseededoutline 0.6135reads as an animal
A four-legged wolf generated with the style setting "low-poly", shown twice: the service's own render on the left and this project's render of the delivered geometry on the right.low-polyoutline 0.6225reads as an animal
A four-legged wolf generated with the style setting "toon", shown twice: the service's own render on the left and this project's render of the delivered geometry on the right.toonoutline 0.6304reads as an animal
A four-legged wolf generated with the style setting "realistic", shown twice: the service's own render on the left and this project's render of the delivered geometry on the right.realisticoutline 0.6007head collapsed
A four-legged wolf generated with the style setting "hand-painted", shown twice: the service's own render on the left and this project's render of the delivered geometry on the right.hand-paintedoutline 0.5314head collapsed
A four-legged wolf generated with the style setting "voxel", shown twice: the service's own render on the left and this project's render of the delivered geometry on the right.voxeloutline 0.6520reads as an animal
A four-legged wolf generated with the style setting "chibi", shown twice: the service's own render on the left and this project's render of the delivered geometry on the right.chibioutline 0.6341reads as an animal
A four-legged wolf generated with the style setting "clay", shown twice: the service's own render on the left and this project's render of the delivered geometry on the right.clayoutline 0.6793reads as an animal
A four-legged wolf generated with the style setting "dark-fantasy", shown twice: the service's own render on the left and this project's render of the delivered geometry on the right.dark-fantasyoutline 0.5472head collapsed
A four-legged wolf generated with the style setting "scifi", shown twice: the service's own render on the left and this project's render of the delivered geometry on the right.scifioutline 0.5772head collapsed
A four-legged wolf generated with the style setting "steampunk", shown twice: the service's own render on the left and this project's render of the delivered geometry on the right.steampunkoutline 0.7224head collapsed
All 12 wolves. The ones marked as collapsed are the ones with the head fused into the chest — including the control, which is the whole problem.

And then a person looked at them and picked a different one

The judge, running the rules it had been given, chose toon on a composite score of 0.7772. The owner of the game looked at all 12 foot soldiers and chose hand-painted, which scored 0.7568 and had in fact been disqualified outright before scoring — on a second gate, one that flags a delivery the service is not confident about. It flagged 5 of them, every one a character, the un-styled control included. That is the wolf rule's failure again, wearing a different rule, and it has its own ticket too.

That is an override, and it is written down as one rather than folded into the numbers. No threshold was retuned to make it win. The report still says toon won.

It is also defensible on the sweep's own finding, which is the part worth sitting with. The experiment established that no style moved a shape past the line it had set for "different at all". When the differences between ids are that far below the threshold, a composite score is ranking noise — and the question left over is not "which is measurably different" but "which do we want to look at". That question is about surface, and surface is exactly what this instrument is blind to. The face in the watchtower proved that, with a number.

So the field is pinned to hand-painted: hand-painted for props, hand-painted for characters, hand-painted for creatures. That is 3 separate entries carrying the same name today, so that changing one of them later is one line and a stated reason rather than a rewrite.

A standing foot soldier generated with the style setting "no style", shown twice: the service's own render on the left and this project's render of the delivered geometry on the right.no styleoutline 1.0000
A standing foot soldier generated with the style setting "no style, reseeded", shown twice: the service's own render on the left and this project's render of the delivered geometry on the right.no style, reseededoutline 0.8689
A standing foot soldier generated with the style setting "low-poly", shown twice: the service's own render on the left and this project's render of the delivered geometry on the right.low-polyoutline 0.8224
A standing foot soldier generated with the style setting "toon", shown twice: the service's own render on the left and this project's render of the delivered geometry on the right.toonoutline 0.8507
A standing foot soldier generated with the style setting "realistic", shown twice: the service's own render on the left and this project's render of the delivered geometry on the right.realisticoutline 0.9282
A standing foot soldier generated with the style setting "hand-painted", shown twice: the service's own render on the left and this project's render of the delivered geometry on the right.hand-paintedoutline 0.8680
A standing foot soldier generated with the style setting "voxel", shown twice: the service's own render on the left and this project's render of the delivered geometry on the right.voxeloutline 0.7712
A standing foot soldier generated with the style setting "chibi", shown twice: the service's own render on the left and this project's render of the delivered geometry on the right.chibioutline 0.6776
A standing foot soldier generated with the style setting "clay", shown twice: the service's own render on the left and this project's render of the delivered geometry on the right.clayoutline 0.7567
A standing foot soldier generated with the style setting "dark-fantasy", shown twice: the service's own render on the left and this project's render of the delivered geometry on the right.dark-fantasyoutline 0.9436
A standing foot soldier generated with the style setting "scifi", shown twice: the service's own render on the left and this project's render of the delivered geometry on the right.scifioutline 0.9022
A standing foot soldier generated with the style setting "steampunk", shown twice: the service's own render on the left and this project's render of the delivered geometry on the right.steampunkoutline 0.9321
The 12 foot soldiers, as they were compared. This is the picture the decision was made from.

One more thing the service cannot do yet

Before any of this, a separate pass poked at 324 combinations of the service's settings to find out which are real and which only exist in the documentation. The most consequential answer was to the question of whether it can produce a model with a skeleton in it — something you can animate.

Every one of the 3 attempts came back 422: "rig is not available in this version of the gateway". A strategy board needs units that walk, so that single refusal is why every model this programme produces has to go through a separate rigging stage afterwards, and why the interesting work is not really about art styles at all.

The whole record, if you want it

Every number on this page was read out of the committed measurements by a script and written into the page's data, so it cannot drift from the file it came from — a test regenerates it and fails on any difference. The prose around them is written by hand, which is why it can be wrong in ways a test cannot catch. If you find something wrong in it, that is a bug worth reporting.

The unabridged version is in the repository: the generated verdict, the pre-registration with its declared amendments, the raw measurements for all 36 models, and the contact sheets these crops were cut from.

← All chapters