Aug 2026 ยท Technique

We built six AI critics to interrogate our own competition entry

A panel of AI agents scored the Iloilo scheme three times and reached the same verdict every round. Here is what they said, including the number that never moved.

The question we started from

Most AI architecture starts from an image and works backwards to a justification. We wanted to start from a question instead, and then build something that would refuse to let us answer it lazily.

The question was this. If precolonial spatial, ecological, structural and cultural knowledge in the Philippines had kept evolving under contemporary conditions, rather than being interrupted, what kind of performing arts centre would it produce now?

That is a good question and a useless one on its own, because nothing stops you answering it with a nice render of a timber roof. The interesting part was building the thing that would push back.

Vernacular is not a style

The premise the whole system rests on: vernacular architecture is not a historic look. It is an intelligence system produced by pressure. Generations of adaptation to climate, material and culture, compressed into building knowledge and passed on because it worked.

If you accept that, then a vernacular building is not one that resembles old buildings. It is one that behaves like an evolved response to a place. And that is testable, which is the whole point.

We held the design against four interacting forces. Environmental performance: heat, humidity, rain, flooding, wind, ventilation, shading. Structural intelligence: material logic, joinery, spans, modularity, seismic response, repairability. Behavioural culture: gathering, hospitality, performance ritual, thresholds, informality. Symbolic culture: memory, ancestry, ritual, land and water relationships, craft meaning.

Every element gets the same question. What pressure produced this form? A roof must answer to rain, heat, structure, acoustics, gathering or symbolism. A screen must do environmental, social or material work. A symbolic gesture must connect to programme, orientation, ritual, ecology or construction. If an element cannot answer, it goes.

Six critics, and one that synthesises

The panel was six agents, each responsible for one layer: Filipino materiality, form and design, culture steward, environment steward, structure, and Filipino vernacular intelligence. Above them sat a Vernacular Intelligence Evaluator whose job was to reconcile their readings into a single verdict.

The important constraint, and the one that makes the whole thing worth doing: they were never asked to generate anything. No forms, no images, no references. Their only job was critique. An agent that proposes tends to fall in love with its own proposal, and you end up with a system that agrees with itself.

Each decision went to the panel alongside our own judgement. A move that satisfied the environment steward but failed the culture steward went back to the drawing board. What survived were the moves doing several jobs at once, which is what the four pressures were selecting for all along.

What the panel actually said

We ran three rounds of testing. Here is what the evaluator returned each time, as weightings across the four pressures.

PressureTest 1Test 2Test 3
Environmental performance0.300.250.28
Structural intelligence0.400.400.42
Behavioural culture0.150.150.15
Symbolic culture0.150.200.15
Dominant forceStructuralStructuralStructural

Structural-dominant, three times out of three. The verdict from the first round:

This fragment is currently structural-dominant. Its strongest intelligence is the visible timber tectonic system: heavy posts, layered beams, exposed joinery, bracket-like connections, replaceable-looking members, and a clear attempt to make construction legible. It is much stronger than a purely picturesque Filipino reference because the woven screen, roof overhang, plinth, planting, and timber frame all imply performance. However, many of those performances are still implied rather than proven.

Vernacular Intelligence Evaluator, iteration test 1

It is not yet fully integrated vernacular intelligence because the behavioural and symbolic systems are still less causally legible than the structure.

Vernacular Intelligence Evaluator, iteration test 3

The number that never moved

Look at the behavioural culture row again. 0.15, 0.15, 0.15. Across three rounds of design development, the one thing the scheme never got better at was the thing a performing arts centre is supposedly for.

That is uncomfortable, and it is exactly why the panel was worth building. Left to ourselves we would have read the improving structural score as progress. The building was getting better at being a building and no better at being a place where people gather, and nothing in our own review process was set up to notice.

The phrase that stuck was implied rather than proven. The scheme looked like it was doing environmental and cultural work. The woven screen implied shading. The deep canopy implied monsoon performance. The plinth implied flood response. Implied. The agents kept declining to award credit for a gesture that resembled an answer.

From critique to procedural model

Once the panel and the architects were satisfied, the scheme moved out of rough massing and into a procedural pipeline. Using Claude Code and Codex, we translated the design logic into Python for Houdini so the entire building could be generated procedurally rather than modelled by hand.

The reason was not speed. It was that testing against four pressures only means something if you can keep testing at full resolution. Any parameter could be adjusted and the whole building regenerated: roof geometries, shading densities, woven structures, airflow patterns, modular assemblies, acoustic surfaces. Variations rather than one fixed model.

This is the part most AI architecture skips. Generating an image of a shading screen is trivial. Generating a shading screen whose density is a parameter you can push until the environmental reading changes is a different exercise, and it is the one that turns a picture into a proposition.

What the system is actually for

It is not a design tool. It never drew anything. What it is, is a discipline: a way of making the difference between a decision and a gesture visible while there is still time to act on it.

The value showed up in what it caught. Where the design drifted toward the picturesque. Where a formal idea was leaning on resemblance rather than reasoning. Where we had convinced ourselves that something looked resolved because it looked good. Those are the failures that are almost impossible to catch from inside your own scheme, and they are the ones a competition jury catches immediately.

Where it fell short

Two honest limitations. The panel was much better at diagnosis than at treatment. It could tell us the behavioural systems were not causally legible; it had very little useful to say about how to fix that, and the suggestions it did offer were generic. The design work stayed human, which is fine, but it means this is not a shortcut.

The second is more interesting. A critic that returns 0.15 three times running is either exactly right or not really looking. We believe it was right, because the criticism matched what we already half knew. But that is a weak test, and a panel you only trust when it agrees with your instincts is not a panel. Calibrating that is the next problem.

Since writing this, the entry took first place. That does not settle the question. A jury rewarded a scheme our own critics kept marking down on exactly the axis a performing arts centre should be strongest, and we would rather hold both of those facts than pick the flattering one.

The scheme is published in full as Iloilo Processional Theatre, including the drawings and the film. The procedural and AI side of how we work is set out under AI architectural visualisation, and the competition process under competition visualisation. The rest of the studio's work is at selected work.

←All of AF_LAB Start a project→