Aug 2026 · Technique

Most of our AI now runs on a computer in the studio

Open-weights models have become good enough to run on a single graphics card. That means client material never has to leave the building. Here is the stack we use, the reasoning, and the one job we still send out.

Start with where the file goes

Clients hand us material that is not public yet. Unbuilt schemes, sites that have not been announced, competition entries under embargo, drawings covered by an NDA. That is normal for this work, and it is the constraint everything else has to fit around.

Most AI tools work one way. You upload your material to a company's servers, their model runs there, and a result comes back. The terms are usually reasonable and the companies are usually serious. But a term is a promise about what someone intends to do with your material, not a fact about where your material is. Once the file has left the building, you are trusting a policy.

There is a second way to run a model, and it has only recently become practical at the quality this work needs. You download the model itself and run it on a computer you own. Nothing is uploaded, because there is nowhere to upload it to.

What open weights actually means

For anyone who has not had reason to care yet: an AI model is a very large file of numbers, produced by a training run that costs millions. Using the model is arithmetic performed on that file. None of that arithmetic requires the internet.

So the only question is whether the company that trained the model publishes the file. When it does, the model is described as open weights. Anyone with a suitable computer can download it and run it, offline, indefinitely.

Open weights is not the same as open source, and it is worth being precise about that. The training data and the training code are usually not published, and the licences differ. Qwen3.8 and Z-Image Turbo are Apache 2.0, about as permissive as it gets. LTX-2.5 is free to use below a revenue threshold. Read the licence before you build a studio on the model, the same as with any other dependency.

Here is what the choice changes in practice.

Hosted modelOpen model, run here
Where client material goesTo the vendor's serversNowhere. It stays on the machine
Who can change the modelThe vendor, whenever they likeNobody. The file does not move
Works with the network offNoYes
What we can tell a clientWhat the terms permitWhat the machine did
What it costsPer image or per secondElectricity, and the card
Where the risk sitsIn a contractIn the room

What we run, and what each one does

The whole stack sits on one workstation with an RTX 5090 in it. Thirty-two gigabytes of memory on the card, which is the number that decides what will and will not run locally.

The models are wired together in ComfyUI. If you have not seen it, ComfyUI is a canvas of boxes joined by wires. Each box does one step: load this model, take this image, generate this many frames, save the result. It looks like a wiring diagram because it is one, and the useful part is that a workflow can be saved, versioned and rerun exactly.

Two video models do the moving image work. LTX-2.5, from Lightricks, is a 22 billion parameter open-weights model that generates picture and audio together and takes a first and last frame as control, which matters when a shot has to begin and end on specific frames of our own geometry. MiniMax H3 was released with open weights in August and produces video and its stereo sound in a single pass, up to 2K and fifteen seconds a clip. Both had ComfyUI support on the day they were published, which tells you how fast this end of the field is moving.

The honest cost of running them here: a clip that a hosted service returns in seconds takes minutes on our own card, and the machine does one job at a time. We schedule around it. It has not yet been the reason a deadline moved.

The half nobody sees

Not all of this is imagery. We run Qwen3.8 locally to build the studio's internal tools: the scripts that rename and sort frames, check a delivery against what was actually ordered, batch a render set, tidy a project folder at the end of a job. Small, dull, constant work, and a good share of the real time saving comes from it.

It runs locally for the same reason as everything else. Those scripts touch project folders and file names, and file names carry client information. Nothing about that is worth sending to a third party to save two minutes writing a script.

The detail pass

Upscaling has a misleading name. It is not enlargement. The model looks at the frame and re-renders it larger, adding detail that was never in the original pixels. Done carelessly that is precisely the failure we sell against: the model invents a mullion, moves a reveal, gives brickwork a texture the specification does not have.

So the pass is deliberately weak. We run Z-Image Turbo, a 6 billion parameter open model from Alibaba's Tongyi lab, at low strength across frames that are already correct. It is small enough to sit on the card several times over, which is what makes it iterable. The job we give it is narrow: grain, material texture, the micro-detail that separates a render from a photograph. Not geometry, not context, not composition.

The test is the same one we apply everywhere else. Put the result next to the source frame. If the outline moved, throw it away.

Where we still send work out

Local does not win everywhere yet, and pretending otherwise would be the sort of claim this piece is arguing against.

For the highest-detail still work, the closed models are still ahead. Nano Banana Pro, Google's Gemini 3 Pro Image model, holds detail across a complicated frame, renders legible text and signage, and keeps a scene consistent across a stack of reference inputs, at a level the open image models have not reached. Where a brief genuinely needs that, we use it and we say so.

When something does go out, it does not go out through a browser tab. We built a small studio tool for it, private to afterform, which calls the model as a verified API request through fal. The distinction is not cosmetic. A public web interface is a consumer product on consumer terms. An API request sits under a contract with a data processing agreement behind it, retention controls we set rather than accept, and terms under which customer inputs and outputs are not used to train models.

So the rule is not sophisticated. If an open model can do the job, it runs here. If it cannot, it leaves through our own tool rather than through somebody's web interface. And where the work touches client material, that path carries two conditions that are not negotiable: nothing in the request is used for training, and the client knows before anything is sent.

What running locally does not fix

Two things this does not solve, both worth saying out loud.

It does not answer what the models were trained on. Open weights means we can hold and inspect the file. It does not mean we chose the data, and these models were trained at a scale nobody consented to individually. Running the file on our own hardware does not change how the file was made. Our answer to that is deliberately narrow: we do not imitate a living artist's signature style on request, and references are translated into an original direction framework rather than copied. That is written down in the AI policy, which we published rather than kept as a trade secret.

It also does not make the compute free. It moves the energy cost onto our own meter, which at least makes it visible. A generation that takes four minutes on a card in the corner of the room is a cost you can feel, which turns out to be a better discipline than an invoice that arrives at the end of the month.

What local does fix is narrow and worth having. Client material stays where it was sent. A workflow that worked in March still works in September, because nothing was deprecated underneath it. And when a client asks what happened to their drawings, the answer is a fact about a machine in a room rather than a summary of somebody's terms of service.

Where this is going

The direction of travel is not subtle. Two years ago, video generated locally was a novelty and the output was unusable at professional standard. It is now a production tool that ships in client work. Every release moves the line, and the number of boxes that still say hosted keeps falling.

We expect to bring more work in-house as the open models improve, still images first. The moment an open image model clears the accuracy bar our planning and competition work needs, that job comes home too and this post gets an edit.

How this sits inside the rest of our approach is set out in the AI policy, and the accuracy standard it refers to under AI architectural visualisation. The work it produces is at selected work. If you have a project where confidentiality is the deciding factor, tell us what is at stake.

←All of AF_LAB Start a project→