2026
LATENT ARCHITECTURE
A Journey Through the Control of Generative AI's Hallucinations
Master's Thesis
2026
Master's Thesis, Prof. Simone Giostra
Politecnico di Milano
ABSTRACT
This thesis investigates the role of generative artificial intelligence in architectural design and asks what becomes of the designer's agency as probabilistic systems are brought into the design process. It addresses three difficulties limiting their adoption: the loss of control associated with opaque, black-box models; the gap between the raster imagery these tools generate and the vector precision of architectural documentation; and the absence of a structured way to evaluate a tool set that changes rapidly. Against these, it categorizes generative tools by function, tests workflows that balance creative exploration with geometric control, and charts a route for AI-assisted design.
The approach is Research through Design, combining literature review, technical analysis, experimental testing, and design practice. Generative tools were first sorted by modality and interface, then assessed across the conceptual, schematic, and realization phases. Comparative testing measured image-generation and image-editing models against four criteria: usability, prompt adherence, consistency, and versatility. These findings were consolidated into an evaluation framework and examined through a case study of an unusual typology, a contemporary Islamic center.
The study finds generative AI most effective as an augmentative partner rather than an autonomous designer. Current models excel at concept exploration, visualization, and rapid iteration, yet remain unreliable for orthographic documentation and precise technical work. The most productive workflows pair human judgment with progressively constrained generation, turning stochastic output into a controlled instrument. On authorship the thesis reaches no final verdict, showing only that it is held or lost according to how the engagement is structured. Its enduring contribution is methodological: a transferable framework for bringing generative AI into practice on terms the architect sets, with control and intent kept in view.
And all research could be summed up in a three-step process:
This thesis is about the use of Generative AI in the Architectural Design process.
I.
State of the Art
II.
Workflow
III.
Demonstration Project
I. State of the Art
Chapters I & II


The technologies discussed here are not parallel fields but nested rings, one sitting inside the next. Artificial intelligence is the outer umbrella, covering all intelligent agents. Within it lies machine learning, which replaces explicit instruction with inductive inference, allowing an algorithm to read thousands of floor plans and derive the rules of architecture itself rather than having a programmer specify the dimensions of a wall. Inside machine learning sits deep learning, built on multi-layered neural networks that learn to represent the world as a hierarchy of concepts, detecting simple edges in the first layers, combining them into shapes, and assembling those into windows or columns further down. Generative AI occupies the innermost ring, a specialized application of deep learning that shifts the objective from classifying existing data to synthesizing new data that plausibly belongs to the same distribution, and it divides again by modality into text-to-text, text-to-image, text-to-audio and text-to-video, with text-to-3D and image-to-3D emerging as the branch that matters most for spatial work.
Useful Terminology:
AI Model: a program trained on data to recognize patterns and generate outputs rather than following hard-coded rules. (Engine)
Prompt: Instructions (input) a user gives an AI model to produce a desired output. (Fuel)
UI: User Interface: The part of a program through which a person interacts with it, such as its screens, buttons, and controls. (Car Body)
LLM: a Large Language Model: An AI model trained on text to understand and generate human language.
Diffusion Model: an AI model that generates images by starting from random noise and refining it step by step into a coherent picture.
Hallucination: a failure state: an instance where a model generates false or nonsensical information


Normally, one might think that typing a prompt into Gemini and getting an image back is a single seamless act, the interface doing the imagining. It is not. Gemini is only a window onto a model, a large trained file of learned patterns that generates the image on a remote server, such as the diffusion model Nano Banana. The same prompt can travel an alternate route through ComfyUI and reach the same kind of model to produce the same kind of output; what changes is not the result but the degree of control the interface exposes. The prompt itself is the input stimulus at the start of both routes, less a description than a set of coordinates steering the model through its latent space toward one specific image.
What about Architects?
“Overall, only 8 percent of firms have implemented AI solutions into their practice, with a further 20 percent currently working on implementing solutions. This is driven significantly more by large firms (50+ employees), the early adopters in this space.” AIA
Architecture is among the fields that theorizes most about AI but doesn’t implement it in practice!


Source: Labor market impacts of AI: A new measure and early evidence
Problems that limits adoption:
Problems?
I.
Slot Machine Effect
Black-box prompting decouples intent from output. The architect drops from creator to curator.
II.
The Integration Gap
Models are trained on images of raster pixels; architecture is built on vector precision. Dimensional amnesia.
III.
Validation Crisis
An AI bubble of competing tools changes Rapidly. Oversaturation breeds analysis paralysis.
II. Workflow
Chapter III



Applications
Image editing models were applied to transform a single conceptual sketch into a wide array of atmospheric variations. By simply adjusting prompts, the sketch evolved from a hyper-realistic visualization into speculative environments (such as a dystopian post-apocalyptic world), allowing the designer to test the emotional resonance of a form under vastly different conditions.
Sketch Realization
Utilizing realtime image generation tools (such as Krea), Basic colored shapes were instantly translated into realized visualizations. As the user draws or modifies shapes and positions, the AI continuously updates a high-fidelity render, offering maximum flexibility and immediate visual feedback during the form-finding process.
Realtime Generation
Utilizing realtime image generation tools (such as Krea), rough conceptual sketches were instantly translated into realized visualizations. As the user draws or modifies strokes, the AI continuously updates a high-fidelity render, offering maximum flexibility and immediate visual feedback during the form-finding process.
Realtime Sketch Realization
Integrating realtime image generation (e.g. Krea) with a parametric Grasshopper model allowed for instant, high-fidelity visualization of algorithmic changes. As the parameters of the 3D model were adjusted, the AI generated realized renders in realtime, effectively closing the feedback loop between abstract parametric logic and tangible visual atmosphere.
Parametric Iteration
A single baseline image was uploaded to an image editing model to create several environmental and atmospheric variations. By adjusting the prompt, the same architecture was visualized at different times of day or within entirely new contextual backgrounds (e.g. moving an urban building into a desert landscape) to test its visual resilience.
Mood Iteration












Using integrated scripting plugins (e.g. Raven), a fully parametric script for a complex staircase was generated in Grasshopper using only a natural language prompt, completely bypassing the manual placement of components. This workflow was also successfully employed to mimic and reconstruct a parametric facade system based purely on a provided reference image.
LLM to Grasshopper







Image to Isometric




Isometric to 3D




UIs























Black Box
UIs
Black box systems are the proprietary, cloud-based platforms built for mass commercial appeal, where the architecture, the training data and the processing are entirely hidden from the user. The architect works through a browser or chat window, typing natural language and receiving a finished image without ever seeing the mechanism that produced it, the mathematics of diffusion deliberately buried beneath a polished consumer interface. The advantage is accessibility: the computation happens on remote server farms, so no specialized hardware is needed, and the models are pre-tuned to yield photorealistic results immediately, which makes them well suited to rapid conceptualization and speculative mood boards. The cost is agency. With no access to the backend, the practitioner cannot enforce specific dimensions, tectonic rules or material constraints, and is left pulling a digital slot machine lever, relying on chance rather than deliberate design. These platforms also run on recurring subscriptions and require confidential project material to be uploaded to third-party servers. Examples cover every modality: Midjourney and DALL-E 3 for images, Runway Gen-2 and Sora for video, ChatGPT and Claude for text and logic, and services such as Luma AI for rapid mesh generation from a text prompt in the browser.









Plugins
UIs
Integrated plugins sit between the two extremes, embedding generative AI directly inside the software the architect already uses. Instead of sending the user off to a separate website, they live as extensions within Revit, Rhino or SketchUp, where they function as context-aware assistants: reading the active viewport, interpreting the camera angle, and analyzing the massing of the 3D model, then using that spatial data as the baseline for generation. The benefit is workflow continuity. Nothing has to be exported, uploaded, generated elsewhere and imported back, and because the AI is tethered to the model itself, the resulting visualizations inherit the correct scale, perspective and geometry, letting material palettes, facade treatments and lighting be tested while modeling is still underway. The cost is dependency. The architect is bound both by the model's limits and by the simplified interface the developer chose to expose, so a hallucinated structural error usually cannot be corrected manually; cloud compute credits keep a meter running, and an update to the host software can break the plugin until a patch arrives. Veras attaches to Revit and Rhino for instant stylized renderings from the active view, ArkoAI and LookX offer similar context-aware visualization inside standard CAD, Raven brings language models into Grasshopper to write parametric scripting nodes from text, and Photoshop's Generative Fill lets entourage and context be inpainted within the usual post-production software.




Local
UIs
At the far end sit the local systems, open-source environments that run entirely on the architect's own hardware. In place of a simplified text box they present a node-based canvas that exposes the whole generative pipeline, with text encoders, samplers, noise schedulers and model weights manually wired together. The black box becomes a transparent switchboard, every variable of the diffusion process open to scrutiny. What this buys is granular control: adapters such as ControlNet allow strict geometric boundaries to be locked down, perspective lines forced, and material application dictated, turning the practitioner from a passive prompter into an active director. Running locally also guarantees data privacy for sensitive projects, removes subscription and compute fees, and opens access to a community library of specialized models that can be customized or trained on a firm's own data. The barrier, however, is steep. Operating these platforms resembles software engineering, demanding familiarity with Python environments, dependency management and neural network architecture, alongside an expensive high-end graphics card to carry the computation. Maintenance is a constant burden, since the pace of community development regularly produces conflicts, broken nodes and unstable updates. ComfyUI is the defining example for image generation and editing, driving models such as Stable Diffusion; LM Studio runs open-source language models like Llama 3 offline for project narratives; local implementations of TripoSR synthesize meshes without cloud dependency; and AnimateDiff nodes animate architectural concepts entirely within the same ecosystem.
Experimental Testing



AI Models


























"Photorealistic daytime wide shot of a modern museum: a single-story rectangular building with a glass curtain wall and exposed concrete frame, set on a clean paved plaza with one bench and two visitors walking by; clear blue sky, soft natural sunlight from the left, crisp shadows, neutral color palette, high resolution, realistic material finishes (reflective glass, rough cast concrete), camera at eye level, 35mm equivalent, slight wide-angle, minimal post-processing."
Image Generation
Realistic - 8 Parameters
Example 01
























"Generate a conceptual axonometric drawing of a cultural pavilion exploring layered public circulation. Use exploded axonometric projection with separated floor plates. Circulation paths highlighted in red linework. Program zones rendered in muted pastel blocks. Structural grid shown as thin grey lines. Transparent façade indicated with dashed outlines. Annotations and diagram labels in small uppercase typography. Background is off-white paper texture. Include shadow cast beneath the exploded layers for depth. Emphasize clarity over realism. Graphic, precise, analytical representation."
Image Generation
Conceptual - 15 Parameters
Example 02
























"Generate a conceptual architectural megastructure hovering over a fractured landscape, composed of interlocking parametric membranes stretched between skeletal frames. Volumes fold, twist, and dissolve into perforated surfaces. Circulation is implied through suspended platforms and hanging capsules. The ground below is split into geometric fissures emitting diffuse light. Structural cables extend beyond the frame, disappearing into void. Interior and exterior conditions blur. Scale is ambiguous. Atmospheric particles drift through the space. Lighting is multi-directional and surreal, casting overlapping colored shadows. Perspective is aerial oblique with exaggerated depth. Composition is asymmetrical yet balanced. Focus on abstraction, diagrammatic clarity, and spatial complexity rather than realism."
Image Generation
Alien - 23 Parameters
Example 03


Image Generation
Evaluation Matrix










"Stage this empty interior in a Japandi loft style while preserving the exact room geometry, window and door positions, ceiling height, and camera angle. Add warm light oak plank flooring. Paint the walls a soft warm beige with one muted clay accent wall at the far end. Install sheer linen curtains on all windows. Add a low modular linen sofa centered in the room with a textured wool area rug beneath it. Place a solid wood coffee table and a slim black metal side table. Add a built-in oak media unit along the back wall. Include a minimalist dining table with four wishbone chairs near the windows. Install sculptural pendant lighting above the dining area and a paper floor lamp near the sofa. Add recessed dimmable ceiling lights with warm temperature. Introduce indoor plants in ceramic pots. Layer natural textures such as rattan, linen, and wool. Add subtle wall art with abstract earth tones. Ensure balanced composition, realistic scale, soft daylight mixed with warm artificial lighting, and high photorealistic material detail."
Image Editing
Staging
Example 01












"Transform the image into a calm winter night scene while keeping the exact same composition, camera angle, perspective, architecture, framing, and all existing people in their precise positions and poses without removing or adding anything. Change the overcast daylight to nighttime with a dark sky softly illuminated by ambient city glow. Add warm yellow light from shop interiors and some residential windows, creating a cozy contrast with cool blue exterior tones. The light should spill naturally onto the wet cobblestone street with realistic reflections. Introduce gentle snowfall, not a blizzard, with visible snowflakes under the lighting and a thin natural layer of snow on cobblestones, rooftops, ledges, and bicycles while keeping textures visible. Keep the people unchanged, with only subtle snow dusting if appropriate. Maintain high photographic realism and cinematic atmosphere."
Image Editing
Time & Weather Change
Example 02




Image Editing
Evaluation Matrix
CONCLUSION



Each tool named is the strongest available for its task at the time of writing; the field moves quickly and specific tools are certain to be overtaken. What outlasts them is the sequence of the six steps, and the kind of tool each calls for.
ROADMAP


The fear that a generative model would turn the architect into a spectator of their own project turns out to be conditional rather than inevitable. An architect can work with a probabilistic tool and still author the result. What decides the matter is the flow of constraint. Give the model little and it returns invention no one asked for; give it the architect's own geometry and it returns something close to what was intended. Authorship survives in the hand that turns that dial, and the single most useful thing to take from this work is the knowledge that the dial exists and can be operated on purpose.
From this follows a less comfortable conclusion. Hallucination is better treated as a material than as a malfunction. The reflex to correct the model's errors is premature, because those errors are worth most at the beginning, when a design is still open and a strange suggestion can carry it somewhere unreached by ordinary means, and they turn into liabilities only later, when precision is owed. EMBER kept five of them. A mass the model misread as a minaret, a skylight it invented over an ablution space, a ceiling it folded where none had been drawn: these were absorbed into the design rather than corrected back toward intention. For a discipline raised on control this is hard to accept. It is nonetheless true that part of the building was authored by keeping what the machine got wrong.
The architect is never a neutral party to this exchange. The tool shapes the choices of whoever directs it even as they constrain the tool, and to deny the influence is only to stop noticing it.
A wider conclusion outruns the project entirely. These tools are fluent in exactly the register architecture has long dismissed as decoration, the rendered and experiential view, and they are clumsy with the orthographic projection the discipline has treated as the real site of design. For the better part of a century the plan was held to generate the building and the perspective merely to sell it. That order is now reversible. A practice could begin in the experiential image, where the tools are strong, and draw its plans and sections afterward, unsettling a hierarchy old enough to pass for nature. Whether the profession takes that road is uncertain; that the road has opened is not.
On the loudest question, whether the architect will be replaced, the answer is for now definite. It will not happen soon, and the reason is precise rather than sentimental. A model can optimize whatever is measurable, daylight and structure and movement, but the claim that it makes a better building falls apart the instant anyone asks what better means, since the answer is cultural and particular and lived, and the model holds all the information about such things while having lived none of it. It is a scholar with no country. The measurable will keep migrating to the machine. The judgment that weighs a building against a way of living will not, and that judgment is the part of architecture least like the work computers were always good at.
None of this arrives without heavy qualification. The tools are shaped by their training, sure with the common and unreliable with the rare, and able to erase the specificity a project rests on, as the upscaling that turned figures in Islamic dress into something else demonstrated plainly. They go stale within months. They still cannot produce a section worth building from, and they remain split between closed systems that are powerful and sealed and open ones that are free and weaker, leaving the architect to choose which compromise to live with. These limits are real. They fence the conclusions above without undoing them.
What is left, once the forecasts and the cautions are set down, is one finding stated without ornament. The machine did not design the building. It enlarged what could be attempted and pressed its collaborator to think harder about intention, yet the building was authored in the decisions it could not make, the choice of what to keep and what to refuse, and the harder question of what a space in this place and for these people ought to be. That is where the discipline still lives. The lasting contribution here is neither a tool nor a model, since both will date within the year; it is a way of keeping that judgment in the architect's hands while the instruments around it grow stronger. Of the two paths, augmentation was always the more demanding. It asks the architect to stay present at every turn, which is precisely how it keeps the architect in the work.
Read Full Thesis:
info@architecture-tareef.com

