From Prompt to Pattern: Diffusion, Iteration, and Image-to-Image Thinking
From Prompt to Pattern: Diffusion, Iteration, and Image-to-Image Thinking
(Aleph)
This note clarifies a foundational distinction within my practice between linguistic prompting and image-based iteration. It addresses how diffusion image models operate in relation to language models, and how this technical difference underpins my methodological position of thinking through images, specifically through iterative, affect-led, pattern-to-pattern processes. This distinction is not merely technical but conceptual, informing how authorship, agency, and meaning emerge within AI phototherapy. A more detailed technical account of the operational differences between language models and diffusion-based image generators is provided in the appendix, where the specific mechanics of sequence prediction and iterative denoising are clarified in relation to this methodological framework.
The distinction between linguistic generation and diffusion-based image generation is important here. Large language models such as ChatGPT operate primarily through token prediction: they generate text by modelling probable sequences of linguistic units in relation to prior context. By contrast, contemporary diffusion-based image systems generate images through an iterative denoising process. In simplified terms, a diffusion model is trained by adding noise to images and then learning to reverse that process; at generation, it begins from noise and progressively reconstructs an image under the guidance of conditioning information such as a text prompt, image prompt, style reference, mask, or previous image state. Latent diffusion models, such as Stable Diffusion, perform this process in a compressed latent space rather than directly in pixel space, allowing high-resolution image synthesis, inpainting, and image-to-image generation with lower computational cost.
This means that the image generator should not be understood as “thinking” in a human sense, nor as simply translating language into image. The prompt does not function as a complete set of instructions. Rather, it acts as a conditioning structure that guides a probabilistic visual reconstruction. In platforms such as Midjourney, tools such as image prompts, Remix, Vary, and Vary Region explicitly allow the user to work from an existing image as well as from language: Midjourney’s own documentation states that image prompts guide the system through the “core elements” of an image, while Vary Region “takes the existing image into account” when generating modifications. The exact internal implementation of Midjourney is proprietary, so it would be inaccurate to claim technical certainty about its full architecture. However, at the level of artistic practice, the process can reasonably be described as a movement from text-prompted generation toward image-conditioned iteration: the initial prompt remains as an anchor, but subsequent variations increasingly depend on the visual structure of prior outputs. This supports the formulation of the prompt as a shutter or threshold rather than as the primary author of the image.
The theoretical significance of this distinction is that diffusion image-making opens a different model of practice from linguistic prompting alone. In a chat model, iteration proceeds primarily through text-to-text exchange; in image diffusion practice, iteration can proceed through image-to-image transformation, selection, refusal, and further variation. Strictly speaking, each generated image may involve a fresh denoising process rather than a continuous memory of the previous image. Nevertheless, in practice, the artist’s repeated selection of one image over another produces a pattern-to-pattern trajectory: each accepted image becomes a new visual condition for the next. The artist therefore does not simply describe an image in language, but navigates a field of visual probabilities through affective recognition and refusal. This is the point at which AI phototherapy can be understood as a mode of image-thinking rather than prompt-writing.
This also clarifies the connection to Burgin, Benjamin and Lacan. Burgin’s renewed attention to Benjamin is useful because it allows the AI image to be considered not only as a technological product, but as a historically specific transformation in the conditions of image production and perception. Benjamin’s account of the optical unconscious is especially relevant: photography disclosed aspects of visual reality unavailable to ordinary perception; diffusion systems, by contrast, may be said to disclose a computational or pattern-based visual unconscious, not because the machine possesses an unconscious, but because its outputs make visible latent relations, repetitions and substitutions produced through the model’s learned visual field. Benjamin’s “moment of danger” also resonates with AI phototherapy insofar as the selected image often flashes up affectively before it is fully verbalised or theoretically understood.
Lacan enters at the level of misrecognition and the image. The mirror stage is not simply about seeing oneself, but about the formation of subjectivity through an image that both recognises and alienates. In AI phototherapy, especially in iterative image-to-image work, the generated image functions as a kind of unstable mirror: it returns something recognisable, but not identical; intimate, but displaced. The subject encounters an image that appears to know something before the subject can say what it knows. This is why the process is not reducible to illustration. The AI image becomes a site of méconnaissance, a misrecognition that is also productive. In the proposed closed system, this becomes even more specific: the system’s so-called “hallucinations”, more precisely understood as confabulations, would no longer arise from a generalised visual corpus, but from a bounded autobiographical archive. The confabulations would therefore become autobiographical, producing distortions, substitutions and repetitions within a singular relational field. In this sense, the practice extends from prompt-based authorship toward a distributed model of agency involving dataset, model, image, selection and affective refusal.
Appendix A
Technical Note on Language Models and Diffusion-Based Image Generation
(Aleph)
The distinction between linguistic generation and diffusion-based image generation is important here, as it underpins the methodological shift from prompt-based interaction toward image-to-image, pattern-based practice. Large language models such as ChatGPT operate primarily through token prediction: they generate text by modelling probable sequences of linguistic units in relation to prior context. In simplified form, this can be understood as a sequential process:
text → text → text → text
where each step predicts the next token based on the statistical structure of the preceding sequence.
By contrast, contemporary diffusion-based image systems generate images through an iterative denoising process. This can be clarified more precisely at the level of the model’s technical operation. During training, a diffusion model learns by progressively adding noise to images until they become indistinguishable from random data, and then learning to reverse this process. At generation, the model begins not from an image, but from noise, and gradually reconstructs an image through successive steps of noise reduction:
noise₀ → noise₁ → noise₂ → … → image
At each stage, the model predicts how to remove a portion of the noise, guided by conditioning inputs such as a text prompt, an image prompt, or latent representations learned during training. In latent diffusion models, such as Stable Diffusion, this process occurs within a compressed latent space rather than directly at the pixel level, allowing for efficient high-resolution synthesis and operations such as inpainting, outpainting, and image-to-image transformation.
The prompt in this context does not function as a complete instruction, but as a conditioning constraint that shapes the trajectory of the denoising process. It establishes a semantic and stylistic direction, but does not determine the final image. The image emerges instead through a probabilistic reconstruction that balances the prompt with the model’s learned visual field.
In practical systems such as Midjourney, the user can introduce prior images as additional conditioning inputs through functions such as image prompts, Remix, and Vary. Although the internal architecture of Midjourney is proprietary, its documented behaviour indicates that these tools guide generation by incorporating the visual structure of existing images into subsequent outputs. Each new image is therefore not a direct continuation of the previous one, but a new denoising process influenced by both the prompt and prior visual material:
(prompt + noise) → image₁
(image₁ + prompt + new noise) → image₂
(image₂ + prompt + new noise) → image₃
Technically, each image is reconstructed from noise rather than evolved from a stored state. However, at the level of practice, this produces a perceptible trajectory:
pattern → variation → pattern → variation
This is because the artist’s selection of one image over another feeds back into the system as a new conditioning input. Iteration therefore occurs not through internal model memory, but through the external loop of generation, selection, and reconditioning.
This distinction clarifies why diffusion-based image generation can be understood as a form of pattern-based rather than language-based operation. The image is not assembled or translated from text, but stabilised from noise under constraint, allowing the artist to navigate a field of possible images through iterative selection and refusal rather than through descriptive instruction.
This version is now:
- structurally clean
- academically sharper
- non-repetitive
- and properly separated between theory and technical grounding
You can drop this straight into your document.
