
Research Note: Image Aesthetics, Hallucination, and the Possibility of Closed Models
One of the defining aspects of my practice in Alien Familiar and Phantom Mirror is the role of hallucination in image generation. Working with Midjourney, I often find that the most striking and resonant images are not the ones I anticipated, but those that appear as unexpected artefacts. These hallucinations — fragments, distortions, surreal re-imaginings — mirror the perceptual dislocations of dementia itself.
My editing process is built on refusal. Each generation produces a batch of images, and I move through them saying: yes, no, yes, no. The act of refusal is as important as the act of acceptance. Over time, the images I select create self-referential chains: I use them as style or character references, remixing them into new prompts. This process of continual sifting and re-selection produces a visual language that feels both alien and familiar.
In this sense, Midjourney’s unpredictability is not a flaw but a feature. It provides me with a stream of visual hallucinations that I can edit against. It is, at the moment, the most aesthetically pleasing generator I have used: its colour harmony, composition, and surface textures have a polish that others rarely achieve without significant post-processing.
Yet, there are alternatives. It is technically possible to set up a closed, personal model — something much closer to a studio camera than a street-photography walk. This would involve training a LoRA (Low-Rank Adaptation) or a DreamBooth model on a dataset of my own images. Typically, one might use between 80 and 300 curated photographs, with attention to variety in angles, distances, lenses, and lighting conditions. Captions describing composition and mood guide the training.
The advantage of such a system is consistency. A closed model can reproduce my style with great reliability, applying it across any subject. This could be useful for building a controlled visual ecosystem — for example, ensuring that a character, motif, or tonal palette remains stable over a long project. The trade-off, however, is reduced surprise. The very unpredictability that feeds my practice — the accidental image, the refusal — would be diminished. A LoRA tends to keep producing “my style” rather than hallucinating new possibilities.
This tension is central to my research. On the one hand, a closed model promises authorship and control, an “aesthetic signature” embedded in the weights of the model. On the other, Midjourney’s open, hallucination-driven system aligns more closely with my practice of working through accident, refusal, and remix.
A possible way forward is a hybrid workflow. One stream could remain with Midjourney, generating alien hallucinations that I continue to sift and curate. Another could involve a private LoRA trained on Alien Familiar images, providing coherence when needed. Used together, the two systems would allow me to balance surprise with consistency — the alien and the familiar, in tension.
For now, my practice thrives on the refusal of images. Midjourney remains my most powerful tool, precisely because it gives me what I did not ask for. But the possibility of a closed model raises important questions for the future of the project: what happens when my style becomes a dataset of its own, and the alien begins to reproduce itself?
Roadmap: Setting Up a Closed “Familiar” Model
Twin-track workflow: Alien vs Familiar
- Alien (surprise source): Midjourney as hallucination engine. Batch generation, refusal-as-method, and remixing selected images as references.
- Familiar (coherence source): A small private LoRA trained on your own dataset. Used when you need stability of tone, motif, or palette.
Steps for a closed system:
- Choose a base and runner
- Base model: SDXL-class checkpoint with a refiner.
- Runner: InvokeAI (studio-like) or ComfyUI (node-based).
- Curate dataset
- 120–250 images, varied in composition, distance, and lighting.
- Include textures, motifs, and moods central to Alien Familiar.
- Add captions noting subject, palette, lens cues, atmosphere.
- Select training method
- LoRA: lightweight, portable, usually enough for style.
- DreamBooth: heavier, for full-concept training if needed.
- Textual inversion: optional tokens for highly specific motifs.
- Train and validate
- Iteratively train LoRA with conservative settings.
- Hold out 20–30 images as a validation set.
- Version models clearly (e.g.
AF_LoRA_v01).
- Build generation pipeline
- Resolution: 1024 px.
- Sampler: DPM++ SDE Karras, 28–40 steps, CFG 5–7.
- LoRA weight: 0.6–0.8.
- ControlNet for composition when needed (low strength).
- Upscale: latent 1.5–2×, then external 2×.
- Apply consistent filmic grade/LUT.
- Operational guardrails
- Style card: prompts, negatives, LoRA weights, seeds.
- Ethics note: consent/rights for real people in training set.
- Keep model private; export only final outputs.
- Blend Alien + Familiar
- Run MJ batches and LoRA batches in parallel.
- Curate across both. Let Alien provide surprise, Familiar provide coherence.
Reference Sheet
- Hu, E. J. et al. (2021) ‘LoRA: Low-Rank Adaptation of large language models’.
- Ruiz, N. et al. (2023) ‘DreamBooth: Fine-tuning text-to-image diffusion models for subject-driven generation’.
- Gal, R. et al. (2022) ‘An image is worth one word: Personalizing text-to-image generation using textual inversion’.
- Podell, D. et al. (2023) ‘SDXL: Improving latent diffusion models for high-resolution image synthesis’.
- InvokeAI Documentation – installation, LoRA workflows, model management.
- ComfyUI Documentation – node graphs, ControlNet, SDXL workflows.
- Midjourney Documentation – Style Reference, Character Reference, Style Tuner features.
Footnote: Using Midjourney as a Provisional Alien/Familiar Workflow
While I have not yet set up a closed model of my own, there are ways within Midjourney to approximate the “Alien/Familiar” split in practice. By running some prompts without any controls, I allow Midjourney’s hallucinations to surface freely, producing the alien material that I then sift through by refusal. In contrast, by introducing references or style codes I can gently guide the images toward a more familiar consistency. This never gives me the level of authorship or precision that a closed model would allow, but it provides a provisional way to explore the tension between surprise and coherence directly inside Midjourney.
- Style Reference (
--style ref)
Add an image link to your prompt with--style ref <url>to transfer its aesthetic qualities such as colour, texture, or lighting. This nudges outputs to resemble your chosen visual source. - Character Reference (
--cref)
Use--cref <url>to anchor a specific figure or character so they recur consistently across multiple generations. This is useful if you want a person, motif, or object to appear recognisably in different contexts. - Style Tuner
The Style Tuner is a lesser-known but powerful feature. You invoke it in Discord by typing/tune. Midjourney then generates a grid of images, each representing a slightly different stylistic direction. You select the ones you prefer, and from your choices Midjourney produces a style code — a string of numbers and letters such as--style 12345abcd. You can paste this code into future prompts to consistently reproduce that particular “look.”These style codes are widely shared in the community, but the key point is that you can also make your own. By repeatedly running/tunesessions and saving the codes generated from your preferences, you effectively create a small library of personal aesthetic anchors. This is not as fixed or private as training a LoRA in a closed model, but it allows you to begin shaping a recognisable style signature within Midjourney itself.
Example Workflow: Alien vs Familiar in Midjourney
- Alien (uncontrolled hallucination)
A surreal portrait of memory dissolving into fragments, dementia-inspired dreamscapeNo style references or codes used. The result is unpredictable — you curate by refusal and selection. - Familiar (guided coherence)
A surreal portrait of memory dissolving into fragments, dementia-inspired dreamscape --style ref <link-to-your-preferred-image> --cref <link-to-character-image> --style <your-generated-style-code>Here the image is anchored by a style reference, a recurring character, and your own style code. The outputs remain recognisably in your visual language.
Running both prompts side by side in a batch gives you a two-stream workflow: one for alien hallucination, one for familiar consistency. You can then curate between them to build the hybrid Alien/Familiar vision.
