Limited Time Offer: Get 50% OFF your first month of Pro & Ultra plans 🎉

Midjourney Prompt Craft: Storyteller Armor for Better Images

Sep 21, 2026

Most disappointing AI images are not the model's fault. They are the result of a prompt that describes a picture instead of directing a scene. Type something like cinematic warrior, beautiful light, 8k, masterpiece and the model has no choice but to improvise — and improvisation is precisely what destroys consistency the moment you need a second, third, or tenth image that belongs to the same world.

The fix is not memorizing magic words. It is adopting a structure. The name many creators use for that structure is storyteller armor: a layered prompt architecture that protects a story's mood, cast, and visual identity no matter how many frames you generate. Once you internalize the layers, you stop gambling on single outputs and start building libraries of usable images.

This guide walks through the layers, the parameters that control them, worked examples, a three-shot action sequence case study, and the mistakes that quietly sabotage otherwise good prompts.

What Storyteller Armor Actually Means

A storyteller armor prompt does more than request an image. It defines a narrative moment, a visual identity, and a technical contract in the same line of text. Those three jobs are separate, and mixing them randomly is why prompts feel unpredictable.

Think of the structure as three concentric layers:

  1. Narrative layer — who or what is in the frame, what they are doing, and what emotional beat the image represents.
  2. Style layer — the medium, palette, era, and rendering language that make the image recognizable as part of a set.
  3. Technical layer — aspect ratio, resolution intent, camera behavior, and negative constraints.

When all three are present, the model receives a coherent brief. When one is missing, it fills the gap with its own defaults, and those defaults are almost never what you wanted for a series.

A weak prompt: a knight in a forest, epic, dramatic.

An armored prompt: medium shot of a weary knight kneeling in a rain-soaked pine forest at dusk, armor dented and mud-streaked, one hand pressed to a wound, cold blue-grey palette with a single amber lantern glow behind her, painterly cinematic realism, shallow depth of field, 35mm lens look. The second version is longer, but it is also predictable, which is the only thing that matters when you need twenty images that feel like siblings.

The practical test: if you handed your prompt to a different person, could they describe the image you expect? If not, the prompt is not armored yet.

Building the Narrative Layer First

Beginners usually start with style words. Professionals start with the story beat, because style without subject produces pretty wallpaper.

Subject, action, and stakes

Name the subject specifically and give them something to do. "A woman" is a placeholder; "a cartographer in her fifties, rolling a torn map against the wind" is a scene. Action verbs also nudge composition: kneeling, reaching, turning away all imply a camera position and a story direction.

Add one detail that implies stakes. A cracked compass, a bandaged hand, an empty chair. These small specifics are what make an image feel authored rather than generated.

Environment as emotional register

The setting is not decoration; it is tone. Rain, dust, fluorescent light, and golden hour each push the viewer toward a different emotional reading of the same character. When building a series, keep environments coherent: a story set in damp northern forests should not suddenly produce sun-bleached deserts unless the narrative actually moves there.

Framing and point of view

Decide the shot scale before you write style words: extreme close-up, medium shot, wide establishing shot. Shot scale is the fastest way to control how a sequence reads when you assemble images later in a video editor or storyboard deck. Wide shots establish, medium shots carry emotion, close-ups land the beat.

The Style Armor: Locking a Look Across Many Images

This is the layer most people skip, and it is the reason their outputs look like samples from ten different artists.

Style armor is a fixed string of descriptors you reuse verbatim in every prompt of a project. It might include medium (oil painting, gouache, 35mm film still, vector illustration), rendering language (photorealistic, stylized, graphic novel ink), palette (muted earth tones, neon cyan and magenta), texture (grain, halation, watercolor bleed), and era (1970s paperback cover, early digital anime).

Write your style armor once. Save it. Paste it into every prompt unchanged. Then vary only the narrative layer. This single habit does more for visual consistency than any parameter or reference feature.

A sample style armor block:

muted ochre and deep teal palette, soft volumetric haze, painterly realism with visible brush texture, 1970s science fiction paperback illustration influence, subtle film grain, high detail on foreground faces

Two cautions. First, don't stack contradictory mediums — asking for photorealistic gouache confuses the model and produces mush. Second, don't exceed roughly six to eight style descriptors; beyond that, later words start competing with earlier ones rather than refining them.

Directing Light and Camera in Words

Lighting and camera language are the highest-leverage vocabulary in the entire craft. They change an image more than any stylistic adjective.

Lighting vocabulary that works

  • Direction: backlit, side-lit, top-down, underlit, rim light
  • Quality: soft diffused, hard directional, dappled, overcast
  • Source: practical lantern, neon sign spill, window light, firelight, screen glow
  • Color temperature: warm tungsten, cold moonlight, mixed color temperatures

If you only add one lighting descriptor per prompt, choose the direction. Direction determines shape, and shape determines whether an image reads as three-dimensional or flat.

Camera language that works

Borrow from real cinematography: low angle, over-the-shoulder, dutch tilt, long lens compression, wide 24mm perspective, shallow depth of field, deep focus. These phrases give the model compositional instructions that no amount of "beautiful" can replace.

Pair one camera term with one lighting term and you have a controllable baseline you can replicate shot after shot. That pairing is essentially your directorial signature.

Parameters as Directorial Controls

Parameters are where prompt craft becomes engineering. Instead of memorizing a full list, understand what each family of controls does and when to reach for it.

Shape and output intent

Aspect ratio is the first decision, not an afterthought. Vertical ratios suit character portraits and mobile-first social formats; wide ratios suit landscapes, crowd scenes, and establishing shots. Decide the ratio from the delivery format — if the image will sit in a vertical video timeline, generate vertical. Cropping a horizontal image later almost always costs you composition.

Randomness and stylization dials

Most modern text-to-image tools expose a stylization strength and some form of variability control. Low stylization stays closer to literal interpretation; high stylization produces more artistic, less predictable results. For a consistent series, keep stylization moderate and variability low. Save high variability for exploration phases, when you are hunting for a direction rather than executing one.

Reference images and character locking

Reference features — whether called character reference, style reference, or image prompting — are the closest thing to a cast list. Use a single approved portrait as the style or character anchor across the project. Keep the reference itself clean: neutral background, even lighting, no extreme expression. A messy reference teaches the model messy habits.

Negative constraints

Negative prompts are less reliable than positive description, but they help with recurring irritants: extra fingers, text artifacts, watermarks, cluttered backgrounds, distorted perspective. Keep the list short and specific. A five-item negative list tuned to your project beats a thirty-item list copied from a forum.

A Reusable Prompt Skeleton You Can Adapt

Here is the structure in a form you can paste and fill in:

[shot scale] of [specific subject + distinguishing detail], [action verb + emotional state],
[environment with weather/time of day], [one narrative detail that implies stakes],
[style armor: medium, palette, texture, era],
[lighting direction + quality], [camera/lens language],
[aspect ratio intent], [negative constraints if the platform supports them]

Fill it in for a science fiction story beat:

medium close-up of a maintenance engineer in her sixties, silver braid pinned tight, grease on her jaw,
tightening a valve while glancing back over her shoulder, anxious but focused,
narrow service corridor of an aging orbital station, condensation on pipes, emergency strobes at the far end,
muted ochre and deep teal palette, soft volumetric haze, painterly realism with visible brush texture, subtle film grain,
hard side light from a leaking red strobe, 50mm lens look with shallow depth of field,
vertical framing, avoid text and watermarks

That prompt is not poetry. It is a work order. And it will produce usable frames on a much higher percentage of attempts, which is the only metric that matters when you are generating dozens of images for a project.

Case Study: A Three-Shot Action Sequence

Say you need three images that read as a continuous action beat. The temptation is to write three completely different prompts. Instead, keep the style armor and lighting identical and change only the narrative layer.

Shot 1 — establish. Wide shot, subject small in frame, environment dominant, motion implied by posture rather than blur. Purpose: place the viewer in the space.

Shot 2 — engage. Medium shot, subject mid-action, tighter framing, same light direction, same palette. Purpose: raise tension without breaking visual continuity.

Shot 3 — land the beat. Close-up, hands or face, extreme detail, same light direction but one new element of color to signal consequence.

Across all three, the style armor string is copy-pasted. The lighting direction never flips. The palette stays in family. The result reads as a sequence rather than three unrelated images, which is the entire point of the framework.

When something goes wrong here, diagnose by layer rather than rewriting everything. If the shots don't feel related, the style armor drifted. If the composition is boring, the camera language is missing. If the emotion is flat, the narrative layer lacked a specific detail.

Keeping Characters Identical Across Many Shots

Character consistency is the hardest problem in AI image workflows and no single trick solves it perfectly. A layered approach works best:

  1. Create a character sheet first. Generate several portraits, pick one, then generate variations of that approved image at different angles. Treat it as casting.
  2. Anchor with references. Use the approved portrait as a reference input for subsequent prompts whenever the platform supports it.
  3. Freeze descriptive details. Hair length, scar placement, clothing color, and silhouette must be described identically in every prompt. Drifting adjectives create drifting faces.
  4. Keep lighting direction stable. Faces read as different people under different key light directions far more than people expect.
  5. Accept controlled variation. Minor differences in a sequence look natural. Chasing pixel-level identical faces wastes time; chasing recognizable identity is achievable.

A practical habit: maintain a plain text file with your style armor block, your character description block, and your negative list. Every prompt is assembled from those blocks plus a fresh narrative line. This makes consistency a copy-paste operation rather than an act of memory.

A Workflow From Idea to Final Frame

A repeatable pipeline keeps quality high and frustration low.

Phase 1 — Explore. Generate broadly with loose prompts and high variability. Goal: find the visual direction, not final images. Ten to twenty quick outputs.

Phase 2 — Lock the armor. Choose the strongest result. Extract its palette, medium, and lighting into a written style block. Write your character block too.

Phase 3 — Assemble prompts. Build each shot from the skeleton, varying narrative and camera only. Generate three to four candidates per shot.

Phase 4 — Select and refine. Pick the best candidate per shot and refine with small edits — one descriptor at a time. Changing three things at once makes it impossible to learn what worked.

Phase 5 — Prepare for delivery. Decide the target medium early. For stills, upscale the keepers and assemble a contact sheet. For motion, keep composition simple in wide shots so a video model can animate them without artifacts, and plan overlaps between frames for smoother interpolation. Tools such as Sora, Runway, and Kling behave far better with clean, uncluttered source frames.

Phase 6 — Document. Save prompts next to outputs. In three weeks, when a client asks for three more images in the same style, your documentation is the difference between an afternoon of work and a full re-discovery process.

Common Mistakes and How to Fix Them

Writing a paragraph of adjectives with no subject. Fix: name the subject and the action in the first ten words.

Changing the style armor mid-project. Fix: copy-paste the block, always. Add new style words only when starting a new project.

Overloading the prompt. More words do not mean more control. Past a point, later clauses dilute earlier ones. Fix: cut anything that does not change what the viewer sees.

Ignoring aspect ratio until the end. Fix: decide the delivery format before generating.

Treating negative prompts as a cure-all. Fix: describe what you want positively and use only two to five targeted negatives.

Refining too many variables simultaneously. Fix: change one descriptor per iteration so you can attribute the difference.

Skipping the character sheet. Fix: cast your characters before you shoot scenes. It takes twenty minutes and saves hours.

Judging a prompt by one output. Fix: evaluate over five to ten generations. Useful prompts are those that produce good results often, not once.

Quick Reference: A Prompt Checklist

Before you hit generate, confirm:

  • The subject is specific and doing something.
  • The environment implies time of day and mood.
  • One narrative detail signals stakes.
  • The style armor block is included verbatim.
  • One lighting direction and quality are stated.
  • One camera or lens term is present.
  • Aspect ratio matches the delivery format.
  • The negative list is short and project-specific.
  • The prompt was saved alongside the output.

If every item is checked and the image still fails, the problem is usually layer conflict — two style words fighting each other — rather than missing information.

FAQ

How long should a prompt be? Long enough to cover the three layers, short enough that you can read it aloud in one breath. For most scenes, forty to eighty words is the practical sweet spot.

Do I need a different prompt for each platform? The structure transfers; the syntax does not. Parameters and reference features differ between Midjourney, DALL·E, Stable Diffusion, Firefly, and Flux. Keep your style armor block and character block in plain language so they port easily, then add platform-specific parameters at the end.

Why do my images look good individually but wrong together? Almost always a drifting style armor or inconsistent lighting direction. Rebuild both blocks and regenerate the set.

Is a prompt framework going to make my work formulaic? The framework controls consistency, not creativity. Your creative decisions live in the narrative layer, where the subject, action, environment, and emotional beat still belong entirely to you.

How do I improve fast? Iterate in batches of five, change one variable per batch, and keep a written log of what moved the needle. Prompt craft is empirical. The people who improve fastest are the ones who keep notes.

The bottom line: storyteller armor is not a secret word list. It is discipline — separating story, style, and technique, then controlling each one deliberately. Do that consistently, and the model stops being a slot machine and starts behaving like a crew you can actually direct.

Alexander

Alexander