How to Write Prompts for Photorealistic AI Art and Video
The difference between an AI image that looks like a pretty illustration and one that looks like a photograph is rarely the model. It is the prompt. Photorealism is not a style you can request by typing the word "realistic." It is the result of describing a scene the way a photographer or cinematographer would think about it: subject, light, lens, atmosphere, and the small imperfections that make an image believable. This guide breaks down the anatomy of a photorealistic prompt and gives you techniques you can apply immediately, whether you are generating still images or short video clips.
Treat the Prompt Like a Technical Specification
The biggest mental shift is to stop treating the prompt as a description and start treating it as a technical specification for a virtual camera, lighting rig, and subject. A photographer does not say "take a nice picture of a person." A photographer says what lens, what aperture, what light direction, what background distance, and what mood.
The same discipline applies to AI. Every word in a well-constructed prompt either adds information about the subject, the environment, the lighting, or the rendering technique. Words that do none of those things are usually wasted space. The goal is density: pack each phrase with concrete, visual information.
A useful habit is to write the prompt in blocks: subject block, environment block, lighting block, camera block, quality block. This structure keeps your thinking organized and makes it easy to change one element without rewriting everything.
Structuring Subject Detail for Believable Results
The subject is the center of your image, and photorealism demands specific detail. Generic labels produce generic faces and objects. Instead of "a woman," write "a woman in her sixties with deep smile lines, short grey hair, and a weathered denim jacket." Instead of "a car," write "a dark green 1970s sedan with a cracked side mirror and rain droplets on the hood."
The trick is to include the kinds of details a real scene has: texture, material, age, and imperfection. Real faces have pores, asymmetries, and skin that catches light unevenly. Real objects show wear. A prompt that includes "scratched," "weathered," "faded," or "slightly worn" moves the result away from the sterile, airbrushed look that makes AI images instantly recognizable.
When your subject is a person, be careful with descriptors that fight each other. "Young" and "wrinkled" produce confused results. Decide on the character first, then describe them consistently. If you plan to generate multiple shots of the same character, keep the character description identical across all prompts and change only the scene and action.
Mastering Lighting, the Hidden Driver of Realism
Lighting is the single most influential element in photorealism, and it is the most commonly ignored. A portrait under flat, even light looks artificial no matter how detailed the face is. The same portrait under directional golden-hour light instantly reads as a photograph.
Learn to name light precisely. "Golden hour" means warm, low-angle sunlight with long shadows. "Overcast" means soft, diffused light with muted contrast. "Rim light" means a light source behind the subject that outlines the edges. "Neon" means strong colored light with deep shadows. "Candlelight" means warm flicker and high contrast. Each of these produces a specific, predictable look.
Also specify where the light comes from. "Side lighting from the left," "backlit with a strong halo," or "top-down studio softbox" all give the model a concrete spatial instruction. The more precisely you describe the light source, position, color, and quality, the closer you get to a believable photograph.
Atmosphere is the partner of lighting. Fog, haze, rain, dust, and steam all scatter light and add depth. A scene described as "morning fog in a pine forest, soft beams of sunlight cutting through the mist" contains both lighting and atmosphere, and it will render dramatically better than "a forest in the morning."
Using Camera Language That Actually Changes the Output
Camera terms are among the most powerful tools in a photorealistic prompt, because models have learned them from millions of captioned photographs. The trick is using them correctly.
Focal length changes the look of the image. A 35mm lens gives a natural, wide field of view. An 85mm lens is the classic portrait length, compressing features and blurring the background. A 200mm telephoto flattens perspective dramatically. Saying "shot on 85mm" is not decoration; it produces a visibly different image than "shot on 24mm."
Aperture controls depth of field. "Shallow depth of field, f/1.8" means the subject is sharp and the background melts into blur. "Deep depth of field, f/11" means everything is in focus. This single choice changes the mood of the image more than almost anything else.
Motion language matters for video prompts. "Slow dolly-in toward the subject," "handheld camera with slight shake," "aerial drone shot descending," and "panning right to follow the subject" all produce different camera movements. In video models, these instructions often work as well as they do in image models.
Negative Prompts and the Art of Saying No
Most tools let you specify what the image should not contain. This is your quality control mechanism. The standard list of things to exclude includes blur, distortion, extra fingers, deformed hands, low quality, watermark, text, and oversaturation.
But negative prompts work best when they are specific to your scene. If you are generating a rainy street scene, you might add "no reflections" only if reflections are hurting the composition. If you are generating a portrait, add "no jewelry" if you do not want accessories. Think about what could go wrong in your particular image, and say no to exactly that.
A warning: overloading the negative prompt with contradictory instructions can confuse the model. Keep the negative list short, focused, and consistent with the positive prompt. If you find yourself fighting the model with negatives, the positive prompt is usually the real problem.
Controlling Variation with Seeds and Settings
Photorealism projects often need consistency across a series. Two settings control this: the seed and the step count. The seed is the random starting point of the generation. If you keep the same seed and the same prompt, you get the same image. Change the seed and you get a variation. This is how you explore: fix the prompt, vary the seed, and collect variations until one is right.
When you find an image you like, note its seed. You can then reuse that seed with slight prompt edits to explore "what if" questions while keeping the base composition stable. This is the professional's trick for iterating without starting from scratch every time.
Step count controls how many refinement passes the model runs. More steps usually mean more detail, up to a point of diminishing returns. Most tools have a sensible default, but if your image looks muddy or unfinished, try increasing the steps. If the output is oversharpened or distorted, try reducing them.
Matching the Model to the Job
Different models have different strengths, and prompt technique should adapt. Image models based on recent Flux architectures are excellent for photorealism and respond well to dense, technical prompts. Older diffusion models may need shorter prompts and more reliance on negative prompting.
Video models differ from image models in an important way: they must maintain consistency over time. When prompting a video model, describe the scene and the camera movement, but keep the description of the subject stable and simple. Adding too many details to a video prompt increases the chance that the model drifts or changes something between frames.
Narrative video models, such as the current generation of Sora, Kling, and Luma models, respond to prompts that describe an action with a beginning, middle, and end. "A chef flips a pancake, catches it in the pan, and smiles at the camera" gives the model a clear sequence to generate. Static descriptions like "a chef in a kitchen" produce much weaker videos.
For pre-production, many creators generate a high-fidelity still image first, refine it until it is perfect, and then use that image as the first frame or reference for the video generation. This workflow gives you the control of image prompting with the motion of video models.
Building a Photorealistic Prompt From Scratch
Here is a complete example that ties the techniques together. Suppose you want a photorealistic portrait of an elderly fisherman.
Weak prompt: "a fisherman portrait, realistic."
Strong prompt: "Portrait of a 70-year-old fisherman with deep wrinkles, grey stubble, and salt-worn skin, wearing a faded yellow raincoat and a knitted beanie, standing on a wooden dock, overcast sky with soft diffused light, light sea fog in the background, shot on 85mm lens at f/2.0, shallow depth of field, natural muted colors, photorealistic, 8k detail, slight film grain."
Every block is doing work. The subject block defines the person precisely. The environment block places him on a dock. The lighting block names the overcast sky and soft light. The camera block sets the lens and aperture. The quality block pushes the rendering toward photographic finish.
When you are starting out, build every prompt this way, even if it feels slow. Within a few sessions, the structure becomes automatic, and you will notice your hit rate of usable images climbing dramatically.
Common Failure Modes and How to Fix Them
If your images look plastic or airbrushed, you are probably missing lighting direction and imperfection details. Add a named light source and texture words like "pores," "grain," or "dust."
If your images are cluttered or confusing, your prompt is trying to do too much. Cut the subject to one or two elements and let the environment breathe.
If faces are distorted, reduce the number of simultaneous subject descriptors and strengthen the negative prompt. Extra fingers and warped faces are classic signs of overloaded prompts or insufficient steps.
If video output drifts between frames, simplify the subject description, keep the camera movement modest, and use a reference image as an anchor.
If results are inconsistent across the same prompt, remember that randomness is part of the process. Fix the seed for consistency and vary the seed for exploration.
Frequently Asked Questions
Do I need to know photography to write good prompts?
It helps enormously, but you do not need formal training. Learning a dozen key terms, such as focal length, aperture, depth of field, and lighting direction, will improve your results more than any other single investment.
Why does my photorealistic prompt still produce an illustration-like image?
Usually because lighting is missing or the subject details are too generic. Add a named light source, specify direction, and include material and imperfection details.
Can I reuse the same prompt across different models?
Roughly, yes, but results vary. Models interpret language differently, so expect to tune the prompt when switching tools, especially the negative prompt and quality block.
What is the best length for a photorealistic prompt?
There is no fixed number, but most strong prompts run between forty and eighty words. Density matters more than length: every phrase should add visual information.
How do I keep a character identical across many images?
Use one canonical character description, keep it identical in every prompt, and generate a reference image early. Then use image-to-image with that reference for every scene.
Should I use the word "photorealistic" in the prompt?
It can help as a quality signal, but it is not a magic word. The real work is done by lighting, camera, and detail language. Use "photorealistic" as the final quality block, not as a substitute for structure.
From Prompt to Photograph-Like Results
Photorealistic prompting is a skill you build, not a setting you toggle. The path is the same for every creator who succeeds with it: learn the vocabulary of photography and cinematography, structure prompts in blocks, describe light before everything else, and iterate with seeds until the output matches your intent. The models improve every few months, but the fundamentals do not change. Light, lens, and detail will always be what separates an image that looks generated from an image that looks real.


