Start Free Now
Limited Time Offer: Get 50% OFF Starter & Basic Yearly Plans 🎉

Video Marketing Terminology: A Practical Glossary for Creators

Oct 6, 2026

Why Video Vocabulary Became a Shared Language

A decade ago, a video marketer and a video editor could work together for months without agreeing on much vocabulary. The marketer talked about reach and cost per view. The editor talked about LUTs and timeline tracks. The two vocabularies rarely collided, and nobody minded.

That separation has collapsed. Generative video tools now put director-level decisions into the hands of people who think in funnels and conversion rates. A performance marketer can sit down, write a prompt, condition a first frame, set a camera path, export three aspect ratios, and read the retention graph the next morning. Every one of those steps uses terminology borrowed from a different discipline: machine learning, cinematography, broadcast engineering, and growth analytics.

When everyone shares vocabulary but nobody shares definitions, the result is predictable. Briefs get misread. Renders get thrown away. Reporting calls turn into debates about what "view" actually means. Terminology is not trivia; it is the interface between intent and output. If the person writing the brief and the person generating the video use different definitions for "consistency," the deliverable will be wrong even if the execution was flawless.

This guide is a working glossary organized by workflow stage rather than by alphabet. Each section covers the terms you will encounter at a specific point in production, explains why the term exists, and shows where misunderstandings usually happen. Use it to onboard new team members, to align a brief template, or simply to sound less confused in your next review call.

The Generation Layer: Language of Models and Prompts

Before a single frame exists, you are already making decisions in a vocabulary borrowed from machine learning. These terms shape what the model can and cannot do.

Model, checkpoint, and foundation model

A model is the trained system that turns input into video. A foundation model is a large, general-purpose model trained on broad data and then adapted to specific tasks. In practice, the distinction matters because foundation models tend to handle a wide range of prompts acceptably, while smaller specialized models can outperform them inside a narrow style or domain. When someone says a tool "has a strong model for product shots," they usually mean a fine-tune or adapter trained on that kind of footage.

Prompt, negative prompt, and prompt weighting

The prompt is your text instruction. A negative prompt lists what you do not want, such as blur, warped hands, or text overlays. Prompt weighting lets you emphasize or de-emphasize parts of the instruction. The practical rule: positive prompts describe subject, action, camera, lighting, and mood in that order. Negative prompts should be short and specific. A negative prompt that lists forty problems usually creates new ones.

Fidelity, adherence, and artifacts

Prompt adherence measures how closely the output matches the instruction. Fidelity usually refers to visual quality and realism. Artifacts are unwanted visual errors: ghosting, melting edges, flickering textures, or limbs that change shape between frames. Confusing adherence with fidelity leads to wasted iterations. If the composition is right but the texture wobbles, you have a fidelity problem. If the texture is beautiful but the subject is wearing the wrong jacket, you have an adherence problem. Different fixes apply.

Sampler, seed, and steps

A seed is the number that determines the starting noise, and reusing it is the fastest way to reproduce a result. Steps refers to the number of refinement passes; more steps can add detail but also cost time and sometimes introduce noise. Samplers are the algorithms that control how noise is progressively removed. For marketing work, the seed matters most: lock it the moment you get something usable.

Render budget and queue priority

Generation consumes compute, whether it is counted in seconds, tokens, or internal units. Treat it the way a production treats camera time: storyboard on cheap settings, then commit to final quality only for approved shots. Teams that skip this step burn most of their capacity on experiments nobody asked for.

Consistency and Control: The Vocabulary of Continuity

Continuity is where AI video stops being a novelty and starts being a production tool. These terms describe the mechanisms that hold a sequence together.

Keyframes and first/last-frame conditioning

A keyframe is a frame you specify directly rather than generate. First-frame conditioning means the model starts from an image you provide. Last-frame conditioning means you specify the destination and let the model interpolate toward it. Chaining these together is the standard method for building a shot that feels continuous: generate a still, use it as the first frame, capture the output's final frame, then use that as the first frame of the next clip.

Character lock, style lock, and reference images

A character lock keeps a person's features stable across shots. A style lock keeps the visual treatment stable, whether that is a film grain, a color palette, or an illustration aesthetic. Both usually rely on reference images. The practical caveat: locks degrade when the character changes pose dramatically or when the camera angle flips more than about ninety degrees. Plan coverage accordingly.

Motion control, camera path, and drag controls

Motion control describes how the model animates the scene. Camera path specifies where the virtual camera travels. Drag controls let you pull a region of the image in a direction to indicate movement. These tools replace the boom arm and the dolly, but they require the same discipline: one clear movement per shot reads better than three competing ones.

Temporal coherence

Temporal coherence is the consistency of an object across time. It is the single biggest quality differentiator between clips that look professional and clips that look uncanny. When you review a generation, watch a single object rather than the whole frame. If the object holds its shape, coherence is good.

Directorial Lingo That Still Earns Its Place

AI did not retire cinematography vocabulary. It made it more useful, because it is now the most efficient way to describe what you want.

Shot sizes and coverage

Extreme wide, wide, medium, close-up, and extreme close-up describe how much of the subject fills the frame. Coverage means the range of shots you capture of a single scene so an editor has options. Even in a fully generated workflow, planning coverage on paper prevents the classic mistake of delivering six medium shots and no way to cut between them.

Camera movement terms

Pan rotates horizontally. Tilt rotates vertically. Dolly moves the camera toward or away from the subject. Truck moves it sideways. Crane lifts it. Handheld implies instability and immediacy. Push in and pull out are the shorthand most prompts respond to best. Naming a movement precisely is faster than describing the feeling you want and hoping the model guesses.

Lighting and color language

Key light, fill light, and backlight describe the three-point setup. Practical refers to a light source visible in the frame, like a lamp. Golden hour and blue hour describe natural light windows. High-key means bright and low-contrast; low-key means dark and moody. These terms transfer directly into prompts and are far more reliable than adjectives like "cinematic," which has been flattened into meaninglessness.

Aspect ratio and framing

Aspect ratio is width-to-height. Rule of thirds places subjects on imaginary grid lines. Headroom is the space above a subject's head, and lead room is space in the direction of movement or gaze. Framing errors are among the most common reasons a technically impressive generated clip fails in a vertical feed.

Post-Production and VFX Terms You Will Actually Use

Most AI-generated footage still passes through a finishing stage. Knowing these terms helps you specify what you need instead of asking for "make it look better."

Upscaling, interpolation, and frame blending

Upscaling increases resolution. Frame interpolation creates new frames between existing ones to raise the frame rate or smooth motion. Frame blending is the cheaper alternative that mixes adjacent frames and often produces smearing. If your output looks like it was filmed through syrup, interpolation artifacts are the likely culprit, not the model.

Rotoscoping, matting, and tracking

Rotoscoping is manually isolating a subject from its background, frame by frame. Matting is the automated or semi-automated version. Tracking follows a point or region across frames so effects stay attached. Modern tools have reduced the labor here dramatically, but the vocabulary remains essential when you describe what needs fixing.

Compositing, layers, and blending modes

Compositing combines multiple visual elements into one frame. Layers stack those elements. Blending modes determine how layers interact, with screen, multiply, and overlay being the most common. When someone asks for a "glow pass" or an "atmospheric haze layer," this is the vocabulary they are using.

Color grading, LUTs, and log footage

Color correction fixes problems; color grading creates a look. A LUT is a lookup table that maps one set of colors to another. Log footage is a flat, low-contrast capture profile that preserves more dynamic range for grading later. Generated footage rarely arrives as true log, but applying a subtle correction pass before a stylistic grade still improves results.

Delivery Specs: Ratios, Codecs, and Captions

Publishing is a technical act. These are the terms that decide whether your file plays correctly on each platform.

Aspect ratios and platform fit

9:16 is vertical, for short-form feeds. 1:1 is square, useful for carousels and some ad placements. 16:9 is landscape, standard for web players and presentations. 4:5 occupies more vertical space than square while still fitting feed layouts, which is why it often performs well in paid social. Generate or crop for the ratio you will actually publish, and check safe areas so interface elements do not cover your text.

Codecs, containers, and bitrate

A codec is the compression method, such as H.264, HEVC, or AV1. A container is the file wrapper, like MP4 or MOV. Bitrate is the amount of data per second; variable bitrate allocates more data to complex scenes. Higher bitrate is not automatically better once the platform re-encodes your file, so keep source quality high and let the platform handle distribution.

Captions, subtitles, and hooks

Captions transcribe dialogue; subtitles assume translation. Open captions are burned into the video; closed captions are a separate track a viewer can toggle. The hook is the opening moment designed to stop the scroll, and its placement within the first one to three seconds is the single most consequential editorial decision in short-form video.

Performance Metrics Without the Fog

Analytics vocabulary is where marketers and creators most often talk past each other. Here is what the numbers actually measure.

Impressions, reach, and views

Impressions count every time content is displayed, including repeats. Reach counts unique people. Views is platform-dependent: some count a view at three seconds, others at thirty, and some only when the entire video plays. Never compare view counts across platforms without checking each definition.

Watch time, average view duration, and hold rate

Watch time is total time viewed. Average view duration divides watch time by views. Hold rate or retention rate describes the percentage of viewers still watching at a given point. Retention curves are the most honest feedback available: a sharp drop at second four tells you the hook failed, while a gradual decline tells you pacing needs work.

CTR, CVR, and thumb-stop rate

Click-through rate measures clicks divided by impressions. Conversion rate measures completed actions divided by clicks. Thumb-stop rate estimates how many people paused their scroll long enough to engage. High CTR with low conversion usually means the creative overpromised. Low CTR with strong retention usually means the thumbnail or first frame undersold good content.

Attribution, incrementality, and blended metrics

Attribution assigns outcomes to touchpoints. Incrementality asks whether the outcome would have happened anyway. Blended metrics combine paid and organic performance into a single view. Vendors love precise attribution because it looks scientific, but incrementality testing is what tells you whether a channel deserves more budget.

Building a Working Glossary for Your Team

A glossary nobody reads is decoration. A glossary wired into your process is leverage.

Start a decision log

Every time your team resolves a terminology disagreement, write it down with a date, the decision, and a one-line rationale. Did you settle on a two-second hook standard? Did you define "consistent character" as locked across a maximum of three shots? That log becomes the fastest onboarding document you own.

Standardize asset and prompt naming

Adopt a naming convention such as campaign_shot-version_ratio_seed. When someone needs to reproduce a result, the filename carries the seed and the ratio. Teams that skip this end up regenerating work they already approved.

Keep a living prompt library

Store prompts that produced approved shots alongside their settings, negative prompts, and seeds. Over a few months this library becomes the real competitive advantage, because it encodes taste in a reproducible form.

Common Terminology Mistakes That Waste Time

Using "cinematic" as a prompt. It means nothing specific. Replace it with a shot size, a lighting setup, and a lens reference.

Confusing resolution with quality. A 4K clip with motion artifacts looks worse than a clean 1080p clip. Fix motion before you upscale.

Comparing retention across different lengths. A sixty-second video and a fifteen-second video have different retention curves by construction. Compare within formats.

Assuming a view is a view. Check each platform's definition before reporting aggregate numbers.

Treating seed as a setting rather than an asset. Lock it, log it, and reuse it.

Skipping the safe-area check. Text placed near the edges of a vertical frame will be covered by captions, buttons, or profile elements on at least one platform.

Confusing interpolation with slow motion. Interpolation guesses frames. True slow motion captures more of them. The difference shows in any scene with fast movement.

FAQ

What is the difference between a prompt and a negative prompt? The prompt describes what you want; the negative prompt describes what to avoid. Keep the negative list short and specific, because a long list of exclusions often creates new visual problems.

Do I need to know cinematography terms to use AI video tools? You do not need formal training, but naming a shot size and a camera movement produces far more predictable results than vague mood descriptions. It is the difference between directing and hoping.

What does temporal coherence mean in plain language? It means objects keep their shape and identity from frame to frame. Good temporal coherence is what separates footage that reads as professional from footage that reads as uncanny.

Is average view duration or hold rate more useful? Hold rate shows where attention breaks, which makes it more actionable. Average view duration gives you a single summary number. Use both: hold rate to diagnose, average view duration to compare.

How many aspect ratios should I export per video? Export what you will actually publish. A vertical cut, a square or 4:5 cut, and a landscape cut cover most placements. Producing more than you can distribute wastes finishing time.

What is the difference between attribution and incrementality? Attribution assigns outcomes to touchpoints you already track. Incrementality tests whether those outcomes would have occurred without the tactic. Attribution explains, incrementality decides.

How do I keep a character consistent across many shots? Use reference images, lock the seed where possible, and limit the number of dramatic angle changes per sequence. Plan coverage that keeps the camera within a range the model handles comfortably, then regenerate only the shots that break.

Should I upscale before or after editing? Edit first at working resolution, lock the cut, then upscale the final timeline. Upscaling before editing multiplies file sizes and slows every subsequent step for no creative benefit.

What is the most underused term in video marketing? Probably "incrementality." Teams spend enormous energy optimizing attributed numbers that may not reflect real business impact, while a simple holdout test could answer the question in a week.

Where should a newcomer start? Learn three things first: shot sizes, retention curves, and export specs. Those three cover the vocabulary that most often causes failed deliverables, and everything else builds on top of them.

Vocabulary is not gatekeeping. It is compression. A phrase like "medium close-up, low-key lighting, slow push in, 9:16 safe" carries a page of intent in a single line, and it means the same thing to everyone on the call. Master the small set of terms that actually change outcomes, ignore the ones that only sound impressive, and your briefs will get shorter while your results get better.

Alexander

Alexander