Captioning and exporting are usually treated as the last two chores of an edit: transcribe the dialogue, then export and hope the file plays everywhere. In practice the two steps are tightly coupled. Burned-in captions change how much bitrate you should spend on fine detail, sidecar caption files change how a player renders timing, and the wrong codec or frame rate can make carefully styled subtitles look blurry, clipped, or slightly out of sync on a phone. What follows is a complete workflow for subtitling video clips and delivering clean MP4 files, with the decision points that actually matter: transcription accuracy, timing rules, caption styling, codec selection, bitrate targets, and the quality checks worth running before anything ships.
Why Subtitling and MP4 Export Are Really One Workflow
Subtitles and encoding decisions push against each other in ways that are easy to miss until a viewer complains.
Caption rendering affects compression. Text edges are high-frequency detail. At a low bitrate, an encoder spends its budget on motion and gradients and treats thin letterforms as noise, which produces shimmering, blocky, or smeared subtitles that are technically present but hard to read. Burned-in captions therefore impose a practical bitrate floor that a caption-free export does not need.
Export settings affect caption layout. Scaling a 1080p sequence down to 720p changes the apparent caption size relative to frame height. Cropping a vertical cut from a horizontal master moves captions into the platform interface zone. Converting frame rate can shift cue timings by a frame or two, which is invisible on a slow interview and very visible on fast dialogue.
The delivery channel dictates both. Some destinations accept a separate caption file, some auto-generate their own and ignore yours, and some autoplay muted with no way to enable captions at all. That last case is the only one where burning text into the picture is genuinely required rather than merely convenient.
The practical consequence: decide delivery targets before you touch transcription or export settings. Everything downstream, from line length to bitrate, follows from that decision.
Caption Formats and Delivery Models Explained
Format choice is often an afterthought, but it determines what styling survives and which players behave predictably.
SRT, VTT, and ASS at a glance
SRT is the lowest common denominator: sequential cue numbers, start and end timestamps, one or two lines of text. Almost every platform accepts it, and it carries no styling at all. WebVTT is the web standard, supports cue settings for position and alignment, and allows styling hooks through CSS. It is the better default for anything played in a browser or a modern app. ASS and SSA support per-character styling, positioning, karaoke effects, and animated text, which makes them popular for burned-in anime and social captions, though support outside dedicated subtitle tools and encoding front ends is uneven. Broadcast pipelines often prefer TTML or IMSC profiles.
Conversion between formats is where quality quietly degrades. Watch for encoding issues (UTF-8 without a byte order mark avoids stray characters), lost positioning when moving from ASS to VTT, and line-break rules that differ between tools. Always re-check the first and last cue after a conversion.
Burned-in, sidecar, or both
A sidecar file is searchable, translatable, toggleable, and free of any bitrate cost, but viewers can switch it off and some platforms ignore it entirely. Burned-in text always appears and works in muted autoplay feeds, but it is permanent, it locks the caption language, and fixing a typo means a full re-export. The pragmatic answer for most creators is both: burn captions into the short social cut where autoplay dominates, and keep a clean master plus a sidecar file for long-form, archives, and future localization.
Building the Caption Pipeline: From Speech to Polished Text
A caption pipeline has four stages, and skipping any of them shows up on screen.
Transcription setup
Automatic speech recognition quality depends more on preparation than on the model. Denoise the audio, apply a gentle high-pass filter, and feed the recognizer a clean mono track at a consistent sample rate. Always set the language hint explicitly rather than letting detection guess on a clip that opens with music. Where the model supports it, load a custom vocabulary with names, product terms, and acronyms; this single step removes most of the tedious corrections. For interviews with multiple speakers, enable diarization so you can attribute lines correctly, and decide early whether you need speaker labels on screen at all. For multilingual clips, transcribe each language track natively rather than translating a transcript, then have a human review the result. Machine translation is a reasonable first pass for a rough cut and a poor final pass for anything published.
Timing and reading speed
Good timing follows a few durable rules. Target roughly 12 to 17 characters per second for general audiences, with an upper bound near 20 for dense technical content. Keep each cue on screen for at least one second and no longer than about six or seven seconds. Leave a small gap, often two frames, between consecutive cues so viewers register a change. Never overlap two cues. Align cue changes with shot changes where possible, because a caption that cuts mid-shot feels like a glitch. Finally, accept that verbatim transcription is usually the wrong goal. Condense redundancy, drop filler words, and preserve meaning over exact wording.
Line breaks and safe areas
Limit captions to two lines and keep each line to roughly 32 to 42 characters depending on font size. Break at clause boundaries, not mid-phrase, and avoid separating an article from its noun. For vertical video, keep text clear of the bottom interface strip and the right-side button column; for horizontal video, keep it above any lower-third graphics. When a frame is busy, a subtle shadow or a semi-transparent backing bar preserves readability without covering faces.
Styling that survives re-encoding
Choose a font weight that holds up after compression, which generally means medium or bold rather than thin. Size captions relative to frame height, commonly around 4 to 5 percent, and test on an actual phone rather than a desktop monitor. High contrast matters more than elegance; if the background is variable, add a thin outline or dark plate. Keep one caption style across a series so viewers recognize your work instantly.
Exporting to MP4 Without Losing Quality
MP4 is a container, not a quality setting. Inside it, codec, bitrate, resolution, frame rate, and audio configuration determine what viewers actually see.
Codec selection
H.264 remains the safest delivery codec: universal playback, predictable behavior, and good hardware support. H.265, also called HEVC, produces noticeably smaller files at equivalent quality but has patchier playback in older browsers and can trigger compatibility headaches. AV1 offers strong efficiency and is well supported in modern browsers, though encoding is slower and older devices may fall back. For a master file, use an editing-friendly intermediate codec and keep the delivery MP4 as a derived output rather than the only copy. When in doubt, H.264 with AAC audio in an MP4 container is the format that plays everywhere without explanation.
Bitrate and quality targets
Constant rate factor gives better results than a fixed bitrate for most exports: for H.264, a CRF between 18 and 23 covers high-quality delivery, with lower values meaning larger files. Use two-pass targeting when a platform imposes a hard size limit. As rough guidance, 1080p delivery sits comfortably between 8 and 12 Mbps, 4K between 35 and 50 Mbps, and social uploads benefit from 10 to 16 Mbps at 1080p because platforms re-encode aggressively. If captions are burned in, add a little headroom so text edges stay crisp through that second encode.
Resolution and frame rate
Never upscale for a delivery export; encode at native resolution and let the platform handle any downscale. Match the source frame rate unless there is a reason to change it, and use a proper conversion method rather than duplicating frames. Variable frame rate footage from screen recorders and phone captures is a common cause of drifting caption sync, so force a constant frame rate before export. If you do change frame rate, re-check cues around fast dialogue.
Audio settings that keep captions honest
Use AAC at 192 to 320 kbps, stereo, 48 kHz, and normalize loudness to your target platform standard, commonly around minus 14 LUFS for web and social. Captions are a backup for comprehension, not a substitute for intelligible dialogue, so fix obvious audio problems before you finalize text.
A Practical End-to-End Workflow
- Lock the picture. Export a single reference master so every later step matches the same timeline.
- Prepare audio for recognition: denoise, normalize, and render a mono reference track.
- Transcribe with an explicit language hint and a custom vocabulary loaded.
- Correct the transcript in one pass, reading for meaning rather than matching every word.
- Split and retime cues, then run a reading-speed pass and shorten anything too fast.
- Style captions consistently and verify safe areas on both vertical and horizontal crops.
- Export one clean master with no burned-in text plus a sidecar caption file in SRT and VTT.
- Export a social cut with burned-in captions and a slightly higher bitrate.
- Watch both exports end to end on a phone with sound off, then with sound on.
- Archive the project file, master, sidecar files, and export settings together.
Quality Control Checklist Before Publishing
The last ten minutes of checking prevent the most embarrassing corrections.
- No cue is shorter than one second or longer than seven.
- No two cues overlap, and no cue sits on screen with zero gap after the previous one.
- Names, brands, and technical terms are spelled correctly throughout.
- Captions never cover a face, a lower-third graphic, or platform interface elements.
- Reading speed is comfortable on a small screen, not just a desktop monitor.
- Text remains legible in the darkest and brightest shots of the clip.
- Audio loudness is consistent between the master and the social cut.
- The sidecar file loads correctly in at least two different players.
- The final MP4 plays without artifacts on both a phone and a desktop browser.
Common Mistakes That Cost Hours
Treating transcription as finished work. Raw machine output is a draft. Punctuation, speaker attribution, and terminology always need a pass.
Exporting at the platform maximum. Higher bitrate does not survive a platform re-encode and only slows your upload. Aim for a quality target, not a size record.
Burning captions into the only copy. You lose the ability to localize, correct, or reuse the text later.
Ignoring variable frame rate. Sync drift that appears in the last third of a clip almost always traces back to a variable frame rate source.
Styling captions before checking safe areas. A style that looks elegant in a 16:9 preview can be cut off in a 9:16 crop.
Using thin fonts over detailed footage. Compression eats thin strokes first.
Skipping the mute test. The single most useful check is watching your own clip with the sound off and seeing whether it still communicates.
Choosing Tools: Decision Criteria
Rather than chasing a specific app, evaluate options against the constraints you actually have. Accuracy in your target languages, including languages with rich morphology where word boundaries are harder, comes first. Custom vocabulary and diarization matter if your content includes names, jargon, or multiple speakers. Format support should include SRT and VTT export at minimum, plus positioning metadata if you publish vertical video. Styling control should let you define reusable presets instead of restyling every clip. Export control should expose codec, CRF or bitrate, resolution, frame rate, and audio settings rather than hiding them behind a single quality slider. Finally, consider batch handling: if you publish several clips a week, automation of transcription and export matters more than any single advanced feature.
FAQ
Should I burn captions in or ship a separate file? Ship both whenever possible. Burn them in for short vertical clips where autoplay is muted and viewers rarely enable captions, and keep a clean master with a sidecar file for everything else.
Which caption format should be my default? Export SRT for maximum compatibility and VTT for web players that support positioning and styling. Keep the project file as the source of truth so you can regenerate either.
How do I stop captions from drifting out of sync? Force a constant frame rate on the source, verify cue timings against the locked timeline, and re-check the final minute of the export, where drift is most visible.
Is a higher bitrate always better for captioned video? Not automatically. Quality targets such as CRF give better results than raw bitrate numbers, and burned-in text needs a modest amount of extra headroom rather than a dramatic increase.
Can I translate captions instead of retranscribing? Translation is fine for a rough internal review, but for published work, retranscribe natively in each language and have a fluent reviewer check both timing and phrasing.
How many lines should a caption have? Two at most. One line is ideal for vertical video, and anything beyond two lines forces viewers to split attention between reading and watching.
What is the quickest way to improve caption quality overall? Load a custom vocabulary before transcription and then run a single reading-speed pass. Those two steps eliminate the majority of complaints.


