Captions burned into the picture survive every platform, every repost and every download. A separate subtitle file does not. For social video that difference decides whether most of your audience reads a word you said, because most social video is watched with the sound off.
The two ways to caption a video
Burned-in captions, sometimes called open captions or hardsubs, are drawn into the video frames themselves. They are part of the picture, exactly like a logo in the corner. Sidecar captions live in a separate file, usually SRT or VTT, which the player reads and draws on top. Broadcast and streaming services use the sidecar approach because it supports multiple languages and lets viewers turn captions off.
| Consideration | Burned in | Sidecar file |
|---|---|---|
| Survives a repost or download | Yes | No, the file is left behind |
| Viewer can turn it off | No | Yes |
| Multiple languages | One per export | Many in one file |
| You control the look | Completely | Barely, the player decides |
| Editable after export | No, re-export required | Yes, edit the text file |
Why social video pushes you towards burning them in
Three things decide it. Social feeds autoplay muted, so a viewer who scrolls past your video sees it before they hear it, and captions are the only thing carrying your message in that first second. Reposting strips a sidecar file, so the moment somebody shares your video or downloads it to post elsewhere, the captions vanish. And platform caption styling is not yours to control, so the same video looks different on each app and sometimes renders in a typeface that fights your brand entirely.
What you give up
Burning captions in costs you three real things and it is worth naming them. A viewer cannot turn them off, which occasionally annoys somebody watching with sound on. One export carries one language, so a second language means a second export. And a typo means re-rendering rather than editing a text file. The first two rarely matter for a short business video. The third is a genuine nuisance, which is why the transcript needs to be editable before the captions are drawn rather than after.
The safe area problem nobody warns you about
Every social platform draws its own interface over your video. Captions positioned by eye in a preview will sit underneath a username, a follow button, a progress bar or a row of icons on at least one platform. The result is a caption you can read in your editor and nobody can read in the feed.
The practical fix is to keep captions inside a safe rectangle well inside the frame, away from the bottom quarter where most platform chrome lives and clear of the right edge where action buttons stack. Yarn positions captions inside that safe area by default and shows the boundary when you move them, so the mistake is visible before you export rather than after you post.
Word-level timing is what makes captions feel alive
Captions that change a whole sentence at a time read as sluggish against fast speech. Timing them at the word level, so each word appears close to when it is spoken, tracks the delivery and holds attention better. This requires transcription that reports where each word falls in time rather than just the text. Yarn uses on-device speech recognition that reports word-level timing, which is also what makes automatic silence trimming possible.
Accuracy still needs a human
Automatic transcription mangles brand names, product names and anything regional. A bike shop called Kerbside becomes curbside, and a supplement called Magnesium Glycinate becomes something unprintable. Any captioning workflow that does not let you fix the transcript before it is drawn into the picture will eventually publish your business name spelled wrong. Edit the transcript first, then render.
The short answer
Burn captions in for social video. Accept that one export means one language and that fixing a typo means re-rendering. Keep the text inside the platform safe area, time it at the word level, and read the transcript before you export.
Last updated 4 September 2026