Closed captions are a separate text track the viewer can switch off; open captions are rendered into the picture and cannot be turned off. Here is what that difference means in practice — and how to produce either one from your video's audio.
Manual translation, hired voice actors, days of waiting. Creators and teams lose time and money on every single video.
Skip the old workflowI keep seeing both terms used as if they mean the same thing, and half the tutorials do not say which one they are actually making.
I added subtitles, turned them off to check the video, and they disappeared — then I uploaded to a platform and they were invisible again.
My captions are stuck in the picture and a client wants them gone, or wants them in another language.
Upload, pick languages, and let the AI handle the rest. No software to install.
A few clicks, that is all.
A few clicks, that is all.
A few clicks, that is all.
A few clicks, that is all.
The distinction is about layers, not about wording. Closed captions live in a separate track that travels alongside the video — SRT, VTT or ASS in most containers, a caption stream in broadcast — and because they are a track, the viewer can switch them off, change the font, or swap in another language. Open captions are drawn into the picture itself: they are pixels, like a logo or a burnt-in timestamp, and no player setting can remove them. That single difference decides most practical questions. On a platform where viewers expect control and you might need multiple languages, closed wins. On TikTok, Reels and Shorts, where a large share of viewing happens muted and there is no caption menu, open captions are the only version that always shows. The catch is that open captions are permanent: if the wording is wrong, or you later want the clips in another language, there is no track to edit. Closed captions can be edited in place, and if you later want them visible everywhere they can be burned in — that conversion is one of the things the workbench does, along with generating the caption text from the audio in the first place.
SRT, VTT or ASS sitting next to the video. Viewers can turn them off, restyle them, or pick a different language, and you can edit the wording at any time without touching the picture.
Burned-in, hardcoded, open — all the same thing: the text is part of the image. It always shows, even muted and on platforms with no caption setting, and it cannot be switched off by anyone.
Closed → open is easy: burn the track into the video. Open → closed is not: burned-in text has to be erased from the frames first, then replaced with a track or with new burned-in text.
Closed captions are a separate text track stored alongside the video, so the viewer can switch them on and off or change the language; open captions are rendered into the video frames themselves, so they are always visible and cannot be turned off. The words can be identical — the difference is the layer they live in.
Yes. Open captions, burned-in subtitles, hardcoded subtitles and 'baked in' all describe text that is part of the picture. Any tool that claims to 'turn off' burned-in subtitles is really doing something else — erasing the text from the frames, which is a video edit, not a caption setting.
Closed captions, normally: YouTube has a caption menu, allows multiple languages per video, and uses your caption track for search and for automatic translation. Upload the SRT in YouTube Studio and viewers who want them can turn them on. If you also want them visible in embeds and shorts where they might be muted, publish a second version with the captions burned in.
Open captions. A large share of feed viewing is silent, and viewers rarely open the caption menu, so the version that always shows is the one burned into the picture. That also means you cannot fix a typo afterwards without re-exporting, so proofread before burning.
Yes, and it is the easy direction: take the SRT/VTT track and render it into the video so the text becomes part of the frames. In the workbench you generate the text from the audio, adjust timing and wording, then produce either a caption file (closed) or a video with the captions burned in (open).
Not by a setting — the text is pixels. The realistic route is to erase the burned-in text from the frames, then add a fresh caption track (or re-burn new text). That is what the caption remover does, and it is also how people replace captions when translating a video that only exists with hardcoded subtitles.
We use local storage only to keep you signed in — no tracking cookies, no third-party ads. See Privacy Policy
SpeakVid does not use tracking or advertising cookies. The only storage used is listed below.