What is the difference between subtitles and captions?
Subtitles vs captions — Subtitles transcribe only spoken dialogue for viewers who can hear but not understand the language; captions also include non-speech sounds and speaker labels for deaf or hard-of-hearing viewers.
People use "subtitles" and "captions" interchangeably in everyday conversation, but in professional video production they refer to different things with different purposes. Knowing the distinction matters when you're publishing content for accessibility, localization, or platform compliance.
What are captions?
Captions are designed for viewers who cannot hear the audio—either because they are deaf or hard of hearing, or because they're watching with the sound off. As a result, captions capture not just the spoken words but also relevant non-speech audio: [Phone ringing], [Ominous music], [Crowd cheering]. They may also identify who is speaking when it isn't obvious from the video. Captions appear in the same language as the original audio.
There are two types of captions:
- Closed captions (CC): The viewer can toggle them on or off. This is what the CC button on streaming platforms controls.
- Open captions: Permanently burned into the video frame—they can't be turned off. Common on social media where autoplay happens with the sound muted.
What are subtitles?
Subtitles assume the viewer can hear the audio but can't understand the spoken language—either because it's a foreign film, or because the content has a strong accent or technical vocabulary the viewer finds hard to follow. Subtitles typically transcribe only the dialogue, omitting sound effects and music descriptions. They're often provided as a separate track alongside the original audio and may be displayed in a different language from what's spoken.
Why the distinction matters
The practical difference comes down to audience intent:
| Captions | Subtitles | |
|---|---|---|
| Primary audience | Deaf / hard of hearing | Non-native speakers, foreign language viewers |
| Includes sound effects | Yes | No |
| Language | Same as audio | Often translated |
| Legal context | Accessibility compliance | Localization |
Many countries have legal requirements for captions on broadcast content and—increasingly—online video. The Americans with Disabilities Act (ADA) in the US and the European Accessibility Act in the EU both create compliance obligations for video content in certain contexts. Meeting those requirements specifically means providing captions, not just subtitles.
How they're produced
Both captions and subtitles rely on the same underlying technology: speech to text to generate a transcript, followed by timing alignment to sync each line with the video. For captions, a human editor typically adds sound effect descriptions. For translated subtitles, the transcript goes through a translation step before timing.
AI has made both faster to produce. Automatic subtitle generation platforms can produce time-coded subtitle files (SRT, VTT) from audio in minutes. For video translation workflows, the translated subtitle track is often produced alongside or as an alternative to a fully dubbed audio track.
Platform naming conventions
Confusingly, different platforms use these terms differently. YouTube calls all text overlays "subtitles" in its UI regardless of whether they include sound effects. Netflix distinguishes between "subtitles" (for hearing people) and "SDH" (subtitles for the deaf and hard of hearing, which are functionally captions). When uploading content, check the platform's specific requirements rather than assuming a universal standard.
Which should you produce?
For accessibility, produce captions (with sound effect descriptions). For localization, produce translated subtitles. For social media where the sound is likely muted, open captions burned into the video ensure your message lands even without audio. Many professional productions produce both—a caption track for their home-language audience and translated subtitle tracks for international viewers.