Glossary
AI terms, in plain language
Short, honest definitions of the words you'll meet across text, image, video and audio AI — each linked to the tool where it matters.
AI avatar (speaking-head video)
An AI avatar is a synthetic video of a person speaking, generated by animating a still photo with AI-driven lip sync and facial motion to match a provided audio or text script.
Read the full entry →AI image upscaling
AI image upscaling is the process of enlarging a low-resolution image using a neural network that predicts and synthesizes missing detail, producing a higher-resolution result than traditional resampling.
Read the full entry →AI voiceover
An AI voiceover is narration generated by a text-to-speech model rather than a human voice actor, producing spoken audio from written scripts with natural prosody and emotion.
Read the full entry →Speaker diarization
Speaker diarization is the process of segmenting an audio recording by speaker identity, answering the question "who spoke when?" across a multi-person conversation.
Read the full entry →Speech to text (STT)
Speech to text (STT) is automatic speech recognition technology that transcribes spoken audio—from recordings or live microphone input—into accurate, editable text.
Read the full entry →Subtitles vs captions
Subtitles transcribe only spoken dialogue for viewers who can hear but not understand the language; captions also include non-speech sounds and speaker labels for deaf or hard-of-hearing viewers.
Read the full entry →Text to image generation
Text to image generation is an AI technique that creates original images from written descriptions, using models trained on vast datasets of image-text pairs to translate words into pixels.
Read the full entry →Text to speech (TTS)
Text to speech (TTS) is AI technology that converts written text into spoken audio using synthetic voices that sound increasingly like real human speakers.
Read the full entry →Video dubbing
Video dubbing is the process of replacing a video's original audio track with a new spoken performance—usually in a different language—while keeping lip movements plausibly synchronized.
Read the full entry →Video translation
Video translation is the end-to-end process of converting video content from one language to another—through subtitles, dubbed audio, or both—so a new audience can fully understand it.
Read the full entry →