Add Subtitles to a Video

Drop your video: subtitles are generated in your browser, synchronized with the speech. Correct them, then export as SRT or burn them in.

100% local processing, nothing is uploaded

How does it work?

Drop your video (MP4, WebM or MOV). Choose the transcription level: the estimated time and the size of the model to download are shown before you start. Speech is recognized in your browser: unlike typical online tools, your video is not sent to any server.

The subtitles appear on the right, synchronized with the video: the current subtitle is highlighted, and clicking a subtitle moves playback to that moment. Correct the text, adjust the start and end, split or merge.

Two ways to add the subtitles: export an SRT file to upload with the video (YouTube, Facebook, LinkedIn, Vimeo), or export an MP4 video with burned-in subtitles, visible everywhere, including on Instagram and TikTok.

Examples

A 5-minute YouTube video

5 minutes of talking to camera give about 80 subtitles. The SRT file (about 7 KB) is added in YouTube Studio › Subtitles: viewers can turn them on or off, and YouTube indexes the text.

A LinkedIn video watched without sound

A 90 s video, with Boxed-style burned-in subtitles, exported in 1080p: the MP4 file weighs about 93 MB (8 Mbit/s). The subtitles stay visible even when the video autoplays without sound.

Frequently asked questions

Is the transcription reliable?
It relies on Whisper, OpenAI’s speech recognition model, which works well in English and many other languages with a clear voice. Three levels are offered: Fast (a draft to proofread carefully), Balanced and Accurate (the most reliable, which requires WebGPU). Proper names, jargon and noisy passages remain the main sources of errors: the transcription is automatic, so proofread it before publishing. The “Find and replace” feature fixes a misheard name everywhere at once.
How long does transcription take?
It depends on your device and the level chosen; the estimated time is shown before you start. With WebGPU (recent Chrome or Edge on a computer), one minute of speech takes from a few seconds to about a minute. Without WebGPU, only the processor works: it is slower, but it works. The model is downloaded only once (80 to 600 MB depending on the level), then kept by your browser.
What is the difference between SRT and VTT?
Both contain the same information: a number or identifier, the start and end times, then the text. SRT (SubRip) is the most widespread: YouTube, Facebook, VLC, Premiere Pro, DaVinci Resolve. WebVTT (.vtt) is the format for websites (the <track> element) and Vimeo; it uses a period before the milliseconds instead of a comma, and starts with the line “WEBVTT”.
Is there a length limit?
Transcription accepts files of 30 minutes at most: beyond that, split the video into parts. The video with burned-in subtitles can last up to 10 minutes, because the MP4 file is created in the browser’s memory. SRT, VTT and text files have no limit.
Are my video or my text uploaded to a server?
No. Transcription is done by your browser, on your device: your video, its sound and the resulting text are not sent anywhere, unlike most online subtitling tools. Only the speech recognition model (Whisper) is downloaded once, from Hugging Face, with its engine from jsDelivr, then cached. Your work (text and timings, never the video) is saved on your device so you can resume it.