Free Automatic Subtitle Generator

Drop a video: speech is transcribed in your browser and split into readable subtitles, then you correct and export them. No server.

100% local processing, nothing is uploaded

How does it work?

Drop your video or audio file, choose a transcription level and the language (or automatic detection), then start. Whisper speech recognition runs in your browser: only the spoken passages are transcribed, thanks to speech detection that ignores silence and music, and every word is timestamped.

The text is then split into readable subtitles: 42 characters per line, two lines at most, displayed for 1 to 7 seconds. An alert flags subtitles that are too fast to read (more than 17 characters per second). These rules can be changed.

Correct them in the list on the right, synchronized with playback: text, start and end, split, merge, find and replace. Finally, export as SRT, WebVTT, text, or an MP4 video with burned-in subtitles.

Examples

A 3-minute interview

A 3-minute interview, with 2 min 40 s of speech, gives about 50 subtitles. At the Balanced level, the model (about 320 MB) is downloaded once; transcription then takes from 30 seconds to 3 minutes depending on the computer. The SRT file weighs less than 5 KB.

A 12-minute tutorial with music

A 12-minute tutorial, with 2 minutes of intro and outro music: only the 10 minutes of speech are transcribed. The software’s name, misheard 14 times, is fixed in one go with “Find and replace”.

Frequently asked questions

Is the transcription reliable?
It relies on Whisper, OpenAI’s speech recognition model, which works well in English and many other languages with a clear voice. Three levels are offered: Fast (a draft to proofread carefully), Balanced and Accurate (the most reliable, which requires WebGPU). Proper names, jargon and noisy passages remain the main sources of errors: the transcription is automatic, so proofread it before publishing. The “Find and replace” feature fixes a misheard name everywhere at once.
How long does transcription take?
It depends on your device and the level chosen; the estimated time is shown before you start. With WebGPU (recent Chrome or Edge on a computer), one minute of speech takes from a few seconds to about a minute. Without WebGPU, only the processor works: it is slower, but it works. The model is downloaded only once (80 to 600 MB depending on the level), then kept by your browser.
Is there a length limit?
Transcription accepts files of 30 minutes at most: beyond that, split the video into parts. The video with burned-in subtitles can last up to 10 minutes, because the MP4 file is created in the browser’s memory. SRT, VTT and text files have no limit.
Are my video or my text uploaded to a server?
No. Transcription is done by your browser, on your device: your video, its sound and the resulting text are not sent anywhere, unlike most online subtitling tools. Only the speech recognition model (Whisper) is downloaded once, from Hugging Face, with its engine from jsDelivr, then cached. Your work (text and timings, never the video) is saved on your device so you can resume it.