Generate an SRT File from a Video

Drop your video: an SRT subtitle file, synchronized with the speech and split according to readability rules, is created on your device.

100% local processing, nothing is uploaded

How does it work?

Drop your video: speech is transcribed locally, word by word, then split into subtitles following professional rules: 42 characters per line, two lines at most, 1 to 7 seconds on screen, with breaks preferably after punctuation and never after an article.

Every subtitle whose reading speed exceeds 17 characters per second is flagged: shorten it or lengthen its display time. The rules can be changed, and the “Re-split” button applies them to the whole file.

Finally, download the .srt file (UTF-8), ready for YouTube, Facebook, VLC, Premiere Pro or DaVinci Resolve, or its WebVTT version.

Examples

An 8-minute tutorial

8 minutes of continuous speech give about 130 subtitles and an SRT file of around a dozen KB. Name it like the video (tutorial.srt next to tutorial.mp4): VLC then displays it automatically.

Fixing the reading speed

A 60-character subtitle displayed for 3 s is read at 20 characters per second: it is flagged. By removing 9 unnecessary characters, it drops to 17 characters per second, the recommended limit.

Frequently asked questions

Is the transcription reliable?
It relies on Whisper, OpenAI’s speech recognition model, which works well in English and many other languages with a clear voice. Three levels are offered: Fast (a draft to proofread carefully), Balanced and Accurate (the most reliable, which requires WebGPU). Proper names, jargon and noisy passages remain the main sources of errors: the transcription is automatic, so proofread it before publishing. The “Find and replace” feature fixes a misheard name everywhere at once.
What is the difference between SRT and VTT?
Both contain the same information: a number or identifier, the start and end times, then the text. SRT (SubRip) is the most widespread: YouTube, Facebook, VLC, Premiere Pro, DaVinci Resolve. WebVTT (.vtt) is the format for websites (the <track> element) and Vimeo; it uses a period before the milliseconds instead of a comma, and starts with the line “WEBVTT”.
How long does transcription take?
It depends on your device and the level chosen; the estimated time is shown before you start. With WebGPU (recent Chrome or Edge on a computer), one minute of speech takes from a few seconds to about a minute. Without WebGPU, only the processor works: it is slower, but it works. The model is downloaded only once (80 to 600 MB depending on the level), then kept by your browser.
Are my video or my text uploaded to a server?
No. Transcription is done by your browser, on your device: your video, its sound and the resulting text are not sent anywhere, unlike most online subtitling tools. Only the speech recognition model (Whisper) is downloaded once, from Hugging Face, with its engine from jsDelivr, then cached. Your work (text and timings, never the video) is saved on your device so you can resume it.