Convert VTT to SRT

Drop your VTT file: the SRT file, readable by VLC, editing software and video platforms, is created on your device.

100% local processing, nothing is uploaded

How does it work?

Drop your .vtt file: it is read and converted on your device, then the .srt file is downloaded straight away, under the same name. Nothing is sent to a server.

The conversion keeps the timings and the text, numbers the subtitles and replaces the periods before the milliseconds with commas. WebVTT-specific elements are removed: header, comments (NOTE), styles, identifiers, position settings and tags (voices, classes, karaoke timestamps).

The resulting SRT file works with VLC, YouTube, Facebook, Premiere Pro, DaVinci Resolve and most editing software.

Examples

The subtitles of an online course

A 45-minute course exported from a platform as WebVTT (about 700 subtitles) becomes an SRT file usable in editing software, in a fraction of a second.

A timing without hours

WebVTT allows short timings such as 01:02.500: they are completed to 00:01:02,500 in the SRT file, i.e. 62.5 s.

Frequently asked questions

What is the difference between SRT and VTT?
Both contain the same information: a number or identifier, the start and end times, then the text. SRT (SubRip) is the most widespread: YouTube, Facebook, VLC, Premiere Pro, DaVinci Resolve. WebVTT (.vtt) is the format for websites (the <track> element) and Vimeo; it uses a period before the milliseconds instead of a comma, and starts with the line “WEBVTT”.
Accented characters in my file display incorrectly: what can I do?
Older subtitle files are often saved in Windows-1252 rather than UTF-8. The tool detects the encoding automatically (UTF-8, UTF-16 or Windows-1252) and always saves in UTF-8, understood by every recent player: accents and special characters are fixed along the way.
What happens to voice and style tags?
SRT supports neither voices (<v Name>) nor style classes: only the words are kept. If the speaker’s name must remain visible, add it to the text with the “SRT editor”, for example “MARY: Hello”.
Are my video or my text uploaded to a server?
No. Transcription is done by your browser, on your device: your video, its sound and the resulting text are not sent anywhere, unlike most online subtitling tools. Only the speech recognition model (Whisper) is downloaded once, from Hugging Face, with its engine from jsDelivr, then cached. Your work (text and timings, never the video) is saved on your device so you can resume it.