Transcribe Video to Text

Get the text of what is said in a video: a meeting, a lecture, an interview. Speech recognition runs on your device, with no server.

100% local processing, nothing is uploaded

How does it work?

Drop the video: the sound track is extracted and analyzed on your device. Speech detection keeps only the spoken passages, then the Whisper model transcribes them. No file and no text is sent to a server: you can transcribe a confidential meeting or interview.

The text is shown sentence by sentence, synchronized with the video: click a sentence to replay the passage, correct it directly, and use “Find and replace” for misheard proper names.

Export as plain text (sentences run together, with a new paragraph at each long pause), as timestamped text (one line per sentence, preceded by its time), or copy the text to the clipboard.

Examples

A 20-minute recorded lecture

A 20-minute lecture, with 17 minutes of speech, gives a text of about 2,500 words, or 5 pages. At the Accurate level on a recent computer with WebGPU, allow 5 to 20 minutes of processing; at the Balanced level, 3 to 15 minutes.

Finding a quote in an interview

In a 45-minute video interview, the timestamped export shows “[00:31:12] …” in front of the sentence you are looking for: just go to 31 min 12 s in the video to hear it again.

Frequently asked questions

Is the transcription reliable?
It relies on Whisper, OpenAI’s speech recognition model, which works well in English and many other languages with a clear voice. Three levels are offered: Fast (a draft to proofread carefully), Balanced and Accurate (the most reliable, which requires WebGPU). Proper names, jargon and noisy passages remain the main sources of errors: the transcription is automatic, so proofread it before publishing. The “Find and replace” feature fixes a misheard name everywhere at once.
How long does transcription take?
It depends on your device and the level chosen; the estimated time is shown before you start. With WebGPU (recent Chrome or Edge on a computer), one minute of speech takes from a few seconds to about a minute. Without WebGPU, only the processor works: it is slower, but it works. The model is downloaded only once (80 to 600 MB depending on the level), then kept by your browser.
Can I transcribe a video in another language?
Yes. You can select English, French, Spanish, German, Italian and a dozen other languages, or let the tool detect the language automatically. The text is transcribed in the spoken language (no translation).
Is there a length limit?
Transcription accepts files of 30 minutes at most: beyond that, split the video into parts. The video with burned-in subtitles can last up to 10 minutes, because the MP4 file is created in the browser’s memory. SRT, VTT and text files have no limit.
Are my video or my text uploaded to a server?
No. Transcription is done by your browser, on your device: your video, its sound and the resulting text are not sent anywhere, unlike most online subtitling tools. Only the speech recognition model (Whisper) is downloaded once, from Hugging Face, with its engine from jsDelivr, then cached. Your work (text and timings, never the video) is saved on your device so you can resume it.