Transcribe Video to Text
Get the text of what is said in a video: a meeting, a lecture, an interview. Speech recognition runs on your device, with no server.
100% local processing, nothing is uploaded
How does it work?
Drop the video: the sound track is extracted and analyzed on your device. Speech detection keeps only the spoken passages, then the Whisper model transcribes them. No file and no text is sent to a server: you can transcribe a confidential meeting or interview.
The text is shown sentence by sentence, synchronized with the video: click a sentence to replay the passage, correct it directly, and use “Find and replace” for misheard proper names.
Export as plain text (sentences run together, with a new paragraph at each long pause), as timestamped text (one line per sentence, preceded by its time), or copy the text to the clipboard.
Examples
A 20-minute recorded lecture
A 20-minute lecture, with 17 minutes of speech, gives a text of about 2,500 words, or 5 pages. At the Accurate level on a recent computer with WebGPU, allow 5 to 20 minutes of processing; at the Balanced level, 3 to 15 minutes.
Finding a quote in an interview
In a 45-minute video interview, the timestamped export shows “[00:31:12] …” in front of the sentence you are looking for: just go to 31 min 12 s in the video to hear it again.