Transcribe Audio to Text

Podcast, voice note, interview: drop the audio file and the speech is transcribed to text in your browser, with timestamps if needed.

100% local processing, nothing is uploaded

How does it work?

Drop your audio file (MP3, WAV or M4A, up to 500 MB and 30 minutes) or a video. Whisper speech recognition runs in your browser: the recording never leaves your device, which suits voice notes, interviews and confidential meetings.

Silence and music are removed before transcription, which speeds it up and prevents text from being invented during gaps. While proofreading, the sound plays along with the text: click a sentence to replay and correct it.

Download the result as plain text, timestamped text, or SRT subtitles if you plan to pair it with a video.

Examples

A 30-minute podcast episode

A 30-minute episode (the tool’s limit), with 28 minutes of speech, amounts to about 4,200 words. The plain text serves as the basis for a blog post or the episode description.

A 2-minute voice note

A 2-minute voice note is transcribed in a few dozen seconds with WebGPU, once the model has been downloaded. The result, about 280 words, can be copied in one click.

Frequently asked questions

Is the transcription reliable?
It relies on Whisper, OpenAI’s speech recognition model, which works well in English and many other languages with a clear voice. Three levels are offered: Fast (a draft to proofread carefully), Balanced and Accurate (the most reliable, which requires WebGPU). Proper names, jargon and noisy passages remain the main sources of errors: the transcription is automatic, so proofread it before publishing. The “Find and replace” feature fixes a misheard name everywhere at once.
How long does transcription take?
It depends on your device and the level chosen; the estimated time is shown before you start. With WebGPU (recent Chrome or Edge on a computer), one minute of speech takes from a few seconds to about a minute. Without WebGPU, only the processor works: it is slower, but it works. The model is downloaded only once (80 to 600 MB depending on the level), then kept by your browser.
Which audio formats are accepted?
MP3, WAV and M4A (AAC) files, as well as the sound of MP4, WebM and MOV videos. iPhone voice memos (M4A) and most recorders work directly. For any other format, convert it to MP3 first.
Is there a length limit?
Transcription accepts files of 30 minutes at most: beyond that, split the video into parts. The video with burned-in subtitles can last up to 10 minutes, because the MP4 file is created in the browser’s memory. SRT, VTT and text files have no limit.
Are my video or my text uploaded to a server?
No. Transcription is done by your browser, on your device: your video, its sound and the resulting text are not sent anywhere, unlike most online subtitling tools. Only the speech recognition model (Whisper) is downloaded once, from Hugging Face, with its engine from jsDelivr, then cached. Your work (text and timings, never the video) is saved on your device so you can resume it.