Transcribe Audio to Text
Podcast, voice note, interview: drop the audio file and the speech is transcribed to text in your browser, with timestamps if needed.
100% local processing, nothing is uploaded
How does it work?
Drop your audio file (MP3, WAV or M4A, up to 500 MB and 30 minutes) or a video. Whisper speech recognition runs in your browser: the recording never leaves your device, which suits voice notes, interviews and confidential meetings.
Silence and music are removed before transcription, which speeds it up and prevents text from being invented during gaps. While proofreading, the sound plays along with the text: click a sentence to replay and correct it.
Download the result as plain text, timestamped text, or SRT subtitles if you plan to pair it with a video.
Examples
A 30-minute podcast episode
A 30-minute episode (the tool’s limit), with 28 minutes of speech, amounts to about 4,200 words. The plain text serves as the basis for a blog post or the episode description.
A 2-minute voice note
A 2-minute voice note is transcribed in a few dozen seconds with WebGPU, once the model has been downloaded. The result, about 280 words, can be copied in one click.